Taste is the ability to look at real work and correctly name what is wrong with it, in the order that matters. That is a skill, it varies enormously from designer to designer, and you can measure it in forty minutes. If you are about to make a design hire and your loop is a portfolio review plus a conversation about values, you are about to spend a few hundred thousand dollars a year on the one signal you never actually checked.

So here is the stance, and it is the point of the piece. Taste is testable. A scored critique of real product work predicts what a designer will produce for you better than any portfolio they can show you, because a portfolio is a record of outputs and a critique is a look at the judgment that generates them. I have run this exercise as the hiring manager and as the design leader dropped into someone else’s loop, across a multi-billion-dollar eCommerce org, an agency, and the AI company I am building now. It is the fastest read I have on whether a person will make the product better or only make it different.

What a portfolio can and cannot tell you

A portfolio is a set of outcomes that survived. It was made under conditions you cannot see, with help you cannot measure, edited by someone with every reason to show you the good version. That is not a criticism of candidates. It is the format working exactly as intended.

The thing you are actually buying is different. You are buying the several dozen judgment calls this person will make every week that you will never review: what to fix first, what to leave alone, when the layout is fine and the copy is the problem, when the whole flow has one step too many. Those calls compound into the product. A portfolio shows you six finished outputs from a function you never get to see. A critique shows you the function.

Those calls have a price tag, and it is a large one. Baymard Institute keeps a running average across fifty studies and puts documented cart abandonment at 70.22 percent (opens in new tab), with the average large ecommerce site able to gain a 35.26 percent lift in conversion through better checkout design alone. Nobody closes that gap with one brilliant redesign. It closes through a long sequence of small correct judgments about which friction matters, which is the exact capacity a critique measures and a portfolio hides.

A portfolio shows you six outputs from a function you never get to see. A critique shows you the function.

The exercise: forty minutes, your own product, ranked

Pick a real flow from your live product. Something shipped, something you have data on, ideally something you already know is imperfect. Not a competitor’s app, not a screen from Dribbble, not a fictional brief. Real work carries texture that invented work does not, and you need a ground truth to score against.

Then ask three questions, in this order, and keep the order fixed for every candidate. What is wrong here? Rank those problems by what they cost the user or the business. Take the top one, and tell me what you would do instead and what that change would cost us.

Give them ten minutes alone with it, writing, before anyone talks. The silent pass matters because the moment you start discussing, you start leading, and a good candidate will read your face and follow it. Then twenty minutes of conversation on the list they wrote. Then ten on the fix.

One rule for the room. They are not redesigning anything. No pushing pixels, no wireframes. If a candidate cannot hold a problem in their head long enough to rank it, moving to a canvas is an escape, not progress.

Write the answer key before they arrive

This is the step every team skips, and skipping it is the difference between a measurement and a mood. Before you run the exercise on a single candidate, run it on yourself and two colleagues who know the product.

Nielsen Norman Group has been prescribing the same shape for inspection work for decades, and it transfers cleanly. For a heuristic evaluation they advise that three to five people should independently evaluate the same interface (opens in new tab), because any single reviewer, however expert, misses things. Have each person list the problems alone, then merge the lists.

Then rank by severity, and borrow a scale that already exists rather than inventing one. NN/g rates usability problems from 0 to 4: 0 is "I don’t agree that this is a usability problem at all," 1 is cosmetic, 2 is minor, 3 is major, 4 is a usability catastrophe. The same guidance is blunt about reliability, noting that severity ratings from a single evaluator are too unreliable to be trusted (opens in new tab) and that the mean of three evaluators is good enough for most practical purposes. If that is true of your own team rating a screen they built, it is certainly true of one hiring manager rating a stranger.

Now you have a ranked key, and the exercise becomes scoreable. Four things to compare: how many of your top three the candidate found, how close their ranking is to yours, whether they attached a cost to each item instead of a preference, and whether they surfaced anything your own team missed. That last one is the strongest positive signal in the whole exercise, and it is the reason to run it on work you know well rather than work you can bluff.

The tells, positive and negative

They rank by cost, not by taste in the decorative sense. Weak critique sounds like a list of preferences: the blue is off, the spacing is tight, the icons are inconsistent. Strong critique sounds like this: nothing on this screen tells me what happens after I press continue, so people who are unsure will stall here, and that is where your drop-off is. Both people noticed things. Only one of them told you what to do Monday.

They name the tradeoff before they call it a mistake. The designers I want say some version of “somebody chose this to protect the signup rate, and I would still change it, here is why that trade is wrong now.” That sentence tells you they assume competence in people they have never met, which is one of the clearest early reads on whether they will work well with your engineers.

They see the boring layer. Contrast, hierarchy, hit targets, focus states, the label that is missing. This stuff is unglamorous and it is where most real damage lives. The 2026 WebAIM Million scan of a million home pages found low contrast text on 83.9 percent of them (opens in new tab), the most common detected failure by a wide margin. A designer with taste flags illegible text without running a tool, and one who walks past it while discussing brand expression has told you something about what they will ship.

The negative tells are just as fast. The candidate who only proposes a global rewrite, rethinking the whole information architecture, is often avoiding the commitment a ranked list requires. The candidate who will not name the worst part of the screen in an interview will not name it in your design review either, and you are hiring partly for that sentence. And the candidate who lists twenty problems without ordering them has shown you they will treat every ticket as equally urgent, which is how roadmaps die.

Anyone can produce a list of twenty problems. Taste is knowing which three of them you would actually spend the quarter on.

Score it like an instrument

Score four dimensions from one to four, written down alone before anyone in the loop talks: coverage against the key, ranking, reasoning, and the quality of the proposed fix. A fix earns a four when it is specific, bounded, and comes with an honest cost. It earns a one when it is a direction rather than a decision.

Use the same flow for every candidate in the same role. Rotating the artifact feels fair and quietly destroys the only thing that makes the numbers mean anything, which is comparability. And do the scoring before discussion, every time, because otherwise the loudest person in the debrief becomes the rubric.

This pairs with, rather than replaces, a build exercise. Critique measures whether someone can see correctly. A ninety-minute problem out of your backlog measures whether they can scope, build, and cut, which I wrote up separately in the interview I run instead of portfolio walkthroughs. If you only have room for one, run the build exercise for makers and the critique for anyone senior, because the higher the level, the more of the job is judgment about other people’s work.

Where taste tests break, and the wine judges who show why

The obvious objection is that judging judgment is circular, and there is real evidence behind it. Robert Hodgson ran replicate samples through a major U.S. wine competition from 2005 to 2008, slipping three pours from the same bottle into flights of thirty for panels of four expert judges. About 10 percent of the judges replicated their own score within a single medal group (opens in new tab), and another 10 percent, on occasion, scored the same wine anywhere from bronze to gold.

Read that as a warning about method, not about expertise. Those judges were scoring an overall impression, alone, against nothing, on a scale with no shared definition of what each point meant. Every one of those conditions is fixable, and an unstructured critique interview reproduces all four of them. The answer key, the fixed questions, the severity definitions, and the independent scoring exist precisely because taste assessed as a general impression is noise, and taste assessed against a known problem set is a measurement.

Two more ways this goes wrong. Choose a flow a smart adult can understand in two minutes, because a critique of a domain that takes three weeks of context measures onboarding speed, not taste. And keep it short, paid, and transparent. Tell candidates in advance what you are measuring and on what scale. You want their real judgment, and stress is very good at hiding it.

When the critique is the wrong instrument

For brand identity, illustration, and motion, the artifact genuinely is the deliverable, and a distinct visual voice is not something you can interrogate out of someone in forty minutes. Judge the work, ask how it was made, and skip this exercise.

For very junior hires, weight it differently. Taste grows from reps, and a new graduate can have real instincts with none of the vocabulary to defend them. They will underperform the critique while being a good bet. For juniors I care more about the proposed fix than the ranked list, because the fix shows raw judgment without requiring the language of a design review.

And the hard one. If nobody on your side can write the answer key, this exercise will hand you false confidence, which is worse than no exercise at all. A loop that cannot produce and defend a ranked list of its own product’s problems will score candidates on agreement, and agreement with a team that does not know is not a signal. That is a real and common situation, especially at a company making its first design hire, and the fix is to borrow judgment before you buy headcount. Bring in an advisor, a design leader you trust at another company, or someone in the seat part time. Writing that key with a team and sitting in their loop is a normal first month of work for me as a fractional head of design, and it is usually cheaper than one mis-hire.

What to change in your loop this week

Keep the portfolio and cut it to fifteen minutes, used as context rather than evaluation. It is a fast way to see what someone has been near and what they care about. It is not the decision.

Add the forty minutes. Same flow for every candidate, key written first by three people, questions in the same order, scored alone before anyone speaks. You can build the whole thing in an afternoon and it will outlive the role you built it for.

The reason this matters beyond the hire is that the market shifted underneath the old signal. Tools now produce a competent-looking screen for anyone, so the artifact stopped carrying information about the person who submitted it, which I argued at length in what I look for when I hire designers now and in the gap is not skill anymore. What did not get commoditized is knowing which screen was worth making. That capacity has a name, people have been calling it a vibe for thirty years, and it takes forty minutes and one honest rubric to measure it.

Run it once on someone you have already hired and rated. If the score matches what you know about how they work, you have an instrument. If it matches better than the interview you have been trusting, you have your answer.