The portfolio walkthrough is the hour to cut from your design hiring loop. It is the hour everybody protects, because it feels like the real evaluation and because it is the part most hiring managers feel qualified to run. It is also the part that stopped working. A candidate can hand you screens that look like a strong mid-level designer’s best week, and from looking at them you cannot tell whether they made the decisions underneath, whether a tool made them, or whether the person narrating them is the person who made them.
So here is what I run instead, and this is the piece to read if you are about to open a design role and your loop is still built around a deck. I hand the candidate a real, ambiguous problem out of a live backlog and I watch three things: whether they can scope it, whether they can build something that actually runs, and whether they can cut. Ninety minutes, any tool, AI encouraged, paid. Nearly everything I need to make the hire shows up in that window. Almost none of it shows up in a walkthrough.
Over my career I have run about 150 design interviews and hired across a multi-billion-dollar eCommerce org, an agency, and the AI company I am building now. I also run this exercise inside other people’s hiring loops when I am their design leader, which means I have watched it work on teams that are not mine and on problems I did not write. The version below is the one that survived.
What a portfolio used to prove, and what it proves now
Craft used to be scarce, so craft was a signal. If the grids were tight and the type was confident, I could trust the person could execute, because producing that work took hours nobody could fake. The artifact was expensive, and expensive artifacts carry information about whoever made them.
The artifact got cheap. The cleanest controlled evidence we have of that shift is on the engineering side: in a randomized experiment run by GitHub and MIT researchers, developers given Copilot completed an HTTP server task 55.8% faster than the control group (opens in new tab). Design tooling followed the same curve a release or two behind. A junior with good taste and a decent prompt now produces something that reads as senior work at a glance.
That does not mean AI designs well. It means the visible artifact stopped being evidence about the person who submitted it. When I look at a beautiful case study now, I am looking at an output whose cost I cannot estimate, made by a process I cannot see, presented by someone who has rehearsed the story eleven times. There is no version of squinting harder that fixes that.
A walkthrough tests how well someone narrates a finished thing. The hire is a bet on what they do with an unfinished one.
The walkthrough is the least reliable instrument in the room
This is not only my read. Personnel selection research spent decades ranking what actually predicts job performance, and in 2022 Sackett, Zhang, Berry and Lievens found that prior meta-analyses had over-corrected their numbers and re-ran them. The revised ranking is worth internalizing: structured interviews came out on top (opens in new tab) with a mean operational validity of .42, and the top five predictors are all job-specific measures. Structured interviews, work samples, job knowledge tests, empirically keyed biodata, and assessment centers, which the authors describe as a form of work sample.
The same paper is honest about the spread, and so am I. That .42 carries an 80% credibility interval from .18 to .66, which the authors say should be read as "plus or minus .24" rather than a promise. No interview method is a machine that outputs correct hires. The useful finding is relative: methods that make a person do something close to the job beat methods that make a person talk about themselves, and the gap is not small.
A portfolio walkthrough is neither of the winners. It is an unstructured conversation about an artifact you cannot verify, run differently for every candidate, scored on a feeling afterward. Nielsen Norman Group has been telling UX teams the same thing for years: pull from a predefined pool of questions asked in the same order (opens in new tab), rate the answers numerically, and write down in advance what a poor answer and a good answer look like. Most design loops I audit do none of that during the hour they weight most heavily.
The brief I hand them
I use a problem we already solved, so I have a real answer to compare against, and I strip it back to how it looked the day it landed. The candidate gets one paragraph, whatever data actually exists (never enough), and one constraint that quietly conflicts with the goal.
The Story Genie version reads roughly like this. A parent is looking at a preview of a personalized hardcover book about sixty seconds after they arrived, and they have not paid yet. Some illustrated pages come back wrong, and the most common failure is the child’s likeness drifting on a single page. Design what happens next. You have ninety minutes, you can use any tool including every AI you like, you can ask me anything, and at the end I want something running or clickable plus the list of what you decided not to do.
Three rules make it work. It has to be a real problem, because invented ones have no texture and candidates can smell it. It has to be underspecified on purpose, because scoping is the skill being measured and a complete brief hides it. And it has to be paid and bounded, because a ninety-minute exercise you compensate is an assessment, while a take-home that eats somebody’s weekend is a filter for who has a free weekend.
Scope: the first twenty minutes tell you most of it
Watch the first five minutes and do not intervene. Some candidates open a design tool immediately and start pushing rectangles. Some write the problem down in their own words first, in plain language, and then start asking. That single tell has predicted more of my hiring outcomes than any portfolio ever did.
Then count the questions, and note what kind they are. The strong ones ask what happens today when a page comes back wrong, how often it happens, what a regeneration costs us, whether the parent has paid yet, and what we want the parent to feel at that moment. The weak ones ask what style we want, whether there is a design system, and how many screens I expect. The first set is trying to find the shape of the problem. The second is trying to find the shape of the deliverable.
My favorite question a candidate can ask is what happens if we do nothing. Almost nobody asks it. The ones who do are the ones who will save you a quarter of engineering time later, because they treat building as a cost to be justified rather than the default response to a ticket.
Build: it has to run, and I am watching how they use the tools
At the end I want a thing I can click, or a thing that executes. Not a narrated prototype and not a flow diagram. Ugly is fine. Half is fine. The point is that the distance between "I would add a retry path here" and a retry path that exists is exactly where the real decisions live: what the empty state says, what happens on the second failure, whether the parent is asked to choose or told what we did. A designer who only makes screens never has to answer those, because the screen ends before the hard part starts. I have written before about why I bet on the designers who build, and this exercise is where that bet gets tested in ninety minutes instead of ninety days.
The other thing I am watching is how they handle their own tools, because "uses AI" is not a skill anymore. METR ran a randomized controlled trial with experienced open-source developers on their own repositories and found they took 19% longer to finish issues when allowed to use AI tools (opens in new tab), having forecast a 24% speedup going in, and still believed afterward that the tools had sped them up by 20%. The tools help enormously in some spots and quietly cost you in others, and the people using them are bad at telling which is which in the moment.
So the signal is not whether a candidate reaches for a model. It is whether they notice when the model is losing. I have watched candidates spend twenty-five minutes re-prompting for something they could have typed by hand in four, cheerfully, never checking the clock. I have watched others abandon a generation after the second bad result, do it manually, and move on. The second group is the one that ships, and you can only see the difference by watching the ninety minutes happen.
The skill is not using AI. It is noticing, in the moment, that the tool is costing you and putting it down.
Cut: the list of what they did not do is the best artifact
The cut list is the part I read most carefully, and it is the part that has no equivalent in a portfolio review, because portfolios are made of things that survived. I want to know what they considered and dropped, and I want the reason to be about the outcome rather than about time.
Good cuts sound like this: I skipped the settings panel because a parent who is sixty seconds from paying will never open it, and I spent that time on making the fix take one tap. Weak cuts sound like this: I ran out of time before the empty states. A candidate who cut nothing did not scope the problem. They built until the clock stopped, which is the same thing they will do in your sprint.
There is a leadership version of this that shows up in the same list. Ask what they would tell the CEO to stop doing, given what they just learned in ninety minutes. The designers who have opinions about where effort should not go are the ones who become worth more than their salary, because most product waste is not bad execution. It is excellent execution of something nobody needed.
The three questions I ask at the end, in the same order, every time
First: what would you need to know to be sure this is right? I am listening for something measurable and specific, like the rate of likeness failures or the drop-off between preview and checkout. "More user research" is a non-answer. So is total certainty.
Second: what did you delete, and what did it cost? This is the cut list said out loud, and it tells me whether the cuts were decisions or accidents.
Third: where did AI make this worse, and how did you catch it? Everyone who used AI for ninety minutes hit a place where it made something worse. A candidate who says it went great either was not paying attention or is managing me. The ones who can name the exact moment the output went sideways and how they spotted it are demonstrating the supervision skill that the whole industry now runs on.
Then I score four things from one to four before any discussion happens: scoping, build, cut, and how clearly they explained the tradeoff. Written down first, alone, then compared. That is the structure part, and skipping it is how a room full of smart people talks itself into the most charming candidate.
When this is not the answer
This exercise is wrong for three situations, and I would rather you know that than run it badly.
If you are hiring for a role where the artifact really is the deliverable, run the walkthrough. Brand identity, illustration, motion design. A distinct visual voice is not something the tools have commoditized, and the portfolio is a genuine work sample there, not a story about one. Ask how the work was made, but judge the work.
If you are hiring a design leader rather than a maker, this brief measures the wrong muscle. The work sample for leadership is a diagnosis and a plan: give them your actual team, your actual roadmap, and ninety minutes to tell you what they would change in the first month and what they would leave alone. Same structure, different job.
And if nobody in your loop can judge the output, this exercise will give you confidence you have not earned. A founder who cannot tell a sharp scoping question from a plausible-sounding one will read the exercise as a vibe check with extra steps, which is worse than a walkthrough because it feels rigorous. If that is you, get someone in the room who has done the job before you run it, whether that is an advisor, a designer you trust at another company, or a design leader you bring in for the loop. I lay out the order those hires should come in, and who evaluates whom, in how to build a design team. Borrow the judgment before you buy the headcount.
What to change in your loop this week
Keep fifteen minutes of portfolio, and use it as context rather than evaluation. It is a fast way to see what someone has been near and what they care about. Just stop treating it as the decision.
Give the ninety minutes back to the exercise. Use the same brief for every candidate in the same role, in the same order, and pay for the time. Score independently before you talk. Read the cut list before you look at the screens. Ask the three questions in the same order so the answers are comparable across people.
The reason this matters past the hire is that you are not buying screens. You are buying the hundreds of small decisions this person will make in threads and standups that you will never review, most of which decide whether the product is good. The screens are now the easiest part of that to fake and the cheapest part to produce. My longer argument for why the old signal broke is in what I look for when I hire designers now, and the underlying shift is in the gap is not skill anymore.
Run this once and you will feel the difference before the ninety minutes are up. You will stop asking whether you like the work and start knowing how the person thinks when the problem is not solved yet, which is the only state a real problem is ever in when you hand it to them.