The slide says NPS 43, up six since the redesign, and somebody is about to use it to decide whether design gets two more people. Do not decide anything on that number. Net Promoter Score is one survey question about whether a customer would recommend your company, folded into a single figure, and on a design deck it is noise wearing a suit. Tie design to activation, retention, and task success rate. If you cannot measure those yet, put no number on the deck at all and say why. Either of those is more credible than a score that drops when your shipping slips.

The mechanics matter, so here they are once. Customers answer one question on a scale from 0 to 10: how likely are you to recommend this company to a friend. Answers of 9 and 10 are promoters, 0 through 6 are detractors, and 7 and 8 are passives who count in the denominator and nothing else. Subtract the detractor percentage from the promoter percentage and you get a number between -100 and +100. Frederick Reichheld introduced it in Harvard Business Review in December 2003 as the one number you need to grow (opens in new tab). Read that title again. The claim was about company growth. It was never a claim about whether your checkout flow works.

The score answers a question nobody on your design team asked

A customer rating your company is rating your pricing, and your stock availability, and the delivery that showed up two days late, and the support call from March, and the sales rep they like. Your redesign is somewhere in that blend, unweighted and unlabeled. So the week your fulfillment partner misses a window, the design number drops. The week finance runs a promotion, the design number rises. Neither week had anything to do with design, and both of them land on the same chart with a trend line drawn through them.

The arithmetic adds its own distortion. Move a block of customers from 6 to 8 and you have made something genuinely better, and the score does not move a point, because a 6 and an 8 both contribute nothing to the promoter count. Move a handful from 8 to 9 and it swings hard. With the sample sizes product teams actually collect, a few dozen responses in a month, the quarter over quarter change you are reading usually sits inside the noise. Nobody prints a confidence interval next to it, which is part of why it survives so long on so many decks.

It cannot see the thing you changed

Even where the score is picking up something real about the experience, it cannot tell you which part. Nielsen Norman Group is direct about this in their assessment of what a customer-relations metric can tell you about your user experience (opens in new tab): "Usability is never entirely captured by subjective scores. We’ve seen many users struggling to complete a task, yet rating a website as highly as someone who had no difficulty whatsoever."

I have watched that gap open up in person. In a fit study on an apparel platform, seven of eight participants could not agree on what the label standard fit meant on a product page. One word, eight different mental pictures, and a sizing decision hanging off it. Every one of those people liked the brand and most of them would plausibly have handed us a 9. The problem only surfaced because we sat and watched them try to pick a size, and the fix was a clearer label and a plain-language definition, not a satisfaction push.

The number also moves with who you happen to be asking. NN/g cites a 2021 Qualtrics XM Institute study of 17,509 consumers across 18 countries which asked people to rate companies they said they liked. In India and Mexico those scores came back at 60 percent or greater. In Japan and South Korea the same exercise produced negative scores, -47 percent and -11 percent, from people describing companies they liked. Open a market in Tokyo and your number falls. Your design did not change. Your sample did.

A number that moves when your prices move, your shipping slips, or your customer mix shifts is not grading your design. It is reporting the weather.

The people who built it added a second number

The strongest version of this argument is not mine. In November 2021 Reichheld came back to Harvard Business Review with Darci Darnell and Maureen Burns for Net Promoter 3.0 (opens in new tab), and the framing there is blunt: as the system spread, "NPS started to be gamed and misused in ways that hurt its credibility," and "unaudited, self-reported Net Promoter Scores undermined the usefulness of NPS." Their fix was to add a hard metric pulled from accounting results, the earned growth rate, which tracks revenue from returning customers and their referrals and can actually be audited. When the inventor bolts a financial number onto his own score because the score went soft, you should not be resting a headcount decision on the score by itself.

The research record pushed the same direction earlier. Timothy Keiningham, Bruce Cooil, Tor Wallin Andreassen, and Lerzan Aksoy tested the growth claim against longitudinal data in the Journal of Marketing in 2007 and could not reproduce the clear superiority the metric was sold on. Writing up the wider picture in MIT Sloan Management Review (opens in new tab), the same four put it this way: "The best metrics have shown only modest correlations to growth, and none of them have shown themselves to be universally effective across all competitive environments." Modest and conditional, at the level of a whole company. Two layers down, at the level of one screen your team redrew, there is nothing left to read.

Three numbers that belong on a design deck instead

Activation first. Pick the moment your product becomes useful to a new person and measure the share of new users who reach it inside a fixed window. First invoice sent. First report shared. First order placed. Design owns most of that path, because nearly everything between signup and that moment is an interface decision. Name the moment out loud before the work starts, in one sentence, with the window attached. Teams that skip that step spend the readout meeting arguing about the definition instead of the result.

Retention second. Take a cohort, define the core action, and report the share still doing it at week four and week twelve. Retention catches the design that demos beautifully and wears badly, which is a failure mode no satisfaction survey has ever caught in time. It is slow, and that frustrates people, so run it next to activation. Activation tells you this month whether the change is working. Retention tells you next quarter whether it held.

Task success rate third, and this is the one most design teams are not reporting and should be. It is the percentage of participants who can actually complete a defined task. Jakob Nielsen and Raluca Budiu call it the simplest usability metric (opens in new tab), and their case for it is the one I would put in front of a CFO: "if users can’t accomplish their target task, all else is irrelevant. User success is the bottom line of usability." Run the same tasks with the same script before and after. Score partial success in named levels rather than pretending every task is binary. Report the confidence interval. Five to eight participants a round is enough to steer with, as long as you say so plainly instead of dressing eight people up as a statistic.

Those three share something NPS does not have. Each is attached to a specific decision, which means each one can tell you that you were wrong. That is the whole test of a metric worth reporting. How to build the case around them without overclaiming is in why most design ROI decks lie, and the baseline and attribution mechanics are in how to measure the ROI of design.

What this looks like on real work

On the product page that carries most of the revenue at a multi-billion-dollar apparel platform, we benchmarked time to completion against a new off-canvas inventory drawer and a filter-by-color pattern rather than asking people how they felt about it. About half the tested customers preferred the drawer, which is a soft number and I treat it as one. The number that carried the recommendation was how long it took someone to answer the question they arrived with, which is whether the thing they want exists in their size and their color. That work is handed to engineering and in testing, not in customers’ hands yet, and I would rather tell you that than round it up.

When I built Story Genie I measured whether people reached the finished book reveal and then bought, not whether they enjoyed themselves along the way. Enjoyment was the hypothesis. Reaching the reveal and paying was the evidence. Had I run a recommend-us survey on that product I would have collected a pleasant number from the people who already liked it and learned nothing at all about the people who left at step three.

Report the number that would have changed your mind if it had gone the other way. A satisfaction score almost never would have.

Who each path is for

If you run design and there is no scorecard at all, this is a quarter of work, not a platform migration. One activation number, one retention cohort, one task success benchmark on your highest-traffic flow. Your analytics tool already holds the first two. The third needs five participants and a script you write in an afternoon.

If NPS is mandated across the company by a CX team or a board, do not fight it and do not try to kill it. Keep reporting it where it belongs, as a relationship number for the whole company, and stop letting it show up in the design section. The sentence that settles this in a meeting is short. NPS tells us how the relationship is doing, these three tell us whether the product got easier to use, and only the second set responds to what my team ships.

If the real problem is that nobody senior owns the standard, that no one in the building can say what design is accountable for this quarter, another dashboard will not fix it. That is a leadership gap. It is the specific thing a fractional head of design is for: a few days a month to set the measures, wire them to the roadmap, and leave behind a scorecard your team runs without that person in the room. The way I make the case to a finance-minded audience is in how I prove design actually moved the business.

And if you truly cannot instrument anything this quarter, put no number up. Bring three recorded sessions of customers failing at something, plus the date by which you will have a baseline. Every good operator I have worked with respects we do not measure this yet, here is the plan. None of them respect a number that comes apart under one question.

When NPS is the right number

There are places it earns the slide, and I would rather name them than pretend the rule is universal. The first is B2B account management with named customers. When the same buyer answers every quarter and your account team calls the detractors by name the next morning, the score is doing real work as a trigger for a conversation rather than as a grade. The comment underneath is the payload. The number is just the routing.

The second is a long-horizon company trend with a large, stable sample. Tracked over years, at the whole-company level, with the country mix held roughly constant, the line does say something about the health of the relationship. It is a slow instrument for a slow question, which is close to what it was built for.

The third is the free-text box. The written answer under the score is often the most useful customer research a company collects, and usually nobody reads it. Tag it, sort it by theme, and you have a standing list of what to test next. NN/g, for their part, still recommends using NPS, and I agree with how they frame it: it "correlates well with perception of usability, is easy to understand and administer, but has limitations for understanding and evaluating UX when used in isolation," and their own advice is to collect behavioral metrics such as task success rates and task times alongside it. Alongside. Not instead of, and never as the grade on design.

Here is the test I would put to any design leader presenting to you, and I would want it applied to me. Ask what result would have told them the work failed. If the only number on offer is a satisfaction score, they have a mood, not a case. If the answer is activation fell, or the task took longer, or the cohort did not come back, you are talking to someone who ties design to the business and is willing to be held to it. That is the person worth hiring, and it is the standard I want to be judged by.