If you are deciding where design hours and design money go next quarter, here is the call sitting in front of you: how much of that should still be aimed at making things, now that making things is close to free. My answer is almost none of it. AI collapsed the cost of producing work that looks right. It did not touch the two things that were always the hard part, deciding what is worth building and standing behind the thing after it ships. That is where the value moved. Your hours and your budget should follow it.

This is not a prediction. I ran an AI-first transformation inside a multi-billion-dollar eCommerce design org, and I founded and run an AI product where agents write and illustrate personalized hardcover kids books with a human supervising every one. So what follows is a map drawn from inside the work: where these tools are genuinely strong, where they fall over, and what a leader should do with that.

What AI is genuinely good at, and I mean genuinely

Start with the half that is easy to skip past, because pretending the tools are worse than they are will cost you more than overhyping them. AI is very good at the middle of the work. Drafts, variations, synthesis, production, the polish pass. Give it a direction and it returns twenty starts instead of one. Give it a pile of transcripts and it returns something you can actually read. Give it a component and it returns the eighth state, the locale version, the export.

One example from real work. At a multi-billion-dollar apparel platform, product descriptions for a launch were written by hand, roughly three to six hours each, across about a hundred products at a time. I built an agent that takes the source data the team already receives and writes the full set in brand voice, at the right length, in language a buyer actually understands. What ran in the hundreds of hours now runs in about one. And the bigger win was not the clock. Before, many different writers with no defined process produced descriptions that ranged from twenty words to two hundred, some in all caps. The output got consistent, so the page got more trustworthy. That is a real user experience improvement delivered by a machine doing the middle of the job.

So the floor came up. Anyone with the tools can now produce something that looks credible, reads cleanly, and ships. If your competitive position rests on producing screens faster than the next team, that position is already gone, and no amount of craft nostalgia gets it back.

The first thing it cannot do: decide what is worth building

A model will answer any brief you hand it, including the wrong one, with total confidence and no flinch. It optimizes for plausible, not correct. It has no stake in your customer, no memory of the thing you tried in 2023 that failed, and no instinct that the request in front of it is a symptom rather than the problem. That gap does not close by prompting harder, because the missing input is evidence about real people and a point of view about what matters.

Nielsen Norman Group looked hard at the tempting shortcut here, the idea of interviewing AI-generated "users" instead of finding real ones. Their conclusion in Synthetic Users: If, When, and How to Use AI-Generated "Research" (opens in new tab) is direct: synthetic users cannot replace the depth and empathy gained from studying and speaking with real people, and the errors they introduce are the kind you cannot detect without doing the real research anyway. AI can organize evidence you already have. It cannot manufacture evidence about the specific, messy experience of your customers.

Now put that next to the new economics. The cost of making fell. The cost of making the wrong thing did not. It went up, because a team that can ship ten features a quarter instead of two now carries ten support surfaces, ten sets of edge cases, ten chances to teach a customer that your product is confusing. Cheap production multiplies whatever decision quality you already had. Good decisions compound faster. Bad ones compound faster too.

The cost of making the wrong thing did not fall with the cost of making. It went up, because now you can make ten of them a quarter.

The second thing it cannot do: carry the consequences

The other half of the value is even simpler and gets discussed even less. When the work ships and the call was wrong, someone has to answer for it. A model does not get called into the room. It does not lose a customer’s trust, sit in the review, or make the fix at ten at night. Accountability is not a task you can hand to a tool, and regulators have been unusually clear about that. Announcing its enforcement sweep on AI-related deception, the FTC put it in one line: there is no AI exemption from the laws on the books (opens in new tab). The company is on the hook for what its systems tell customers, full stop.

That is the legal floor, and the design version of it is broader. Somebody has to own the pricing screen that misleads by accident, the flow that quietly fails a screen reader, the pattern the model reproduced because it is common on the web and not because it is right for your customer. Common is exactly what a model is trained to give you. Deciding that common is not good enough here, and being the name attached to that decision, is a human act.

This is the part that makes design leadership worth paying for right now. Not the ability to produce. The willingness to say this is the standard, this is what we are shipping, and I own it if I am wrong. I wrote a version of this in AI will not replace designers, it will replace the design process, and it has only gotten more true: the role survives, the busywork it was hiding behind does not.

Speed is not the outcome, and the data on that is blunt

Here is where leaders get fooled, and it is worth knowing before you sign a tooling contract on the strength of a demo. Feeling faster and being faster are different things. METR ran a randomized controlled trial with experienced open-source developers on their own repositories. The developers expected AI to speed them up by 24 percent. They were 19 percent slower with the tools (opens in new tab), and afterward they still believed they had been sped up by 20 percent. That is a forty point gap between what people felt and what actually happened, measured on people who use these tools daily.

The organizational version shows the same shape. Google’s 2025 DORA report (opens in new tab), drawn from nearly 5,000 technology professionals, found AI adoption near universal at 90 percent and more than 80 percent of respondents believing it raised their productivity. It also found that AI adoption continues to have a negative relationship with software delivery stability. Their framing is the useful one: AI does not fix a team, it amplifies what is already there. Strong systems get stronger. Weak ones get louder.

Read those two studies together and the conclusion is not that AI does not work. It is that adoption is not the achievement, and self-reported speed is not evidence. If you want to know whether AI made your team better, you have to measure the outcome the work was supposed to move, which is the same discipline good design always required. I have argued that most AI transformations are theater for exactly this reason. Tools bought, process untouched, nobody measuring.

Adoption is not an achievement. AI amplifies the judgment already in the building, and it amplifies the absence of it just as fast.

Where the hours should actually go

So here is the reallocation, concretely. Take the hours your team spent producing variations, resizing, redlining, and hand-assembling decks, and assume most of that is now machine work. Do not backfill it with more production. Move it into three places.

First, evidence and framing. Time with real customers, real session data, real support tickets, turned into a sharp statement of the problem before anyone opens a canvas. This is the input the model cannot generate and the thing that determines whether everything downstream was worth doing.

The feature my team won a company hackathon with started this way. It was not an idea from a meeting. It came from about fifty customer mentions across two years of research, all asking for the same thing in different words. The build took a day. Finding the problem took two years of listening.

Second, the standard. Someone has to look at the fast, plausible, fine output and say that is not right yet, and be specific about why. That is taste applied as a job function, and it is the difference between a team that ships a lot and a team that ships well. When making is cheap, the edit is the craft.

Third, ownership. Every shipped decision needs a name on it. Not a committee, not a tool, a person who will explain the call and fix it if it lands badly. That single practice does more for quality than any model upgrade, because it puts a human standard between the generator and the customer. It is also the whole argument in execution is cheap and ideas are expensive: the bottleneck moved from making to deciding, so staff the deciding.

Who each path is for

Names on the options, because the right answer depends on which problem you actually have.

If your bottleneck is genuinely volume, a catalog, a locale matrix, a hundred marketing variants, buy the tooling and hire a strong senior maker who can drive it. You do not have a judgment problem. You have a throughput problem, and this is the best decade in history to have one.

If your bottleneck is that nobody senior owns what gets built or what good looks like, tooling will make it worse, faster. That is a leadership gap, and it is the case for bringing in a fractional head of design who sets the standard, makes the calls, and builds the team’s judgment while you get the output of someone who has run this before. This is the shape I see most often at companies between fifteen and a hundred and fifty people: plenty of design capacity, no one accountable for the direction of it.

If you have a defined scope and a real deadline, a product design firm or a consultant is the cleaner deal. You are buying execution and a standard for a fixed window, not an ongoing seat at the table. That is a legitimate purchase and it is often the honest answer.

When the answer is not this

Three cases where I would tell you to ignore most of the above. If you are pre-product-market-fit with a handful of users, do not build a judgment apparatus. You do not have enough signal to be right yet, so ship, watch, and learn. The fastest cheap experiment is worth more than the best-framed opinion at that stage.

If you are a solo founder or a two-person team, you are already the accountability. Adding a design leader before you have a team to lead is buying process you do not need. Use the tools hard, keep your own standard high, and hire when the volume of decisions genuinely exceeds one person.

And do not read any of this as permission to stay away from the tools. The designer who says AI cannot do judgment, and uses that as a reason not to learn what it can do, is making the exact mistake this piece is warning about. Judgment gets sharper by working close to the material, not by supervising from a distance. The people I trust most on this are the designers who build, because they can feel when the output is wrong instead of approving whatever came back.

Strip it down and the map is short. AI took the middle of the work, and it took it fairly. What is left is deciding what deserves to exist and owning it once it does. Those were always the expensive parts of design. Now they are the only parts that are scarce, which means they are also where the return is. Point your hours there.