AI-first is one of the most abused phrases in the industry, so let me say plainly what it means when it is real. It is not buying seats, running a workshop, or bolting a chat box onto your product. AI-first is wiring the systems you already run, your design system, your product data, your internal tools, directly to agents, so that work which used to take a team weeks takes minutes instead. Everything short of that plumbing is theater. It looks like progress, but it is theater.
I am not describing this from a keynote. I led design inside a multi-billion-dollar org through this exact shift, and I founded and shipped an AI company where the agents are the product, not a feature on top of it. What I actually do with AI is not write more clever prompts. I build the connective tissue between real systems and models, and then I watch the hours drop. This piece is a concrete account of that tissue: tokens that emit front-end code, MCPs wired over live product data, product pages spun up on demand, and what going AI-first costs the person who has to lay the pipe.
The reason this needs saying is that the prompt has become the whole conversation, and the prompt is the least valuable thing you can automate. If the most interesting thing your team does with AI happens inside a text box, you have not gone AI-first. You have bought a faster typewriter and told the board it was a transformation.
The prompt is the cheap part
Prompt craft is real, and on day one it feels like the skill. You learn to phrase a request so the model gives you something usable, and it feels like leverage because the alternative was a blank page. That is a fine place to start. It is a terrible place to stop, because the phrasing is the fastest-commoditizing layer in the whole stack.
Look at where the models are going on their own. METR, a research group that measures what AI agents can actually do, found that the length of task a frontier model can finish autonomously with 50 percent reliability has been doubling roughly every seven months (opens in new tab). The raw capability is improving without you lifting a finger. So if your advantage is a clever wording that coaxes a bit more out of the model, you are standing on the one part of the system that gets cheaper and more forgiving every quarter. The prompt tricks you hoard today are next year’s defaults, shipped in the base model.
A more clever prompt saves you a sentence. Wiring your real systems to an agent saves you the week.
The durable advantage is not how you ask. It is what the model can reach when you do. A model with no access to your systems is a very smart contractor locked out of the building. The work of AI-first is unlocking the doors: making your design system, your product data, and your internal tools legible and callable, so an agent can do real work against the real thing instead of hallucinating a plausible version of it. That work is plumbing, and plumbing is exactly what the theater skips.
Tokens that emit the front-end
Start with the design system, because it is the clearest case. Most teams treat a design system as a library of screenshots and a Figma file people copy from. That version is useless to a model. The version that matters treats the system as data: design decisions written as named values a machine can read. A design token, in the words of the W3C group standardizing the format, is information associated with a human readable name, at minimum a name and value pair (opens in new tab). Your brand blue stops being a swatch someone eyeballs and becomes color.text.primary with a defined value, referenced everywhere.
Once the system is tokens instead of pictures, it can emit front-end code. The same spec defines translation tools that convert token data into platform-specific source code developers can use, and names the ones teams already run, Style Dictionary and Terrazzo. Point those at your tokens and one source of truth spits out CSS variables, iOS constants, and Android resources at once. Now an agent generating a screen is not guessing at your spacing scale or inventing a third shade of blue. It is composing from the same tokens your production build compiles, so the output is on-brand and consistent by construction, not by a designer catching drift in review three days later.
This is the difference between a model that decorates and a model that builds. When the system is legible, the agent produces real front-end that matches the product because it is drawing from the product’s actual vocabulary. When the system is a folder of images, the model produces something that looks roughly right and quietly diverges from everything you ship. I wrote separately about why the design system is the foundation the whole AI-first process stands on. The short version: tokenize the system or the speed you gain from AI just accelerates the chaos.
MCPs wired over real product data
The second pipe is the one that connects an agent to your live data. This is what the Model Context Protocol is for. Anthropic introduced it in November 2024 as an open standard that enables developers to build secure, two-way connections between their data sources and AI-powered tools (opens in new tab). In practice it is a thin, standard wrapper you write over the APIs and actions you already have, described so a model knows what each one does and when to call it.
I built an MCP server that wires an agent pipeline to the tools it needs to do actual work. The payoff is not abstract. When an agent can query real product data, real inventory, real customer records, through a defined tool instead of being handed a stale spreadsheet in the prompt, it stops making things up and starts operating on the truth. A support flow that needed a person to open three dashboards and stitch the answer together becomes one request the agent resolves against the same endpoints those dashboards call. I unpacked how this actually works in what an MCP actually is, because the plumbing is easy to mystify and cheap to build once you stop believing the marketing.
The point for a design leader is this: the interface work does not end at the screen anymore. Wiring the model to real data is design work, because you are deciding what the agent can see, what it can touch, and how it explains itself when it acts. Skip that wiring and your AI-first initiative is a very expensive way to generate confident guesses about data the model was never allowed to read.
Product pages spun up on the fly
Put the two pipes together, tokens that emit front-end and MCPs that reach live data, and you get the thing that still surprises people. Real product pages, spun up on demand, populated with real content, in the time it used to take to schedule the kickoff. Not a mockup. Not a Figma frame someone will rebuild in code next sprint. The actual page, in the actual system, drawing on the actual data.
At Story Genie, the company I founded, agents write and illustrate a full personalized hardcover children’s book, and the parent sees the entire finished book, every page, before they pay a cent. The preview builds in about sixty seconds. That is a product page and its contents generated on the fly, per customer, from a pipeline wired end to end. No human studio could ever show you a complete, custom, illustrated book about your specific child before you decide to buy, because producing it by hand would cost more than the sale. The plumbing is what makes the impossible thing routine.
AI-first is not a purchase you announce. It is a set of pipes you lay, and the work only drops from weeks to minutes once they connect.
The plumbing is the job
Here is the part leaders want to delegate and cannot. Someone has to lay the pipe, and if the most senior person in the room has never wired a token to a build or a tool to an agent, they cannot tell the difference between real plumbing and a convincing demo of it. That is how theater gets funded: the people approving the budget cannot see which parts are load-bearing. I stay close to the build for exactly this reason, and I made the case for it in why I am still a leader who builds.
When the pipes connect, the numbers stop being incremental. Inside that multi-billion-dollar org, we built an agent against one workflow and took it from around 320 hours to about one. That is not an efficiency gain you round up in a slide. It is a change in what a single person can carry, and it only happened because the systems were wired, not because anyone found a magic prompt. The hours freed did not vanish and neither did the people. They moved up the stack, into judgment and taste and deciding what was worth building at all, the work the production grind used to crowd out.
Who this is for, and when it is not this
This is for leaders who own real systems and are willing to make them reachable, or willing to sit close enough to the people who do that they can tell wiring from theater. If that is you, stop grading your AI-first effort by how many licenses you bought or how good your team is at prompting. Grade it by one question: what can an agent now do against your real systems that it could not do last quarter? If the answer is nothing, you have bought a costume, not a capability.
There is a real case where this is the wrong project, and I want to be clear about it. If you are a small team with a handful of one-off tasks, do not build a plumbing project. The wiring has a fixed cost, and if you are only ever going to generate a page twice, the token pipeline and the MCP server cost more than the hours they save. Just pick a few tools your team will actually open, use them well, and prompt your way through it. The plumbing earns its keep when the same work repeats, when the same reports get pulled, the same pages get built, the same data gets stitched, over and over, forever. Automate the repeated thing, not the rare one.
And if your organization is not honest about the difference between motion and progress, the plumbing will not save you, because it will never get built. That failure mode is common enough that I wrote a whole piece on it, why most AI transformations are theater. The tell is always the same. When a leader describes their AI-first shift entirely in terms of tools purchased and prompts refined, and cannot name a single system they wired to an agent, you are watching the performance, not the work. The leverage was never in the prompt. It was always in the pipe, and someone has to lay it.