An MCP demo almost always impresses, and it almost never ships. It reaches the live data, it answers in plain language, the room nods along. Then someone asks the question that kills it. What happens when the model quotes a price under your floor, or cites a product that does not exist, or writes copy that breaks every brand rule you have? In the demo, nothing, because the demo never had a bad day. In production, that is a Tuesday. The entire distance between a prototype and something a team trusts with real work comes down to one decision: where the rules live.
Here is the stance, and I am not going to hedge it. Most MCP demos die because they trust the model with rules that belong in the server. If a constraint actually matters, your brand voice, a price floor, never inventing data, it cannot live in a prompt that asks the model nicely to behave. It has to live in the tool itself, in code the request runs through and cannot route around. Bake the constraint into the server and the agent physically cannot break it. That is the whole difference between something that demos well and something you put in front of a customer.
I did not arrive at this from a conference talk. I built the agent pipeline at the AI company I founded, and I built an MCP over a product catalog that is now a V1 in pilot. Between the two, I have watched the model try nearly every wrong thing a model can try. If you want the plain-language version of what an MCP even is before we get into building a good one, I wrote that in what an MCP actually is. This piece is about the harder question that comes next: how you build one a team will actually use instead of one that wins a demo and quietly dies.
The demo trusts the model. The V1 does not.
A demo is easy for a reason. It only has to work once, in front of a friendly audience, on a happy path someone chose on purpose. Nobody feeds it the weird input. Nobody asks for the discount that would sink the margin. The model looks reliable because it was never given a chance to be anything else. A V1 is hard for the opposite reason. It runs thousands of times, against inputs you did not script, and a model is a probabilistic system. On a long enough timeline it will produce the confident wrong answer, not because it is broken but because that is what sampling from a distribution does.
This is not a knock on the model. It is a fact about where the model sits. The MCP spec is explicit that tools are model-controlled (opens in new tab), meaning the model decides which tool to call and what arguments to pass based on its read of the moment. That is exactly the power you want, and exactly why you cannot assume the arguments will be sane. The model choosing freely is the feature. Your server treating every one of those choices as unverified input is the safeguard that makes the feature usable.
If a rule matters, it cannot live in a prompt that asks the model nicely. It has to live in code the request runs through and cannot route around.
Where the guardrails actually live
The protocol authors are blunt about whose job this is, and it is worth reading in their own words. On security, the spec says servers MUST validate all tool inputs, implement proper access controls, rate limit tool invocations, and sanitize tool outputs (opens in new tab). Note the word. Not the model should try. The server must. Every guardrail that matters is described as work the tool does, on the server side, on every call, whether the model cooperated or not.
So take the three constraints teams ask about most and put them where they belong. A price floor is not an instruction in a system prompt. It is a check in the pricing tool that rejects or clamps any number under the minimum before it can ever reach a customer, and returns the reason so the model can correct. No invented data is not a plea to please only use real products. It is a tool that queries the catalog and returns exactly what it finds, and when a product does not exist it returns an error, not a plausible guess, because the tool is the only place the answer is allowed to come from. Brand rules are not a paragraph of tone guidance the model may forget by the third paragraph. They are a validation step that checks the output against the real rules before the server hands anything back.
Once the constraints live in the tool, something quietly changes about the whole system. You stop needing the model to be trustworthy. You need it to be useful, which is a far lower and far more realistic bar. The model can propose anything it wants, and the proposals that would break a rule simply do not make it out of the server. That is what people mean, or should mean, when they say a system is production-grade. Not that the model never errs. That the model’s errors cannot reach the customer.
What to wrap, and what to leave out
The instinct when you build your first MCP is to expose everything. You already have an API, so you wrap every endpoint, hand the model the full keyboard, and call it powerful. It is a trap, and Anthropic named it directly in their guidance on building tools for agents: a common error is tools that merely wrap existing software functionality or API endpoints (opens in new tab). A raw endpoint is shaped for a programmer who reads the docs and holds the whole system in their head. The model has neither. It needs verbs shaped for the job, not a mirror of your backend.
So I wrap the actions the work actually needs, and I shape each one. Fewer tools, named for what they accomplish, with the constraints built in and the arguments narrowed to what is safe to vary. The same guidance makes the other half of the point: a tool should return only high-signal information back to the agent. A raw endpoint that dumps fifty fields buries the model in noise and burns the context it needs for the actual task. A good tool returns the five fields that matter, already filtered, already clean. You are not exposing your API to the model. You are designing an interface for a very fast colleague who cannot read the manual.
Errors are a feature, not a failure
The part first-time builders get wrong is treating a rejected call as a dead end. It is the opposite. When your tool refuses a bad request, that refusal is your best chance to steer the model back on course. Anthropic makes this explicit: when a tool call raises an error, for example during input validation, you can prompt-engineer your error responses to clearly communicate specific and actionable improvements (opens in new tab). A tool that answers a low bid with "price 4.00 is below the floor of 9.99, raise it and try again" does not just block the mistake. It teaches the model to fix it on the next call, without a human stepping in.
This is the loop that makes a server-enforced system feel smart instead of rigid. The server holds the hard line, the error explains the line, and the model adjusts. You get the flexibility people love about agents and the reliability people need from production at the same time, because the two jobs are split. The model handles the open-ended part, working out what the user wants and how to get there. The server handles the non-negotiable part, making sure that whatever the model tries, the rules hold.
The model is just the newest untrusted client. You would never let the browser enforce your price floor. Do not let the model either.
None of this is new, and that is the point
If this sounds familiar, it should. We learned it a decade ago on the web and wrote it down. OWASP, the standard reference for application security, puts it plainly: input validation must be implemented on the server-side (opens in new tab), because any validation you do on the client can be circumvented. No serious team trusts the browser to enforce a business rule, because the browser is under the user’s control, not yours. The model is the same story wearing new clothes. It is the newest untrusted client, and the fix is the one we already know: the server is where the rules are enforced, every time, no exceptions.
That is why I am not nervous about the model getting better, and you should not be either. A stronger model makes the useful part more useful. It does not change where the guardrails go. The teams that internalize this ship MCPs their colleagues actually reach for, because the tool is safe to use without babysitting. Skip it and you get a beautiful demo that does not survive contact with a real customer, and the model takes the blame for doing exactly what a probabilistic system does.
Build it once, and point it inward too
The server you build for customers is often the best internal tool your team never asked for. The same tool that feeds product data to a storefront can feed it to your designers. Instead of copying images and text from a browser into a design file one product at a time, a designer asks for the products and they land in the file, real and current. The rules you built into the server come along for free, so the mockup shows the real price and the real name, not a placeholder someone forgets to swap. And the source does not have to be your live site. Point the same server at a PIM, the product information system most retailers already run, or a digital asset library, and the strategy holds. What matters is one trusted source, wired once, that both your customers and your team pull from.
What makes it compound is a cache that remembers. Save every pull into a shared component library, so the next designer who needs that product gets it instantly instead of hitting the API again. Then add two checks that keep the shortcut honest. A staleness check flags anything that has not been refreshed in too long. A drift check flags a saved product that no longer matches the source, like a new price, a retired color, or a swapped image. Without those checks, a cache is just a faster way to show old data. With them, the tool gets faster the more people use it and stays right while it does.
It is also the cheapest way to earn trust in the server before customers lean on it. Your own team uses it every day and tells you what breaks, while a fix still costs nothing. By the time it faces the outside, the people down the hall have already found the rough edges.
Who this is for, and when it is not this
This is for anyone about to put an agent in front of real work, real money, a real brand, or a real customer. If a wrong answer costs something, and a human is not going to inspect every single output, the constraints belong in the server, full stop. Build the tool so the expensive mistakes are impossible, not merely discouraged, and you have something a team will trust with the actual job instead of a toy they demo once and shelve.
There is a real case where this is too much, and I want to name it so you do not over-build. If your MCP is an internal tool for one or two power users who read every output before it goes anywhere, the heavy guardrails are ceremony you do not need yet. A thin wrapper and an attentive human is a perfectly good V1 at that scale. The same is true for genuinely low-stakes, read-only, exploratory work where the worst case is a shrug and a retry. The server-side rigor earns its cost the moment the mistakes get expensive and the human steps out of the loop. Below that line, keep it simple. Above it, there is no shortcut. Put the rules in the server, and build the thing your team can actually trust.