
Short answer: AI agents fail in production for engineering reasons, not model reasons. The demo had ten clean inputs and a happy path. Real users bring long conversations, vague requests, failing APIs and edge cases. Seven gaps cause almost every failure: tool errors, context overflow, no eval set, no tracing, too much permission, no budgets and no durable state. Find your gap from the symptom, then close it.
This guide is for founders, CTOs and engineering leads whose agent impressed everyone in the demo and now creates support tickets. It starts with the symptom you are seeing, because "it times out", "it calls the wrong tool" and "the bill tripled" point to different fixes. It is written from the checks we run when we take over an agent that already has users, including the 11-agent system we built that tests mobile apps on real devices.
What "production-ready" means for an AI agent
A production-ready AI agent is one where you can say, for any user request, what it did, why, what it cost and whether it was correct. In practice that means four things: an eval set run on every change, a trace for every request, scoped permissions on every tool, and hard limits on steps, tokens and cost.
If any of those four is missing, the agent is still a demo with users.
We call the method in this guide the Demo-to-Production Gap Check: start from the symptom, map it to one of seven gaps, run the first check, then fix the gap and add a test case so it cannot come back.
Start from the symptom
Most teams describe the problem in one line. This table maps that line to the gap that usually causes it and the first check to run.
| What you are seeing | Most likely gap | First check (under a day) |
|---|---|---|
| "It times out" or hangs | Tool errors, no step budget | Open 10 slow traces: which tool call or loop eats the time? |
| "It calls the wrong tool" | Too many tools, vague tool descriptions, context overflow | Count tools sent per call; read the descriptions of the two tools it confuses |
| "It makes things up about our data" | Tool returned an error or nothing, model improvised | Search traces for tool results that were empty, truncated or a stack trace |
| "It followed the rules in testing, now it ignores them" | Context overflow pushed instructions out, or too many rules | Measure prompt tokens on long conversations; check what gets cut first |
| "The bill tripled" | No step, token or cost limit | Sort runs by cost; look at the top 1% |
| "It did something it should never do" | Permissions too wide | List every write action each tool can take and with which credentials |
| "It double-charged or created duplicates" | No idempotency, no durable state | Check whether a retried run repeats side effects |
| "We cannot tell what happened" | No tracing | Try to replay yesterday's worst conversation. If you cannot, start with gap 4 |
If you cannot run the first check in the last column, that is your answer: you need tracing before anything else.
Why AI agents fail in production when the demo worked
A demo is a performance. Someone picks inputs that work, runs them once, and the model is good enough on a happy path that it looks finished. A Hacker News commenter described internal agent projects this year this way: "every one of them I've seen has been underwhelming and wrong", largely because the agents are not expert in the company's own systems and data.
Production is different in five ways at once:
| Demo | Production |
|---|---|
| 5 to 10 hand-picked inputs | Thousands of real inputs, many vague or wrong |
| Short, fresh conversations | Long threads that fill the context window |
| APIs that answer instantly | Timeouts, rate limits, 500s, partial data |
| One user | Concurrent users, retries, double clicks |
| Someone watching | Nobody watching until a customer complains |
None of these are model problems, which is why swapping GPT for Claude, or the other way round, rarely fixes them.
The math that makes long agents fragile
Every step an agent decides on its own is a chance to go wrong, and the chances multiply. If each step succeeds 95% of the time:
| Steps the model decides | Success at 95% per step | Success at 99% per step |
|---|---|---|
| 3 | 85.7% | 97.0% |
| 10 | 59.9% | 90.4% |
| 20 | 35.8% | 81.8% |
These are calculated as the per-step rate raised to the number of steps (for example 0.95 to the power 10 is 0.599). Real steps are not fully independent, so treat this as a direction, not a forecast. The direction is clear: an agent that needs 20 free decisions per task will fail often even if each decision looks good in isolation. The most reliable fix is fewer model-decided steps, which we cover at the end.
If this table describes your agent, the fix is usually a redesign of a few steps, not a weekend of prompt tuning. That is the kind of work our engineers do every week; you can talk to the founders and bring one trace that went wrong.
Gap 1: tool calls fail and the agent improvises
In a demo every API answers. In production the CRM times out, the search returns nothing, a required field is missing. If the tool hands back a raw stack trace or an empty string, the model guesses what happened and carries on, often confidently wrong.
What to put in place:
- A timeout on every tool, shorter than your user-facing timeout. Serverless limits make this worse; we covered that in Vercel function timeouts for AI apps.
- Typed, readable errors. Return
{ "error": "customer_not_found", "hint": "ask the user for their account email" }, not a stack trace. The model can act on the first and hallucinates around the second. - Retry rules in code, not in the prompt: retry idempotent reads with backoff, never blindly retry writes.
- Clear tool descriptions. Say when to use each tool and when not to. Two tools with overlapping descriptions are the most common cause of wrong tool calls we see.
Gap 2: long conversations overflow the context
The demo conversation was four turns. Real users come back to the same thread for days, paste documents and ask follow-ups. Eventually the system prompt, tool definitions, history and tool results stop fitting, and something gets cut. Often it is the instructions that kept the agent safe, which is why an agent "passes every eval and then ignores its rules".
Fixes:
- Count tokens before every call and decide explicitly what to drop or summarise.
- Pin the instructions and tool definitions; summarise old turns, never the rules.
- Trim large tool results before they enter the context. A 4,000-row export is for code to process, not for the model to read.
- Send fewer tools per step. Twenty tool definitions in every call cost tokens and raise the odds of a wrong choice. Route first, then expose only the tools that step needs.
- Keep rules few and specific. A long list of rules is harder to follow than five that matter.
Gap 3: no eval set, so nobody knows what "working" means
Without a test set, every prompt change is a coin flip and every model upgrade is a surprise. You need an eval set before you can improve anything.
Build it from real traffic, or from realistic requests if you have not launched. Each case states the request, the expected outcome and the tools that must and must not be called. Here is one case in the shape we use:
id: cancel-last-order-keep-subscription
request: "cancel my last order but keep the subscription"
user_context:
account: acct_demo_204
orders: [1182, 1175]
expect:
outcome: "order 1182 cancelled, subscription untouched, confirmation sent"
must_call: [get_orders, cancel_order]
must_not_call: [cancel_subscription, refund_payment]
max_steps: 6
max_cost_usd: 0.05Track four numbers on every run: task success rate, wrong or forbidden tool calls, steps per task and cost per task. Run the set in CI on every change to prompts, tools or models. OpenAI's evals guide describes the same discipline. Start with 50 cases, and turn every production failure into a new one, so the same bug cannot ship twice.
Gap 4: no tracing, so failures are invisible
When a user says "the bot did something weird yesterday", can you open exactly what happened? If not, you are debugging from screenshots.
Log one trace per user request, with every model call (prompt, model, tokens, latency) and every tool call (arguments, result or error, duration), plus the final outcome. The OpenTelemetry semantic conventions for generative AI give a standard shape for these spans, so they go to the observability stack you already run instead of another dashboard.
Then review traces every week. Sort by failures, longest runs and most expensive runs. That review is where most real fixes come from, and it feeds new cases into the eval set from gap 3.
Gap 5: the agent can do too much
In a demo, giving the agent an admin token is convenient. In production it means one confused step can change or delete real data. OWASP lists this as LLM06:2025 Excessive Agency and names three root causes: excessive functionality, excessive permissions and excessive autonomy.
Scope every tool:
- Separate read tools from write tools, and give each its own credentials.
- Rate each write tool by risk. OpenAI's guide suggests rating tools by read-only versus write access, reversibility, required permissions and financial impact, then adding checks or a human before high-risk ones.
- Require human approval for irreversible or expensive actions until your evals and traces prove they are safe.
- Treat text from tools, web pages and documents as untrusted. OWASP LLM01:2025 Prompt Injection describes indirect injection, where instructions hidden in an external source, such as a website or file, reach the model. An agent with wide permissions turns that into real actions.
Gap 6: no budgets, so cost and latency drift
Agents loop. A model that cannot find what it needs will call the same search five times, try another tool, then start again. In a demo that costs cents. At 10,000 requests a day it shows up in the bill and in response times.
Set hard limits in code:
- A maximum number of steps and tool calls per task, with a graceful handoff when the limit is hit.
- A maximum number of tokens per run.
- A cost alert per task, not only per month.
- A smaller, faster model for routing and simple steps, and the large model only where judgement is needed.
OpenAI's guide also recommends escalating to a human when an agent exceeds failure thresholds, for example after several attempts to understand what the user wants.
Gap 7: no durable state, so retries repeat side effects
Long tasks get interrupted: a deploy restarts the worker, a function times out at minute four, a user refreshes. If the agent's progress lives only in memory, the task either disappears or starts again from the top and repeats what it already did, such as sending the email twice or creating a second order.
Fixes:
- Save the task's state after each step that changes something, so a restart resumes instead of repeating.
- Put an idempotency key on every tool that creates, sends or charges, so a repeated call is ignored. It is the same idea as in Next.js server actions idempotency.
- Run long tasks in a background job or workflow runner, not inside a web request.
The structural fix: make the agent smaller
Most of the seven gaps shrink when the agent has less to decide. Anthropic's Building effective agents separates workflows, where LLMs and tools are orchestrated through predefined code paths, from agents, where the model directs its own process, and says the most successful teams use simple, composable patterns. OpenAI's guide makes the same point about growing a single agent by adding tools before reaching for multi-agent systems.
In practice: if a step is always the same (look up the account, check the plan, fetch the last five orders), write it as code. Let the model handle the parts that need judgement, such as understanding the request, choosing between options and writing the reply. Going back to the table above, cutting a task from 20 model-decided steps to 5 moves it from roughly 36% to 77% at 95% per step. A support agent that is mostly fixed workflow with a few model decisions is easier to test, cheaper and more predictable than one that improvises everything.
The production checklist
| Area | Check | Done? |
|---|---|---|
| Tools | Timeouts, typed errors, retry rules, clear descriptions | |
| Context | Token counting, pinned instructions, trimmed tool results, small tool sets | |
| Evals | 50+ real cases with expected tools and outcomes, run in CI | |
| Tracing | One trace per request with every model and tool call | |
| Permissions | Read and write split, risk-rated tools, approval on irreversible actions, untrusted tool text | |
| Budgets | Step, token and cost limits per task, alerts per task | |
| State | Saved progress per step, idempotency keys on writes, long tasks off the web request | |
| Change control | Model and prompt versions pinned, upgrades gated by evals |
Model versions deserve their own row because providers retire them on a schedule; see OpenAI's September 2026 model shutdowns.
A two-week plan to get from demo to production
| Days | Work | Done when |
|---|---|---|
| 1 to 2 | Add tracing to every model and tool call | You can replay any conversation from yesterday |
| 3 to 4 | Run the symptom table on the last 100 traces | The top 3 failure causes are named, with counts |
| 5 to 7 | Build the first 50 eval cases and run them in CI | You have a baseline success rate and cost per task |
| 8 to 10 | Fix the top cause: tool errors, context or permissions | The eval success rate goes up and no forbidden calls remain |
| 11 to 12 | Add step, token and cost limits and human approval on risky writes | No run can exceed its budget or act irreversibly alone |
| 13 to 14 | Move fixed steps into code and re-run the evals | Fewer model-decided steps, same or better success rate |
What this guide covers, and where it does not apply
This guide covers LLM agents that call tools through APIs inside a product or an internal system: support agents, operations agents and SaaS copilots. It does not cover model training or fine-tuning, voice agents (latency and turn-taking are their own problem), or tuning retrieval for RAG answers.
Some advice matters less in two cases:
- One-shot generation with no tools. If the model writes a summary and nothing else happens, there is no agent. Gaps 5 to 7 barely apply; gaps 2 to 4 still do.
- Batch jobs where a person reviews every output. Human review covers much of gap 5, so spend the effort on evals and tracing first.
Two limits on what we say here. The compounding table assumes each step succeeds independently, which real steps do not, so read it as a direction. And our view comes from building and taking over client agents, not from a controlled study across many companies.
One tradeoff we make on purpose: we push teams towards fixed workflows with a few model decisions instead of a fully autonomous agent, because workflows can be tested and their cost is predictable. The price is flexibility on unusual requests. We handle that with a fallback: when a request does not fit the workflow, the model gets a wider, budgeted path or the case goes to a human.
For teams in Europe, the UK and the UAE
Two additions if you sell outside the US:
- Traces contain personal data. Prompts and tool results often hold names, emails and order details. Under GDPR in the EU and UK, and the UAE's Personal Data Protection Law, decide what you log, mask what you do not need, set a retention period and keep traces in a region your data agreements cover.
- Human oversight is part of the design. Approval steps on high-impact actions help reliability, and they answer the oversight questions European and Gulf enterprise customers ask in security reviews.
Bring us the trace that worries you
Most AI agents fail in production for the reasons above, and almost all of them are fixable. If your agent already has users and you cannot answer "what did it do on this request, and why?", the gap list above tells you where to start, and the two-week plan tells you in what order. If you would rather have senior engineers do it with you, we are a small team with founders from Microsoft, IIT Kanpur and a $1B+ startup, and an ex-Google engineer on the team. We have built multi-agent systems, including an 11-agent setup that tests mobile apps on real devices, and a voice AI calling platform.
Talk to the founders for 30 minutes and bring one trace or one conversation that went wrong. You can see what we build, from AI agents and RAG to dedicated senior AI engineers, on the homepage.
FAQ
Frequently asked questions
- Why does our AI agent work in testing but fail in production?
- Testing used a handful of clean inputs on a happy path. Production adds long conversations that overflow the context window, vague requests, slow or failing APIs, concurrent users and inputs nobody imagined. Without an eval set built from real traffic and a trace of every run, you cannot see which of these is breaking.
- Why do AI agents fail in production most often?
- In our experience the most common causes are tool calls that fail silently, context windows that fill up on long conversations, and the absence of an eval set, so nobody notices regressions. Wide permissions and missing cost limits cause fewer failures but the most expensive ones.
- How do we test an AI agent before release?
- Build an eval set from real or realistic requests, each with the expected outcome, the tools that must and must not be called, and a step limit. Run it on every change to prompts, tools or models, and track task success rate, wrong tool calls, steps per task and cost per task.
- Should we use a multi-agent system to fix reliability?
- Usually not at first. OpenAI's guide to building agents recommends getting one agent working by adding tools before splitting into several, and Anthropic advises starting with simple, composable patterns. Every extra agent adds handoffs and steps, and each step is another chance to fail.
- How do we stop AI agent costs from exploding?
- Set a maximum number of steps and tool calls per task, cap tokens per run, cache stable context such as instructions and tool definitions, use a smaller model for routing and simple steps, and alert on cost per task instead of waiting for the monthly bill.
- What should we log for an AI agent in production?
- One trace per user request that holds every model call (prompt, model, tokens, latency), every tool call (arguments, result or error, duration) and the final outcome. The OpenTelemetry semantic conventions for generative AI give a standard shape, so the traces fit the monitoring stack you already run.
Sources
- 01Anthropic: Building effective agents
- 02OpenAI: A practical guide to building agents (PDF)
- 03OpenAI API: Working with evals
- 04OpenTelemetry: Semantic conventions for generative AI systems
- 05OWASP Top 10 for LLM Applications: LLM01:2025 Prompt Injection
- 06OWASP Top 10 for LLM Applications: LLM06:2025 Excessive Agency
- 07Hacker News discussion: why internal agent projects underwhelm
#AI agents #AI agents in production #why AI agents fail #LLM evals #observability #OpenTelemetry #tool calling #prompt injection #LLM cost #reliability
Written by

Akshit Ahuja
Co-founder, Chief Speed Officer
Ex-Spinny, a $1B+ startup / Backend and AI systems
Backend systems specialist who thrives on building reliable, scalable infrastructure. Akshit handles everything from API design to third-party integrations, ensuring every product HeyDev ships is production-ready.


