Skip to content

AI

Why AI agents fail in production, and what to build around them

The recurring reasons LLM agents break after the demo: compounding errors, bad tools, lost context, cost loops and missing evaluation. Fixes you can ship.

By · Published · 3 min read

Short answer: agents fail in production for boring reasons. Small per-step error rates multiply across a long chain, tools are described badly, context fills up, a loop never ends, nobody measures quality, and the task did not need an agent in the first place. Fix them with shorter workflows, stricter tools, hard limits, logging, and tests built from real failures.

Why do errors compound?

If each step in a chain succeeds 95 percent of the time, ten steps succeed about 60 percent of the time, because 0.95 to the tenth power is roughly 0.6. A demo shows you the run that worked. Production shows you all of them.

The remedy is fewer steps. Replace model decisions with code wherever the decision is deterministic. If you always validate the invoice number with a regex, do it in code and give the model the result.

Is it a workflow or does it need an agent?

A workflow is a fixed path where a model fills in some steps. An agent chooses its own path in a loop. Most business tasks fit a workflow: classify the ticket, extract fields, look up the order, draft the reply, ask a person to approve. It is easier to test, cheaper and easier to explain. Use a free-running agent only when the number and order of steps truly cannot be known in advance.

What goes wrong with tools?

  • Vague names and descriptions. The model picks the wrong tool or invents arguments.
  • Too many tools. With forty tools in context, selection accuracy drops. Give each task the five it needs.
  • Raw API wrappers. Returning a 4,000-line JSON blob wastes context and hides the answer.
  • Errors as stack traces. Return a short message that says what to change. "date must be YYYY-MM-DD" lets the model recover.
  • No idempotency. The model retries a call, and you create two orders. Accept an idempotency key.

What happens when context fills up?

Long sessions accumulate tool output, and early instructions get lost or ignored. Symptoms are an agent that forgets the goal, repeats work, or contradicts an earlier decision. Keep outputs short, summarise old turns into a few lines of state, and store facts in a database the agent can query instead of keeping them in the prompt.

Why do agents loop and burn money?

An agent that cannot solve a step will often try the same call again. Without a limit, it runs until your budget does. Set a maximum number of steps, a token budget and a wall-clock timeout per run. When a limit hits, stop and hand the case to a person with the trace attached. See monitoring LLM costs for how to catch this early.

Why does nobody notice quality dropping?

A model update, a prompt edit or a changed upstream data format can degrade results without raising an error. You need an evaluation set. Start small: collect twenty to fifty real inputs with the answer you would accept, run them on every change, and compare. Add every production failure to the set. This is the single habit that separates teams that improve from teams that guess.

What should you log?

The full trace for each run: input, each model call with its prompt and output, each tool call with arguments and result, latency, token counts and the final outcome. Without it you cannot reproduce a bad run. LLM observability covers the tooling.

Where should a human stay in the loop?

Any action that is hard to undo: sending money, messaging a customer, deleting data, changing a record that other records depend on. Show the proposed action and its reasoning, let a person approve or edit it, and record the decision. Over time the approval rate tells you which steps are safe to automate.

A production-readiness list

  • Step, token and time limits are set.
  • Every tool has a clear description, a strict schema and idempotency where it writes.
  • Irreversible actions need approval.
  • Traces are stored and searchable.
  • An evaluation set runs before each change.
  • There is a manual fallback when the agent gives up.

None of this is exciting, and all of it is what makes an agent usable on a Tuesday afternoon with real customers.

References

Author

Raktim Ranjit is a software engineer and the founder of NodeDR Infotech. He builds and maintains the software described here.

Have something in mind?

Let’s build something useful.

Tell me about the idea, product, or workflow you’re working through.

Tap to say hello