Skip to content

AI

How to monitor and cut LLM costs in production

Where LLM spend actually comes from, how to attribute it to features and users, the alerts worth setting, and the changes that reduce a bill without hurting quality.

By · Published · 3 min read

Short answer: record input tokens, output tokens and the model name for every call, tag each call with the feature and user it served, and compute cost from your provider's current price list. Alert on spend per day and on the most expensive single request. Then cut cost by shortening prompts, caching, routing easy work to a smaller model and capping agent loops.

Where does the money go?

Cost is roughly input tokens times the input price plus output tokens times the output price. Output tokens are usually priced several times higher than input tokens. Three things inflate the bill without anyone deciding to spend more.

  • Context growth. A chat that resends the full history on every turn costs more each turn. Turn twenty may carry ten times the input of turn two.
  • Retrieval stuffing. Adding ten documents "just in case" to every prompt.
  • Loops. An agent retrying a failing step for fifty rounds.

How do you attribute cost?

Your provider bill gives one number. You need to know which feature caused it. Attach metadata at the call site, and store it with the token counts.

await logCall({
  feature: "invoice-extraction",
  userId,
  model: res.model,
  inputTokens: res.usage.input_tokens,
  outputTokens: res.usage.output_tokens,
  traceId,
});

Store the token counts, not a computed dollar value. Prices change, and you want to recompute history. Keep a small price table keyed by model and effective date.

Which queries should you run?

-- spend by feature, last 7 days
SELECT feature,
       sum(input_tokens  * p.in_price  / 1e6 +
           output_tokens * p.out_price / 1e6) AS usd
FROM llm_calls c JOIN model_prices p USING (model)
WHERE c.created_at > now() - interval '7 days'
GROUP BY feature ORDER BY usd DESC;

Run the same grouped by user, and list the ten most expensive single traces. The top one is often a bug.

Which alerts are worth setting?

  • Daily spend above a fixed ceiling, sent to a channel a person reads.
  • Any single trace above a threshold, for example five times your median.
  • Spend per active user above a limit, which catches abuse and runaway sessions.
  • A sharp rise in average input tokens per call, which usually means a prompt or retrieval change.

Also set a hard limit at the provider if they offer one, and per-user rate limits in your own code. An alert tells you after the money is gone. A limit stops it.

How do you reduce cost without hurting quality?

  • Trim the prompt. Remove instructions the model follows anyway and examples that do not change behaviour. Test the shorter version against your evaluation set.
  • Use prompt caching. If your provider supports it, put the stable part of the prompt first so repeated prefixes are billed at a lower rate.
  • Route by difficulty. Send classification, extraction and simple rewriting to a small model. Keep the large model for hard cases.
  • Cap output. Set a max token limit and ask for a structured, short answer.
  • Summarise history. Replace old turns with a short running summary.
  • Cache identical requests. Many apps get the same question repeatedly.
  • Batch offline work. Many providers discount asynchronous batch jobs.
  • Cap loops. A maximum step count is cost control as much as reliability.

How do you know a cheaper change did not break things?

Run your evaluation set on both versions and compare. Without it you are trading an invisible quality loss for a visible saving. If you do not have one yet, read what LLM observability is and start collecting real cases.

What about self-hosting a model to save money?

It can pay off at steady high volume, and it adds work: GPUs, serving software, updates and monitoring. For most small products the hosted API is cheaper once you count your own time. Compare on total cost, not on the per-token price alone. If you want to try it, see how to self-host an LLM.

Author

Raktim Ranjit is a software engineer and the founder of NodeDR Infotech. He builds and maintains the software described here.

Have something in mind?

Let’s build something useful.

Tell me about the idea, product, or workflow you’re working through.

Tap to say hello