Skip to content

AI

What is LLM observability, and what should you log?

LLM observability means recording enough about every model call to debug, evaluate and cost it. What to capture, what to redact, and how to start with a table and a few queries.

By · Published · 3 min read

Short answer: LLM observability is the practice of recording each model call and each step around it, so that when an answer is wrong, slow or expensive you can see exactly what went in, what came out and what it cost. At minimum that is the prompt, the response, the model and parameters, token counts, latency, errors, and a trace ID that links the steps of one user request.

Why is normal logging not enough?

A web request either works or throws. A model call returns something plausible even when it is wrong. There is no stack trace for "the summary left out the refund". You need the actual inputs and outputs to judge it, and you need to be able to find the bad ones among thousands.

What should you capture for every call?

  • Timestamp, environment and a trace ID shared by every step of one request.
  • Model name, version and parameters like temperature and max tokens.
  • The full prompt, including system prompt, retrieved documents and tool definitions, or a stored reference to them.
  • The full response, including tool calls the model asked for.
  • Input tokens, output tokens and, if your provider reports it, cached tokens.
  • Latency to first token and total latency.
  • Status: success, refusal, timeout, rate limit, schema validation failure.
  • The prompt template version, so you can tie a change in quality to a change in text.
  • User feedback if you collect it, such as a thumbs-down or an edit.

What about agents and multi-step flows?

Use spans. A user request is a trace. Each model call, each retrieval and each tool call is a span inside it, with a parent. Reading a trace top to bottom shows where time went and where the chain went wrong. OpenTelemetry has semantic conventions for generative AI spans, which means you can use your existing tracing backend instead of adopting a new product just for this.

What should you not store?

Prompts contain whatever users typed, and users type personal data. Decide what you keep before you start. Redact emails, phone numbers, card numbers and government IDs at write time. Set a retention period, 30 days is common for raw prompts, and keep only aggregates beyond that. If you operate under a data protection law, a log full of raw prompts is personal data and needs the same care as a database.

Can you start without a vendor tool?

Yes. One table in PostgreSQL is enough to learn what you need.

CREATE TABLE llm_calls (
  id            bigserial PRIMARY KEY,
  trace_id      uuid        NOT NULL,
  parent_id     bigint,
  created_at    timestamptz NOT NULL DEFAULT now(),
  model         text        NOT NULL,
  prompt_ver    text,
  input_tokens  int,
  output_tokens int,
  latency_ms    int,
  status        text        NOT NULL,
  prompt        jsonb,
  response      jsonb,
  feedback      smallint
);
CREATE INDEX ON llm_calls (trace_id);
CREATE INDEX ON llm_calls (created_at);

Then write three queries and look at them weekly: the slowest ten percent of calls, the most expensive traces, and every call with negative feedback or a validation failure.

What do you do with the data?

  • Debug. Open the trace for a complaint and read what the model actually saw.
  • Build an evaluation set. Copy bad cases into a test file with the answer you would accept, and run it on every prompt change.
  • Control cost. Group by feature and prompt version. See how to monitor LLM costs.
  • Spot drift. Watch the validation failure rate and the average output length. A jump means something upstream changed.

When do you need a dedicated platform?

When several teams share the data, when you need side-by-side prompt comparison with a nice interface, or when volume makes your own table awkward. Until then, a table you understand beats a dashboard you do not.

References

Author

Raktim Ranjit is a software engineer and the founder of NodeDR Infotech. He builds and maintains the software described here.

Have something in mind?

Let’s build something useful.

Tell me about the idea, product, or workflow you’re working through.

Tap to say hello