AI
What is LLM observability, and what should you log?
LLM observability means recording enough about every model call to debug, evaluate and cost it. What to capture, what to redact, and how to start with a table and a few queries.
By Raktim Ranjit · Published · 3 min read
Short answer: LLM observability is the practice of recording each model call and each step around it, so that when an answer is wrong, slow or expensive you can see exactly what went in, what came out and what it cost. At minimum that is the prompt, the response, the model and parameters, token counts, latency, errors, and a trace ID that links the steps of one user request.
Why is normal logging not enough?
A web request either works or throws. A model call returns something plausible even when it is wrong. There is no stack trace for "the summary left out the refund". You need the actual inputs and outputs to judge it, and you need to be able to find the bad ones among thousands.
What should you capture for every call?
- Timestamp, environment and a trace ID shared by every step of one request.
- Model name, version and parameters like temperature and max tokens.
- The full prompt, including system prompt, retrieved documents and tool definitions, or a stored reference to them.
- The full response, including tool calls the model asked for.
- Input tokens, output tokens and, if your provider reports it, cached tokens.
- Latency to first token and total latency.
- Status: success, refusal, timeout, rate limit, schema validation failure.
- The prompt template version, so you can tie a change in quality to a change in text.
- User feedback if you collect it, such as a thumbs-down or an edit.
What about agents and multi-step flows?
Use spans. A user request is a trace. Each model call, each retrieval and each tool call is a span inside it, with a parent. Reading a trace top to bottom shows where time went and where the chain went wrong. OpenTelemetry has semantic conventions for generative AI spans, which means you can use your existing tracing backend instead of adopting a new product just for this.
What should you not store?
Prompts contain whatever users typed, and users type personal data. Decide what you keep before you start. Redact emails, phone numbers, card numbers and government IDs at write time. Set a retention period, 30 days is common for raw prompts, and keep only aggregates beyond that. If you operate under a data protection law, a log full of raw prompts is personal data and needs the same care as a database.
Can you start without a vendor tool?
Yes. One table in PostgreSQL is enough to learn what you need.
CREATE TABLE llm_calls (
id bigserial PRIMARY KEY,
trace_id uuid NOT NULL,
parent_id bigint,
created_at timestamptz NOT NULL DEFAULT now(),
model text NOT NULL,
prompt_ver text,
input_tokens int,
output_tokens int,
latency_ms int,
status text NOT NULL,
prompt jsonb,
response jsonb,
feedback smallint
);
CREATE INDEX ON llm_calls (trace_id);
CREATE INDEX ON llm_calls (created_at);Then write three queries and look at them weekly: the slowest ten percent of calls, the most expensive traces, and every call with negative feedback or a validation failure.
What do you do with the data?
- Debug. Open the trace for a complaint and read what the model actually saw.
- Build an evaluation set. Copy bad cases into a test file with the answer you would accept, and run it on every prompt change.
- Control cost. Group by feature and prompt version. See how to monitor LLM costs.
- Spot drift. Watch the validation failure rate and the average output length. A jump means something upstream changed.
When do you need a dedicated platform?
When several teams share the data, when you need side-by-side prompt comparison with a nice interface, or when volume makes your own table awkward. Until then, a table you understand beats a dashboard you do not.
References
Author
Raktim Ranjit is a software engineer and the founder of NodeDR Infotech. He builds and maintains the software described here.