LLM Observability: What to monitor in production
A production guide to LLM observability - the signals that matter, a runnable OpenTelemetry trace, the dashboard and alerts worth having, what to keep, and a complaint-to-root-cause walkthrough.
Observability is the difference between knowing a request was slow and knowing which retrieval step, tool call or model hop caused it. It’s the data layer the rest of LLMOps depends on - you can’t evaluate, control cost or run an incident without it.
The two questions
Prototype: “Did it work when I tried it?” Production: “When a request misbehaves at 2am, can I see exactly why - and is p95 latency holding under real load?”
In the demo you watch one request succeed. In production you need to reconstruct any request after the fact, and watch aggregate health across all of them.
What breaks without it
- Every issue is a guess. A user reports a bad answer; with no record of what the model saw, you’re debugging blind.
- Multi-step systems are opaque. An agent returns the wrong result and you can’t tell which of six steps went wrong.
- Slow creep goes unseen. p95 latency drifts up with traffic; averages hide it until users churn.
- Cost is a quarterly surprise. Without per-request token accounting you learn about a runaway feature from the invoice.
- Evals have no raw material. The best eval cases are real production failures - which you can only mine if you logged them.
The signals that matter
- Request / response logs - every input and output, redacted as needed. You cannot debug what you did not record.
- Traces - for chains and agents, a tree of spans following one request across every retrieval, tool call and model hop. See what to log in an LLM trace for the field-by-field schema.
- Latency - p50 and p95, ideally split by step (retrieval vs generation), not just an average.
- Token usage - per request and per feature, feeding straight into cost control.
- Failure modes - timeouts, refusals, malformed output, tool errors, empty/low-score retrieval.
- Outcome signals - user feedback (thumbs, escalations) tied back to the trace.
Instrument it
The cheapest path is a drop-in SDK. Langfuse, for example, captures latency, tokens and cost on every call with no extra code:
# pip install langfuse - drop-in replacement for the OpenAI client
from langfuse.openai import OpenAI
client = OpenAI()
resp = client.chat.completions.create(
model=MODEL, # your pinned model id
messages=[{"role": "user", "content": question}],
)
# Latency, token usage and cost now land in your trace automatically.
If you want vendor-neutral telemetry that lives in your existing stack, follow the
OpenTelemetry GenAI semantic conventions (linked below) - a standard set of
gen_ai.* span attributes. A chat call becomes a span named after the operation
and the model:
span: chat <model>
gen_ai.operation.name = "chat"
gen_ai.provider.name = "openai"
gen_ai.request.model = "<model>"
gen_ai.usage.input_tokens = 1240
gen_ai.usage.output_tokens = 310
gen_ai.response.finish_reasons = ["stop"]
Standardising on these attributes means your traces, metrics and dashboards
aren’t locked to one vendor. Two caveats: the conventions are still marked as
in development, and they have changed - gen_ai.provider.name replaced the
older gen_ai.system - so pin the version your instrumentation emits and
re-check it when you upgrade.
A trace you can run
Auto-instrumentation covers the model call. The spans that make a trace
useful, such as retrieval, the prompt version and tool calls, are yours to
add. This runs as-is
with pip install opentelemetry-sdk (the model and retriever are stand-ins, so
it needs no API key):
# pip install opentelemetry-sdk
from opentelemetry import trace
from opentelemetry.sdk.trace import TracerProvider
from opentelemetry.sdk.trace.export import SimpleSpanProcessor
from opentelemetry.sdk.trace.export.in_memory_span_exporter import InMemorySpanExporter
exporter = InMemorySpanExporter() # in production: an OTLP exporter to your backend
provider = TracerProvider()
provider.add_span_processor(SimpleSpanProcessor(exporter))
trace.set_tracer_provider(provider)
tracer = trace.get_tracer("support-bot")
MODEL = "my-pinned-model"
def retrieve(question): # stand-in for your retriever
return [{"id": "policy/refunds#window", "score": 0.83}]
def call_model(prompt): # stand-in for your provider SDK
return {"text": "Refunds close 30 days after purchase.",
"input_tokens": 1240, "output_tokens": 310, "finish_reason": "stop"}
def answer(question, prompt_version="rag-answer-v4"):
with tracer.start_as_current_span("answer") as root:
root.set_attribute("app.prompt_version", prompt_version)
with tracer.start_as_current_span("retrieve") as span:
docs = retrieve(question)
span.set_attribute("app.retrieval.doc_ids", [d["id"] for d in docs])
span.set_attribute("app.retrieval.top_score", docs[0]["score"])
with tracer.start_as_current_span(f"chat {MODEL}", kind=trace.SpanKind.CLIENT) as span:
span.set_attribute("gen_ai.operation.name", "chat")
span.set_attribute("gen_ai.provider.name", "openai")
span.set_attribute("gen_ai.request.model", MODEL)
out = call_model(question)
span.set_attribute("gen_ai.usage.input_tokens", out["input_tokens"])
span.set_attribute("gen_ai.usage.output_tokens", out["output_tokens"])
span.set_attribute("gen_ai.response.finish_reasons", [out["finish_reason"]])
return out["text"]
answer("Can I get a refund after 45 days?")
spans = {s.context.span_id: s for s in exporter.get_finished_spans()}
for s in sorted(spans.values(), key=lambda s: s.start_time):
depth = 0 if s.parent is None else (1 if spans[s.parent.span_id].parent is None else 2)
print(" " * depth + s.name, dict(s.attributes))
Its actual output (with opentelemetry-sdk 1.45):
answer {'app.prompt_version': 'rag-answer-v4'}
retrieve {'app.retrieval.doc_ids': ('policy/refunds#window',), 'app.retrieval.top_score': 0.83}
chat my-pinned-model {'gen_ai.operation.name': 'chat', 'gen_ai.provider.name': 'openai', 'gen_ai.request.model': 'my-pinned-model', 'gen_ai.usage.input_tokens': 1240, 'gen_ai.usage.output_tokens': 310, 'gen_ai.response.finish_reasons': ('stop',)}
Note the split: the model span carries only standard gen_ai.* attributes, so
any GenAI-aware backend can read it, while your own context lives under an
app.* prefix that will never collide with a future convention.
The dashboard worth having
Most LLM dashboards show everything and answer nothing. Five panels cover the questions you’ll actually ask:
| Panel | Shows | Answers |
|---|---|---|
| Latency by step | p50 and p95 for retrieval, model and tools, stacked | ”Where did the time go?” |
| Errors by type | Timeouts, rate limits, refusals, malformed output, tool errors | ”Is it us, the provider or the model?” |
| Tokens and cost per request | Distribution per feature, not a total | ”Did a change make every request dearer?” |
| Retrieval health | Top score distribution, empty-retrieval rate | ”Is the context still finding anything?” |
| Outcomes | Thumbs-down, escalations, retries per 1,000 requests | ”Are users actually worse off?” |
For streaming interfaces, add time to first token alongside total latency -
it’s what users feel. The OpenTelemetry conventions define client metrics for
both, gen_ai.client.operation.duration and
gen_ai.client.operation.time_to_first_chunk.
Alerts that predict pain
Alert against your own baseline, not a fixed number - a 3-second p95 is fine for one feature and an outage for another. A starting set:
| Alert | Condition (tune to your traffic) | Why |
|---|---|---|
| Error spike | Error rate above 2× its 7-day median for 10 minutes | Catches provider incidents and broken deploys |
| Latency regression | p95 above 1.5× its 7-day median for 15 minutes | Slow creep users notice before you do |
| Empty retrieval | Share of requests with no chunk above the score threshold doubles | Index or embedding problem upstream |
| Cost drift | Cost per request up 30% day over day | A prompt or context change nobody priced |
| Feedback drop | Thumbs-down rate above 1.5× baseline over a day | Quality regression that errors don’t show |
Every alert should link straight to a filtered list of the traces behind it - an alert that makes on-call go looking is half an alert.
What to keep, and for how long
Keeping every full prompt and response forever is expensive and a privacy liability; keeping nothing leaves you blind. A split that works:
- Metadata for 100% of requests - ids, prompt version, model, tokens, cost, latency, scores, outcome. It’s small, and it’s what dashboards and alerts need.
- Full payloads for every failure - errors, thumbs-down, escalations, low-grounding answers. These are your future eval cases.
- Full payloads for a sample of successes - enough to spot-check quality and build baselines, on a shorter retention period.
- Redaction before storage, and a retention policy you can state - see what to log in an LLM trace.
From complaint to root cause
What good observability looks like in practice. A customer reports a wrong answer about refund windows:
- Find the trace. Search by customer id and time - you land on one request, not a log file.
- Read the spans. Retrieval returned three chunks with a top score of 0.41, well below the usual 0.8, and none of them from the refund policy.
- Check the prompt version. Same as last week; the prompt isn’t the suspect.
- Check the index. The retrieval-health panel shows the empty-retrieval rate jumped two days ago, when the help centre was restructured - the policy page moved and the re-index didn’t pick up the new URL.
- Fix and keep. Re-index, then add the question to the eval set with the expected source, so a future move fails a test instead of a customer.
Without the retrieval span, the same investigation starts with “the model is hallucinating” and ends with a prompt change that fixes nothing.
From signal to action
Signals you don’t act on are decoration:
- Alert on the things that predict user pain first - error-rate spikes and p95 latency regressions.
- Tie traces to evals - pipe real failures into your eval set so the same failure becomes a permanent test.
- Make it replayable - being able to re-run a production request with its exact inputs turns a vague report into a five-minute fix.
Minimal vs mature
| Aspect | Minimal | Production-grade |
|---|---|---|
| Capture | Request/response logs | Full span traces (OTel / SDK) |
| Latency | Average | p50/p95, split by step; time to first token |
| Cost | Total spend | Per request, per feature |
| Alerts | None | Relative to baseline, linked to traces |
| Retention | Everything, forever (or nothing) | Metadata always, payloads by outcome |
| Debugging | Read logs | Search + replay a single trace |
| Feedback loop | - | Failures flow into the eval set |
Tools and where to go next
The observability category of the directory covers Langfuse, LangSmith, MLflow, Arize Phoenix and others - filterable by open-source, self-hostable and OpenTelemetry support. Once you can see your system, close the loop with evaluation and pressure-test the rest with the Production Checklist.