← All articles
Observability

LLM Observability: What to monitor in production

A production guide to LLM observability - the signals that matter, a runnable OpenTelemetry trace, the dashboard and alerts worth having, what to keep, and a complaint-to-root-cause walkthrough.

June 2, 2026 · updated September 27, 2026 · 7 min read · observability · monitoring · tracing

Observability is the difference between knowing a request was slow and knowing which retrieval step, tool call or model hop caused it. It’s the data layer the rest of LLMOps depends on - you can’t evaluate, control cost or run an incident without it.

The two questions

Prototype: “Did it work when I tried it?” Production: “When a request misbehaves at 2am, can I see exactly why - and is p95 latency holding under real load?”

In the demo you watch one request succeed. In production you need to reconstruct any request after the fact, and watch aggregate health across all of them.

What breaks without it

  • Every issue is a guess. A user reports a bad answer; with no record of what the model saw, you’re debugging blind.
  • Multi-step systems are opaque. An agent returns the wrong result and you can’t tell which of six steps went wrong.
  • Slow creep goes unseen. p95 latency drifts up with traffic; averages hide it until users churn.
  • Cost is a quarterly surprise. Without per-request token accounting you learn about a runaway feature from the invoice.
  • Evals have no raw material. The best eval cases are real production failures - which you can only mine if you logged them.

The signals that matter

  • Request / response logs - every input and output, redacted as needed. You cannot debug what you did not record.
  • Traces - for chains and agents, a tree of spans following one request across every retrieval, tool call and model hop. See what to log in an LLM trace for the field-by-field schema.
  • Latency - p50 and p95, ideally split by step (retrieval vs generation), not just an average.
  • Token usage - per request and per feature, feeding straight into cost control.
  • Failure modes - timeouts, refusals, malformed output, tool errors, empty/low-score retrieval.
  • Outcome signals - user feedback (thumbs, escalations) tied back to the trace.

Instrument it

The cheapest path is a drop-in SDK. Langfuse, for example, captures latency, tokens and cost on every call with no extra code:

# pip install langfuse  -  drop-in replacement for the OpenAI client
from langfuse.openai import OpenAI

client = OpenAI()
resp = client.chat.completions.create(
    model=MODEL,  # your pinned model id
    messages=[{"role": "user", "content": question}],
)
# Latency, token usage and cost now land in your trace automatically.

If you want vendor-neutral telemetry that lives in your existing stack, follow the OpenTelemetry GenAI semantic conventions (linked below) - a standard set of gen_ai.* span attributes. A chat call becomes a span named after the operation and the model:

span: chat <model>
  gen_ai.operation.name          = "chat"
  gen_ai.provider.name           = "openai"
  gen_ai.request.model           = "<model>"
  gen_ai.usage.input_tokens      = 1240
  gen_ai.usage.output_tokens     = 310
  gen_ai.response.finish_reasons = ["stop"]

Standardising on these attributes means your traces, metrics and dashboards aren’t locked to one vendor. Two caveats: the conventions are still marked as in development, and they have changed - gen_ai.provider.name replaced the older gen_ai.system - so pin the version your instrumentation emits and re-check it when you upgrade.

A trace you can run

Auto-instrumentation covers the model call. The spans that make a trace useful, such as retrieval, the prompt version and tool calls, are yours to add. This runs as-is with pip install opentelemetry-sdk (the model and retriever are stand-ins, so it needs no API key):

# pip install opentelemetry-sdk
from opentelemetry import trace
from opentelemetry.sdk.trace import TracerProvider
from opentelemetry.sdk.trace.export import SimpleSpanProcessor
from opentelemetry.sdk.trace.export.in_memory_span_exporter import InMemorySpanExporter

exporter = InMemorySpanExporter()  # in production: an OTLP exporter to your backend
provider = TracerProvider()
provider.add_span_processor(SimpleSpanProcessor(exporter))
trace.set_tracer_provider(provider)
tracer = trace.get_tracer("support-bot")

MODEL = "my-pinned-model"

def retrieve(question):  # stand-in for your retriever
    return [{"id": "policy/refunds#window", "score": 0.83}]

def call_model(prompt):  # stand-in for your provider SDK
    return {"text": "Refunds close 30 days after purchase.",
            "input_tokens": 1240, "output_tokens": 310, "finish_reason": "stop"}

def answer(question, prompt_version="rag-answer-v4"):
    with tracer.start_as_current_span("answer") as root:
        root.set_attribute("app.prompt_version", prompt_version)
        with tracer.start_as_current_span("retrieve") as span:
            docs = retrieve(question)
            span.set_attribute("app.retrieval.doc_ids", [d["id"] for d in docs])
            span.set_attribute("app.retrieval.top_score", docs[0]["score"])
        with tracer.start_as_current_span(f"chat {MODEL}", kind=trace.SpanKind.CLIENT) as span:
            span.set_attribute("gen_ai.operation.name", "chat")
            span.set_attribute("gen_ai.provider.name", "openai")
            span.set_attribute("gen_ai.request.model", MODEL)
            out = call_model(question)
            span.set_attribute("gen_ai.usage.input_tokens", out["input_tokens"])
            span.set_attribute("gen_ai.usage.output_tokens", out["output_tokens"])
            span.set_attribute("gen_ai.response.finish_reasons", [out["finish_reason"]])
        return out["text"]

answer("Can I get a refund after 45 days?")

spans = {s.context.span_id: s for s in exporter.get_finished_spans()}
for s in sorted(spans.values(), key=lambda s: s.start_time):
    depth = 0 if s.parent is None else (1 if spans[s.parent.span_id].parent is None else 2)
    print("  " * depth + s.name, dict(s.attributes))

Its actual output (with opentelemetry-sdk 1.45):

answer {'app.prompt_version': 'rag-answer-v4'}
  retrieve {'app.retrieval.doc_ids': ('policy/refunds#window',), 'app.retrieval.top_score': 0.83}
  chat my-pinned-model {'gen_ai.operation.name': 'chat', 'gen_ai.provider.name': 'openai', 'gen_ai.request.model': 'my-pinned-model', 'gen_ai.usage.input_tokens': 1240, 'gen_ai.usage.output_tokens': 310, 'gen_ai.response.finish_reasons': ('stop',)}

Note the split: the model span carries only standard gen_ai.* attributes, so any GenAI-aware backend can read it, while your own context lives under an app.* prefix that will never collide with a future convention.

The dashboard worth having

Most LLM dashboards show everything and answer nothing. Five panels cover the questions you’ll actually ask:

PanelShowsAnswers
Latency by stepp50 and p95 for retrieval, model and tools, stacked”Where did the time go?”
Errors by typeTimeouts, rate limits, refusals, malformed output, tool errors”Is it us, the provider or the model?”
Tokens and cost per requestDistribution per feature, not a total”Did a change make every request dearer?”
Retrieval healthTop score distribution, empty-retrieval rate”Is the context still finding anything?”
OutcomesThumbs-down, escalations, retries per 1,000 requests”Are users actually worse off?”

For streaming interfaces, add time to first token alongside total latency - it’s what users feel. The OpenTelemetry conventions define client metrics for both, gen_ai.client.operation.duration and gen_ai.client.operation.time_to_first_chunk.

Alerts that predict pain

Alert against your own baseline, not a fixed number - a 3-second p95 is fine for one feature and an outage for another. A starting set:

AlertCondition (tune to your traffic)Why
Error spikeError rate above 2× its 7-day median for 10 minutesCatches provider incidents and broken deploys
Latency regressionp95 above 1.5× its 7-day median for 15 minutesSlow creep users notice before you do
Empty retrievalShare of requests with no chunk above the score threshold doublesIndex or embedding problem upstream
Cost driftCost per request up 30% day over dayA prompt or context change nobody priced
Feedback dropThumbs-down rate above 1.5× baseline over a dayQuality regression that errors don’t show

Every alert should link straight to a filtered list of the traces behind it - an alert that makes on-call go looking is half an alert.

What to keep, and for how long

Keeping every full prompt and response forever is expensive and a privacy liability; keeping nothing leaves you blind. A split that works:

  • Metadata for 100% of requests - ids, prompt version, model, tokens, cost, latency, scores, outcome. It’s small, and it’s what dashboards and alerts need.
  • Full payloads for every failure - errors, thumbs-down, escalations, low-grounding answers. These are your future eval cases.
  • Full payloads for a sample of successes - enough to spot-check quality and build baselines, on a shorter retention period.
  • Redaction before storage, and a retention policy you can state - see what to log in an LLM trace.

From complaint to root cause

What good observability looks like in practice. A customer reports a wrong answer about refund windows:

  1. Find the trace. Search by customer id and time - you land on one request, not a log file.
  2. Read the spans. Retrieval returned three chunks with a top score of 0.41, well below the usual 0.8, and none of them from the refund policy.
  3. Check the prompt version. Same as last week; the prompt isn’t the suspect.
  4. Check the index. The retrieval-health panel shows the empty-retrieval rate jumped two days ago, when the help centre was restructured - the policy page moved and the re-index didn’t pick up the new URL.
  5. Fix and keep. Re-index, then add the question to the eval set with the expected source, so a future move fails a test instead of a customer.

Without the retrieval span, the same investigation starts with “the model is hallucinating” and ends with a prompt change that fixes nothing.

From signal to action

Signals you don’t act on are decoration:

  • Alert on the things that predict user pain first - error-rate spikes and p95 latency regressions.
  • Tie traces to evals - pipe real failures into your eval set so the same failure becomes a permanent test.
  • Make it replayable - being able to re-run a production request with its exact inputs turns a vague report into a five-minute fix.

Minimal vs mature

AspectMinimalProduction-grade
CaptureRequest/response logsFull span traces (OTel / SDK)
LatencyAveragep50/p95, split by step; time to first token
CostTotal spendPer request, per feature
AlertsNoneRelative to baseline, linked to traces
RetentionEverything, forever (or nothing)Metadata always, payloads by outcome
DebuggingRead logsSearch + replay a single trace
Feedback loop-Failures flow into the eval set

Tools and where to go next

The observability category of the directory covers Langfuse, LangSmith, MLflow, Arize Phoenix and others - filterable by open-source, self-hostable and OpenTelemetry support. Once you can see your system, close the loop with evaluation and pressure-test the rest with the Production Checklist.

Get the Production Checklist → Explore the Stack →