← All articles
Governance

LLM incident response template

A ready-to-adapt incident response template for LLM applications - severity levels, the first 15 minutes, mitigation levers unique to LLMs, and a post-mortem structure.

June 15, 2026 · updated September 27, 2026 · 3 min read · governance · incident-response · template

LLM apps fail in ways classical services don’t: a prompt edit tanks quality, a provider ships a model update overnight, an injection makes the agent misbehave, or a traffic spike triples the bill. Your normal incident process mostly applies - but the levers are different. Here’s a template to adapt.

Severity levels

Define these before you need them:

SevExampleResponse
SEV1Harmful output, data leak, or agent taking wrong actions at scaleAll-hands, immediate mitigation, comms
SEV2Quality regression affecting many users; cost runawayOn-call owns, mitigate within the hour
SEV3Degraded latency, elevated error/refusal rateTriage in business hours

The LLM-specific additions to your usual list: harmful/unsafe output, prompt-injection exploitation, hallucination spike, and cost runaway.

The first 15 minutes

1. STOP THE BLEEDING
   - Roll back the last prompt / model / config change (one step - you versioned it)
   - If an agent is taking actions: disable the high-risk tools / force approval mode
   - If cost runaway: enable hard rate limits / budget cap

2. CONFIRM SCOPE
   - Which feature, which prompt_version, since when?
   - Pull traces for affected requests; check what changed (deploy log, provider status)

3. COMMUNICATE
   - Open the incident channel, assign a lead, post a one-line status

The first move is almost always revert to the last known-good version. The ability to do that in one step is why versioning and one-step rollback are non-negotiable in the Production Checklist.

Mitigation levers unique to LLMs

  • Roll back the prompt/model/config - the fastest fix for a quality or behaviour regression.
  • Switch the model / provider - if a provider update or outage is the cause and you have a gateway with fallback.
  • Tighten guardrails - raise input/output filtering, force human-in-the-loop for the affected workflow.
  • Disable tools or enable approval mode - for agents taking wrong actions.
  • Cap cost - hard rate limits and budget alerts to stop a runaway bill.
  • Fail safe - degrade to a canned response or human handoff rather than serving bad output.

Four playbooks

The four incidents that are specific to LLM systems each have a first move, a question that confirms the cause, and a fix that sticks:

IncidentFirst moveConfirm the causeMake it stick
Quality regressionRoll back the last prompt, model or retrieval changeDid the eval pass rate on sampled traffic drop right after a change in the deploy log? If nothing changed on your side, check the provider’s changelogAdd the failing inputs to the eval set; gate the next change on them
Provider outage or silent model changeFail over to the evaluated secondary, or degrade to a safe responseProvider status page, error codes by provider, and whether the response model id or fingerprint changedPin model versions; alert on fallback events and on unexpected model ids
Prompt injection exploitedDisable the affected tools or force approval modeTrace the injected text to its entry point - user input, a retrieved document, a tool resultAdd the payload to the red-team set; gate the tool in code, not in the prompt
Cost runawayEnforce the budget cap at the gateway; rate-limit the top callersCost per request versus its 7-day median: one caller, one feature, or every request?Per-user and per-feature caps; a drift alert on cost per request

The right-hand column is the part teams skip once the pressure is off, and it’s the only part that stops the same incident from paging someone next month.

What you need in place beforehand

An incident is the wrong time to discover you can’t see anything:

  • Traces with prompt_version, inputs and outputs, searchable.
  • A deploy/change log so you can correlate the incident with a change.
  • A subscription to your providers’ status pages.
  • A documented rollback path that on-call has actually practised.

Post-mortem structure

Blameless, and tuned for LLM causes:

- Summary & impact (who, how many, how long)
- Timeline (change → symptom → detection → mitigation → resolution)
- Root cause: prompt | model/provider | retrieval | data | infra | misuse
- Detection gap: why didn't an eval / alert catch this earlier?
- Action items:
    - Add a regression case to the eval set
    - Add / tune the alert that would have caught it
    - Fix the rollback or guardrail gap

The most valuable output of any LLM incident is a new eval case and a new alert - so the same failure can’t recur without something noticing. Tie this into your governance layer and the regulated workflow blueprint if you operate under audit requirements.

Get the Production Checklist → Explore the Stack →