LLM Governance: Audit trails, approvals and explainability
A practical guide to governing LLM applications - audit trails, approval gates, risk tiers, an evidence pack, a retention schedule, and what the EU AI Act asks for - so you can answer for the system, not just run it.
Most of the LLMOps stack is about making a system work. Governance is about being able to answer for it - to a customer, a regulator, or your own incident review. It’s the layer teams notice last and need most once an LLM makes decisions that matter.
The two questions
Prototype: “Does it give a good answer?” Production: “Who approved this, what exactly did it do, and can you prove it after the fact?”
A demo is judged on output quality. A governed system is judged on accountability: a record of what happened and why, a human in the loop where the stakes demand one, and the ability to reproduce any decision.
What breaks without it
- No audit trail. A customer disputes an outcome and you have no record of the input, the context, or what the model returned. You can’t investigate, let alone defend it.
- Unexplainable decisions. “The AI decided” is not an answer a regulator accepts. Without recorded rationale and sources, every decision is a black box.
- Unreviewed high-risk actions. The model takes an irreversible action - a refund, an account change, a clinical suggestion - with no human sign-off.
- Silent drift. A provider updates the model under you; last month’s decisions are no longer reproducible because nothing was version-pinned.
- Retention and privacy breaches. Sensitive inputs sit in logs forever with no policy - a liability waiting for a subject-access request.
- No owner. When something goes wrong, no one is accountable per layer.
The controls that matter
- Audit trail - capture inputs, retrieved context, outputs and decisions immutably, with who/what/when.
- Human-in-the-loop & approval gates - a designed review step for high-risk actions, not an afterthought.
- Explainability & reproducibility - record the rationale and sources behind a decision, and pin model + prompt versions so it can be reproduced.
- Data retention & privacy - a defined, enforced policy for how long records live and who can read them.
- Incident process - a path to triage, contain and learn from model failures.
- Accountability - a named owner for each layer of the stack.
Frameworks like the NIST AI Risk Management Framework and the OWASP Top 10 for LLM Applications (both cited below) are the reference points to map these controls against.
If you operate in the EU
The EU AI Act turns two of these controls into law for high-risk systems. Article 12 (record-keeping) requires that high-risk AI systems technically allow the automatic recording of events - logs - over the system’s lifetime. Article 14 (human oversight) requires that they be designed so natural persons can effectively oversee them while they’re in use. That is the audit trail and the approval gate from the list above, written into regulation.
The dates moved in 2026, so check them rather than remembering them. Per the European Commission, the Act became applicable on 2 August 2026, with exceptions. The AI Omnibus - Regulation (EU) 2026/1744, in force since 27 July 2026 - deferred the high-risk obligations to 2 December 2027 for the stand-alone use cases listed in Annex III (such as employment, education and critical infrastructure) and to 2 August 2028 for AI embedded in products covered by Annex I. Transparency rules - including telling people when they are interacting with an AI system (Article 50) - were among those taking effect in August 2026.
Whether your system is high-risk is a legal question, not an engineering one. But the engineering is the same either way, and a deferred deadline is extra time to build the logging and oversight properly, not a reason to skip them.
Instrument it: the decision record
Governance becomes real when every consequential decision writes an immutable record:
{
"decision_id": "dec_8842",
"request_id": "req_9931",
"actor": "support-agent-v3",
"model": "claude-sonnet-5",
"prompt_version": "v12-locked",
"inputs_redacted": true,
"sources": ["policy/refunds#window", "ticket/55831"],
"rationale": "Within 30-day window; auto-approved per policy R-12",
"human_review": "not_required",
"outcome": "refund_approved",
"retention_class": "7y",
"timestamp": "2026-06-08T09:14:02Z"
}
The prompt_version, model and sources fields are what make the decision
reproducible and explainable; human_review and retention_class are what
make it accountable and compliant.
Instrument it: the approval gate
For high-risk actions, separate deciding from acting - and require a gate:
def execute(decision):
if decision["risk"] == "high" and not decision.get("approved_by"):
return enqueue_for_human_review(decision) # nothing happens yet
record_audit(decision) # immutable log first
return perform_action(decision)
Nothing high-risk takes effect until a human (or an explicit policy) approves, and the audit record is written before the action, not after.
Tier your use cases
Not every LLM feature needs the same controls, and treating them all as high-risk is how governance becomes the team everyone routes around. Tier each use case once, write the tier down, and let it decide the controls:
| Tier | Typical use | Human review | Record kept | Change process |
|---|---|---|---|---|
| Low | Internal drafting, summaries a person reads before use | None required | Standard traces | Eval gate |
| Medium | Customer-facing answers, internal decisions with a human in the loop | Sampled review each week | Decision record per answer | Eval gate + owner sign-off |
| High | Actions with legal, financial or safety effect; regulated decisions | Approval before every action | Immutable decision record, long retention | Eval gate, sign-off, documented change record |
Two questions place most use cases: can the output take effect without a person reading it first? and would a wrong output harm someone who can’t easily undo it? Two yeses mean high. Revisit the tier whenever the feature gains a new tool or a new audience - that’s when use cases quietly move up.
The evidence pack
When a customer, auditor or regulator asks how a system is governed, you want to hand over a folder, not start an investigation. Keep it current per system:
- System description - what it does, for whom, its tier, and what it must never do.
- Model and prompt inventory - pinned model ids and prompt versions in production, with dates (the same list the migration playbook asks you to keep).
- Evaluation results - the latest eval run, the eval set’s composition, and the history of release gates passed or failed.
- Oversight design - which actions need approval, who approves, and the approval log.
- Incident log - what went wrong, what was changed, and the eval case added afterwards (see the incident template).
- Data handling - what is logged, what is redacted, retention periods, and who can read the traces.
Most of it should come straight out of systems you already run - the eval harness, the trace store, the change log - which is the real test of whether your governance is operational or just documented.
Retention, written down
“Forever” is not a retention policy, and neither is “whatever the logging tool defaults to”. A schedule per record type makes the trade-off explicit:
| Record | Contains | Retention (example) |
|---|---|---|
| Request metadata | Ids, versions, tokens, cost, scores | 13 months, for year-over-year comparison |
| Full prompts and responses | Customer text, retrieved context | 30-90 days, redacted, restricted access |
| Failure payloads kept as eval cases | Minimised, anonymised examples | While the case is in the eval set |
| High-tier decision records | Inputs, sources, rationale, approver | As long as the decision can be challenged |
The periods here are examples, not advice - your legal and data-protection obligations set the real numbers. What matters is that each record type has one, and that deletion actually runs.
Minimal vs mature
| Aspect | Minimal | Production-grade |
|---|---|---|
| Audit | Basic logs | Immutable, access-controlled trail |
| Review | Ad-hoc | Mandatory sign-off for high-risk |
| Explainability | None | Rationale + sources per decision |
| Versioning | Latest model | Pinned model + prompt versions |
| Retention | Default / forever | Schedule per record type, deletion enforced |
| Risk | Every use case treated the same | Tiered once, controls follow the tier |
| Evidence | Assembled when someone asks | Evidence pack kept current per system |
| Ownership | Unclear | Named owner per layer |
Where this lives in a real system
Governance isn’t a bolt-on - it’s wired through the architecture. See the regulated LLM workflow reference architecture for where the policy gate, human-review queue and audit store sit, and the governance items in the Production Checklist for the bar to clear. For the bigger picture, what is LLMOps? frames how governance relates to the rest of the stack.