How to calculate hallucination rate
A practical method for measuring LLM hallucination (faithfulness) rate in production - how to define it, sample it, judge it, and track it over time.
“The model hallucinates” is an anecdote. “Our hallucination rate rose from 3.0% to 5.5% after the last prompt change, measured on 1,000 sampled responses” is something you can act on. Here’s how to turn the anecdote into a number - and how to tell whether a change in that number is real.
Define it precisely first
Hallucination rate only means something once you pin down what counts. The most useful, measurable definition in production is faithfulness:
An output is unfaithful if it asserts something not supported by the source material it was given (retrieved context, tools, or the prompt).
This is sharper than “is it true?” - you’re measuring grounding, not omniscience. For a RAG system that’s exactly the right question. Decide up front whether you’re scoring per response or per claim (claim-level is stricter and more informative).
The measurement loop
1. Sample N responses from production (with their retrieved context)
2. For each, judge: is every claim supported by the context?
3. hallucination_rate = responses_with_any_unsupported_claim / N
4. Report it with a confidence interval, then track it over time
and slice by feature / risk_category
1. Sample representatively
Pull a random sample from real traces, including the context the model actually saw. Random matters: cherry-picked samples flatter you. Size matters too - 100-200 responses is enough to spot a large problem, but not a shift of a couple of points (see step 4).
2. Judge with a rubric
For each response, the judge (human or LLM-as-judge) answers a narrow question:
{
"response": "...",
"context": "...the documents the model was given...",
"verdict": "supported | unsupported",
"unsupported_claims": ["specific claim not in context"]
}
Resist a “partially supported” verdict. By the definition above, one unsupported claim makes the response unfaithful - a middle category only gives the headline number somewhere to hide. If you want more resolution, score per claim instead.
An LLM-judge prompt that works: “Given ONLY the context below, is every factual claim in the response supported by it? List any claim that is not.” Keep the judge’s job binary and grounded - don’t ask it to assess truth, only support.
In code, that judge is small - return structured JSON so you can aggregate it:
import json
JUDGE_MODEL = "..." # an exact, pinned model id - never a "latest" alias
def faithfulness_judge(answer: str, context: str) -> dict:
prompt = (
"Using ONLY the context, is every claim in the answer supported?\n"
'Reply JSON: {"verdict": "supported|unsupported", "unsupported": [...]}\n\n'
f"Context:\n{context}\n\nAnswer:\n{answer}"
)
out = client.chat.completions.create(
model=JUDGE_MODEL,
messages=[{"role": "user", "content": prompt}],
response_format={"type": "json_object"},
)
return json.loads(out.choices[0].message.content)
# hallucination_rate = unsupported_count / judged_count
3. Calibrate the judge
Before you trust the LLM-judge, have a human label ~30 of the same cases and compare. If they agree most of the time, automate; if not, tighten the rubric. Recalibrate when you change the judge model.
4. Put an interval on it
A rate measured on a sample is an estimate, and at the low rates you’re hoping for, the uncertainty is large relative to the number. Report a confidence interval, and before announcing that the rate “rose”, check the change could not be noise. Standard library only:
from math import sqrt, erf
def wilson(k, n, z=1.96):
"""95% Wilson score interval for k unfaithful responses out of n judged."""
p = k / n
centre = (p + z * z / (2 * n)) / (1 + z * z / n)
half = z * sqrt(p * (1 - p) / n + z * z / (4 * n * n)) / (1 + z * z / n)
return centre - half, centre + half
def two_proportion_p(k1, n1, k2, n2):
"""Two-sided p-value: did the rate really change between two samples?"""
p = (k1 + k2) / (n1 + n2)
se = sqrt(p * (1 - p) * (1 / n1 + 1 / n2))
z = abs(k1 / n1 - k2 / n2) / se
return 1 - erf(z / sqrt(2))
for n in (200, 1000):
before, after = round(0.03 * n), round(0.055 * n)
lo1, hi1 = wilson(before, n)
lo2, hi2 = wilson(after, n)
print(f"n={n}: before {before/n:.1%} [{lo1:.1%}-{hi1:.1%}] "
f"after {after/n:.1%} [{lo2:.1%}-{hi2:.1%}] "
f"p={two_proportion_p(before, n, after, n):.3f}")
Its actual output:
n=200: before 3.0% [1.4%-6.4%] after 5.5% [3.1%-9.6%] p=0.215
n=1000: before 3.0% [2.1%-4.3%] after 5.5% [4.2%-7.1%] p=0.006
The same move from 3.0% to 5.5% is noise at 200 samples - a result that size or bigger turns up about one time in five with no real change - and a real regression at 1,000. So size the sample to the change you need to detect: a couple of hundred responses catch a broken release, and it takes around a thousand to see a prompt edit that costs you two or three points.
5. Compute and slice
Headline number aside, slice by feature and risk_category. A 3% overall rate can
hide a 15% rate on the one workflow where a wrong answer is expensive.
Turn it into a guardrail
Once you can measure it, you can defend it:
- Gate releases on it - block a prompt or model change that pushes the rate up, the same way you’d block on a failing test.
- Add a runtime grounding check - a lightweight output check that flags
unsupported answers before they reach the user, logged as
grounded_answer: false. - Alert when the production rate drifts, which often signals a retrieval or freshness problem upstream rather than the model itself.
Don’t chase zero
Some unfaithfulness is unavoidable, and the cost of driving it to zero (refusing more, retrieving more, slower, pricier) may not be worth it. Pick a target appropriate to the risk: near-zero for regulated or medical use, more relaxed for low-stakes drafting. The point isn’t a perfect score - it’s a number you watch, gate on, and improve deliberately. Build the sampling into your eval dataset and it becomes part of every release.