LLM Evaluation: How to test prompts, RAG and agents
A production-grade guide to LLM evaluation - what breaks without it, what to measure, how to write an LLM-as-judge and an eval runner, and how to gate releases the way you gate code on tests.
Evaluation is the highest-leverage practice in LLMOps and the one teams most often skip. Get it right and every other change - a new prompt, a cheaper model, a different retriever - becomes a measured decision instead of a gamble. Get it wrong and you are shipping on vibes.
The two questions
Every layer of the stack is the gap between a prototype question and a production one. For evaluation:
Prototype: “Does it answer correctly in the demo I just ran?” Production: “Does it pass evals on real cases, every time we change something?”
The demo answers one input, once. Production has to answer thousands, after each edit, without regressing the cases that matter. An eval dataset - a curated set of inputs with checkable expectations - is what closes that gap.
What breaks in production without it
These are the failure modes an eval set prevents:
- You can’t tell if a change helped. A prompt tweak “feels better”; you ship it; a week later a different cohort is worse. With no baseline, you never find out which change did it.
- Silent regressions. A model or provider update subtly shifts behaviour. No eval gate means no alarm until users complain.
- A global pass rate that lies. “92% accuracy” can hide a 40% pass rate on the one workflow - refunds, medical, legal - where being wrong is expensive.
- Over-fitting. If you only ever tune against the same 20 cases, you optimise for them and regress everything else.
- Unrepresentative data. An eval set of inputs you imagined tests a product that doesn’t exist. Real traffic is harder and weirder.
What to measure
Pick metrics tied to how the system actually fails, not generic “quality”:
- Task success / accuracy - did it do the job? (exact match, or rubric.)
- Faithfulness - is the answer grounded in the provided context? The core metric for RAG.
- Safety & refusal - does it refuse what it should, and only that?
- Relevance & tone - usually judged by an LLM-as-judge with a rubric.
Always slice these by a risk_category so a healthy average can’t hide a broken
high-stakes category.
Instrument it: the eval row
Keep each case declarative, and split its checks in two: literal checks that a string or regex can decide on its own, and judge questions for anything that depends on what the answer means:
{
"input": "Can I get a refund after 45 days?",
"expected_behavior": "State the 30-day window, decline, offer a next step",
"literal": {
"must_match_all": ["(?i)\\b(30|thirty)[- ]days?\\b"],
"must_not_match": ["(?i)\\b(i|we)('ve| have)? (approved|processed|issued) (a|your|the) refund\\b"]
},
"judge": [
"Does the answer state or imply a refund window other than 30 days? Expect: no",
"Does the answer grant, promise or hint at an exception? Expect: no"
],
"risk_category": "customer_support_policy",
"critical": true
}
Literal checks are free and deterministic, so use them for everything a string can genuinely settle - and nothing else. A check like “must not include invented exceptions” looks literal but isn’t: no model ever writes that phrase, so it passes exactly the answers it was meant to catch. Whether an exception is being offered is meaning, which is a judge question. The eval dataset guide shows how to test each check against hand-labelled answers before it gates anything.
Instrument it: the LLM-as-judge
For subjective qualities (faithfulness, helpfulness), a second model scores the answer against a narrow rubric. Keep its job binary and grounded - return JSON so you can aggregate it:
import json
JUDGE_MODEL = "..." # an exact, pinned model id - never a "latest" alias
def faithfulness_judge(answer: str, context: str) -> dict:
prompt = (
"Using ONLY the context, is every claim in the answer supported?\n"
'Reply JSON: {"verdict": "supported|unsupported", "unsupported": [...]}\n\n'
f"Context:\n{context}\n\nAnswer:\n{answer}"
)
out = client.chat.completions.create(
model=JUDGE_MODEL,
messages=[{"role": "user", "content": prompt}],
response_format={"type": "json_object"},
)
return json.loads(out.choices[0].message.content)
Calibrate the judge before you trust it: have a human label ~30 cases and compare. If they agree, automate; if not, tighten the rubric. Re-calibrate when you change the judge model - which is also why the model id is pinned. The canonical study, Judging LLM-as-a-Judge with MT-Bench (linked in the sources below), found strong judges agreeing with human preferences over 80% of the time, and also documented their position, verbosity and self-preference biases.
Run it like a test suite
An eval that runs manually gets skipped. A minimal runner for the literal layer (the judge questions run the same way, one yes/no call each):
import re
def literal_pass(answer, literal):
return all(re.search(p, answer) for p in literal.get("must_match_all", [])) and not any(
re.search(p, answer) for p in literal.get("must_not_match", [])
)
def run_evals(dataset, answer_fn):
results = []
for case in dataset:
answer = answer_fn(case["input"])
results.append({"case": case, "passed": literal_pass(answer, case.get("literal", {}))})
pass_rate = sum(r["passed"] for r in results) / len(results)
critical_failures = [r for r in results if r["case"].get("critical") and not r["passed"]]
return pass_rate, critical_failures, results
Then gate it in CI so a regression can’t merge:
# .github/workflows/eval.yml (sketch)
on:
pull_request:
paths: ['prompts/**', 'src/**', 'evals/**']
jobs:
eval:
steps:
- run: python run_evals.py --data evals/
# fail the build if pass_rate drops below threshold
# or any case tagged "critical" regresses
This is the whole point: evaluation gives an LLM app the same safety net unit tests give normal code.
Run it more than once
The same input can pass on one run and fail on the next, so a single pass per
case measures luck as much as quality. Run each case several times at your
production settings - three to five is a practical start - and report two
numbers: the average pass rate, and the share of cases that passed on every
run. The gap between them is your flakiness, and flaky cases deserve a look of
their own: an answer that’s right four times in five is wrong for one user in
five. For critical cases, gate on the all-runs number.
Comparing two versions
Absolute scores drift with the judge; direct comparison is steadier. To decide between two prompt or model versions, show a judge both answers to the same input and ask which better meets the rubric - then ask again with the order swapped, because judges favour a position (the MT-Bench study cited below documents this):
def compare(case, answer_a, answer_b):
first = judge_prefers(case, answer_a, answer_b) # "first" | "second" | "tie"
second = judge_prefers(case, answer_b, answer_a)
if first == "first" and second == "second":
return "A"
if first == "second" and second == "first":
return "B"
return "tie" # the judge changed its mind when the order changed - no signal
Count only the verdicts that survive the swap. If most cases come back “tie”, the two versions are closer than any single score would suggest - which is itself a useful answer.
Minimal vs mature
Ship the minimal column to get a signal today; the mature column is what makes it trustworthy.
| Aspect | Minimal | Production-grade |
|---|---|---|
| Dataset | 20–50 real cases | 100s, growing from every incident |
| Scoring | Literal regex checks, tested | + LLM-as-judge, calibrated |
| Coverage | Top intents | Sliced by risk_category |
| Cadence | Run before a big change | Gated in CI on every change |
| Regressions | Spotted by hand | Block the deploy automatically |
| Non-determinism | One run per case | Repeated runs; critical cases must pass every run |
| Version choice | Compare two scores | Pairwise judge with swapped order |
Evaluating RAG and agents
- RAG: separate retrieval quality from generation quality - most “wrong answer” bugs are retrieval bugs. Measure whether the right context was fetched before you judge the answer. See RAGOps.
- Agents: evaluate at the step level, not just the final answer. Use the trace to check the agent chose the right tool with the right arguments - a correct final answer can still hide a wrong, expensive path.
Tools and where to go next
You don’t have to build the harness yourself - the evals category of the directory covers Braintrust, LangSmith, DeepEval, OpenAI Evals and Giskard, filterable by open-source, self-hostable and more.
Then turn it into a habit: the evaluation items in the Production Checklist are the bar to clear, and the Maturity Score tells you how your eval practice compares across the rest of the stack. For a guided route in, the learning page lists a free short course on evaluating and debugging generative AI.