← All articles
Evaluation

LLM Evaluation: How to test prompts, RAG and agents

A production-grade guide to LLM evaluation - what breaks without it, what to measure, how to write an LLM-as-judge and an eval runner, and how to gate releases the way you gate code on tests.

June 1, 2026 · updated September 27, 2026 · 6 min read · evaluation · testing · evals

Evaluation is the highest-leverage practice in LLMOps and the one teams most often skip. Get it right and every other change - a new prompt, a cheaper model, a different retriever - becomes a measured decision instead of a gamble. Get it wrong and you are shipping on vibes.

The two questions

Every layer of the stack is the gap between a prototype question and a production one. For evaluation:

Prototype: “Does it answer correctly in the demo I just ran?” Production: “Does it pass evals on real cases, every time we change something?”

The demo answers one input, once. Production has to answer thousands, after each edit, without regressing the cases that matter. An eval dataset - a curated set of inputs with checkable expectations - is what closes that gap.

What breaks in production without it

These are the failure modes an eval set prevents:

  • You can’t tell if a change helped. A prompt tweak “feels better”; you ship it; a week later a different cohort is worse. With no baseline, you never find out which change did it.
  • Silent regressions. A model or provider update subtly shifts behaviour. No eval gate means no alarm until users complain.
  • A global pass rate that lies. “92% accuracy” can hide a 40% pass rate on the one workflow - refunds, medical, legal - where being wrong is expensive.
  • Over-fitting. If you only ever tune against the same 20 cases, you optimise for them and regress everything else.
  • Unrepresentative data. An eval set of inputs you imagined tests a product that doesn’t exist. Real traffic is harder and weirder.

What to measure

Pick metrics tied to how the system actually fails, not generic “quality”:

  • Task success / accuracy - did it do the job? (exact match, or rubric.)
  • Faithfulness - is the answer grounded in the provided context? The core metric for RAG.
  • Safety & refusal - does it refuse what it should, and only that?
  • Relevance & tone - usually judged by an LLM-as-judge with a rubric.

Always slice these by a risk_category so a healthy average can’t hide a broken high-stakes category.

Instrument it: the eval row

Keep each case declarative, and split its checks in two: literal checks that a string or regex can decide on its own, and judge questions for anything that depends on what the answer means:

{
  "input": "Can I get a refund after 45 days?",
  "expected_behavior": "State the 30-day window, decline, offer a next step",
  "literal": {
    "must_match_all": ["(?i)\\b(30|thirty)[- ]days?\\b"],
    "must_not_match": ["(?i)\\b(i|we)('ve| have)? (approved|processed|issued) (a|your|the) refund\\b"]
  },
  "judge": [
    "Does the answer state or imply a refund window other than 30 days? Expect: no",
    "Does the answer grant, promise or hint at an exception? Expect: no"
  ],
  "risk_category": "customer_support_policy",
  "critical": true
}

Literal checks are free and deterministic, so use them for everything a string can genuinely settle - and nothing else. A check like “must not include invented exceptions” looks literal but isn’t: no model ever writes that phrase, so it passes exactly the answers it was meant to catch. Whether an exception is being offered is meaning, which is a judge question. The eval dataset guide shows how to test each check against hand-labelled answers before it gates anything.

Instrument it: the LLM-as-judge

For subjective qualities (faithfulness, helpfulness), a second model scores the answer against a narrow rubric. Keep its job binary and grounded - return JSON so you can aggregate it:

import json

JUDGE_MODEL = "..."  # an exact, pinned model id - never a "latest" alias

def faithfulness_judge(answer: str, context: str) -> dict:
    prompt = (
        "Using ONLY the context, is every claim in the answer supported?\n"
        'Reply JSON: {"verdict": "supported|unsupported", "unsupported": [...]}\n\n'
        f"Context:\n{context}\n\nAnswer:\n{answer}"
    )
    out = client.chat.completions.create(
        model=JUDGE_MODEL,
        messages=[{"role": "user", "content": prompt}],
        response_format={"type": "json_object"},
    )
    return json.loads(out.choices[0].message.content)

Calibrate the judge before you trust it: have a human label ~30 cases and compare. If they agree, automate; if not, tighten the rubric. Re-calibrate when you change the judge model - which is also why the model id is pinned. The canonical study, Judging LLM-as-a-Judge with MT-Bench (linked in the sources below), found strong judges agreeing with human preferences over 80% of the time, and also documented their position, verbosity and self-preference biases.

Run it like a test suite

An eval that runs manually gets skipped. A minimal runner for the literal layer (the judge questions run the same way, one yes/no call each):

import re

def literal_pass(answer, literal):
    return all(re.search(p, answer) for p in literal.get("must_match_all", [])) and not any(
        re.search(p, answer) for p in literal.get("must_not_match", [])
    )

def run_evals(dataset, answer_fn):
    results = []
    for case in dataset:
        answer = answer_fn(case["input"])
        results.append({"case": case, "passed": literal_pass(answer, case.get("literal", {}))})
    pass_rate = sum(r["passed"] for r in results) / len(results)
    critical_failures = [r for r in results if r["case"].get("critical") and not r["passed"]]
    return pass_rate, critical_failures, results

Then gate it in CI so a regression can’t merge:

# .github/workflows/eval.yml (sketch)
on:
  pull_request:
    paths: ['prompts/**', 'src/**', 'evals/**']
jobs:
  eval:
    steps:
      - run: python run_evals.py --data evals/
      # fail the build if pass_rate drops below threshold
      # or any case tagged "critical" regresses

This is the whole point: evaluation gives an LLM app the same safety net unit tests give normal code.

Run it more than once

The same input can pass on one run and fail on the next, so a single pass per case measures luck as much as quality. Run each case several times at your production settings - three to five is a practical start - and report two numbers: the average pass rate, and the share of cases that passed on every run. The gap between them is your flakiness, and flaky cases deserve a look of their own: an answer that’s right four times in five is wrong for one user in five. For critical cases, gate on the all-runs number.

Comparing two versions

Absolute scores drift with the judge; direct comparison is steadier. To decide between two prompt or model versions, show a judge both answers to the same input and ask which better meets the rubric - then ask again with the order swapped, because judges favour a position (the MT-Bench study cited below documents this):

def compare(case, answer_a, answer_b):
    first = judge_prefers(case, answer_a, answer_b)   # "first" | "second" | "tie"
    second = judge_prefers(case, answer_b, answer_a)
    if first == "first" and second == "second":
        return "A"
    if first == "second" and second == "first":
        return "B"
    return "tie"   # the judge changed its mind when the order changed - no signal

Count only the verdicts that survive the swap. If most cases come back “tie”, the two versions are closer than any single score would suggest - which is itself a useful answer.

Minimal vs mature

Ship the minimal column to get a signal today; the mature column is what makes it trustworthy.

AspectMinimalProduction-grade
Dataset20–50 real cases100s, growing from every incident
ScoringLiteral regex checks, tested+ LLM-as-judge, calibrated
CoverageTop intentsSliced by risk_category
CadenceRun before a big changeGated in CI on every change
RegressionsSpotted by handBlock the deploy automatically
Non-determinismOne run per caseRepeated runs; critical cases must pass every run
Version choiceCompare two scoresPairwise judge with swapped order

Evaluating RAG and agents

  • RAG: separate retrieval quality from generation quality - most “wrong answer” bugs are retrieval bugs. Measure whether the right context was fetched before you judge the answer. See RAGOps.
  • Agents: evaluate at the step level, not just the final answer. Use the trace to check the agent chose the right tool with the right arguments - a correct final answer can still hide a wrong, expensive path.

Tools and where to go next

You don’t have to build the harness yourself - the evals category of the directory covers Braintrust, LangSmith, DeepEval, OpenAI Evals and Giskard, filterable by open-source, self-hostable and more.

Then turn it into a habit: the evaluation items in the Production Checklist are the bar to clear, and the Maturity Score tells you how your eval practice compares across the rest of the stack. For a guided route in, the learning page lists a free short course on evaluating and debugging generative AI.

Get the Production Checklist → Explore the Stack →