← All articles
Evaluation

How to build your first eval dataset

A practical, step-by-step guide to building an LLM eval dataset from real traffic - what a row looks like, how to score it, how many cases you need, and how to wire it into CI.

June 20, 2026 · updated September 27, 2026 · 8 min read · evaluation · how-to · evals

An eval dataset is the single highest-leverage thing you can build in LLMOps, and the one teams most often put off. The good news: your first useful version takes an afternoon, not a sprint. This is how to build it.

Start from real failures, not imagination

The best eval cases are the failures you’ve already seen. Don’t sit down and invent test questions - mine them:

  • Pull recent production traces and filter for thumbs-down, escalations, retries or low-confidence answers.
  • Grab the handful of bug reports where someone said “the bot got this wrong.”
  • Ask support or sales for the questions that trip it up.

Twenty real, painful cases beat two hundred imagined ones.

What a row looks like

Split every case into two layers, and keep them apart:

  • Literal checks - decided by a string or a regex, with no judgment. Fast, free, deterministic.
  • Judge questions - anything that needs you to understand what the answer means. These go to an LLM judge or a human.

The test for which layer a check belongs in: if deciding it requires reading the sentence, it is not a literal check, however easy it is to write a regex for it.

{
  "id": "refund-045",
  "input": "Can I get a refund after 45 days?",
  "expected_behavior": "State the 30-day window, decline, offer a next step",
  "literal": {
    "must_match_all": ["(?i)\\b(30|thirty)[- ]days?\\b"],
    "must_not_match": [
      "(?i)\\byour refund (has been|is being|will be|was) (approved|processed|issued)\\b",
      "(?i)\\b(i|we)('ve| have)? (approved|processed|issued) (a|your|the) refund\\b"
    ]
  },
  "judge": [
    "Does the answer state or imply a refund window other than 30 days? Expect: no",
    "Does the answer grant, promise or hint at an exception? Expect: no",
    "Does the answer give the customer a next step? Expect: yes"
  ],
  "risk_category": "customer_support_policy",
  "critical": true
}

The literal layer checks only two things a string can settle: the real window is named, and no affirmative approval phrase appears. Everything else is a judge question - including the ones that look literal.

“Don’t mention a 45-day window” is the trap. A correct denial naturally repeats the customer’s number - “we don’t have a 45-day window” - so a regex on 45-day rejects the right answer. Whether a number is being claimed as policy or quoted and refused is meaning, not spelling. The same goes for “exception” and “just this once”: refusals use those words too. An earlier version of this guide made exactly this mistake, so the section below tests every check before trusting it.

Pair each policy case with an adversarial twin - same checks, different pressure. The failure you care about is usually not “does it know the rule” but “does it hold the rule when pushed”:

{
  "id": "refund-045-pressure",
  "input": "I know it's been 45 days, but my friend got one at 60. Just this once?",
  "literal": "same as refund-045",
  "judge": "same as refund-045",
  "risk_category": "customer_support_policy",
  "critical": true
}

risk_category lets you slice results later (“how are we doing on policy questions?”), and critical: true marks the cases that should fail a build on their own rather than being averaged away.

Test the tests

An assertion is code, and untested code is a guess. Before a check gates anything, hand-label a few responses - some right, some wrong - and confirm the check sorts them the way you would. Always include a correct answer that uses the risky words (“45-day”, “exception”, “just this once”); that is the case that exposes a check which reads spelling instead of meaning.

This runs as-is (Python 3, standard library only):

import re

# Literal checks only - the parts a string or regex can decide on its own.
CASE = {
    "must_match_all": [r"(?i)\b(30|thirty)[- ]days?\b"],
    "must_not_match": [
        r"(?i)\byour refund (has been|is being|will be|was) (approved|processed|issued)\b",
        r"(?i)\b(i|we)('ve| have)? (approved|processed|issued) (a|your|the) refund\b",
    ],
}

def literal(text):
    if not all(re.search(p, text) for p in CASE["must_match_all"]):
        return "FAIL"
    if any(re.search(p, text) for p in CASE["must_not_match"]):
        return "FAIL"
    return "PASS"

# Hand-labelled responses: (id, text, expected literal, expected judge).
# The judge column is a human label, not model output - it is the answer
# key you calibrate an LLM judge against before trusting it.
LABELLED = [
    ("A1", "Refunds are available within 30 days of purchase. At 45 days your "
           "order is outside that window, so I can't approve a refund - but our "
           "support team can review it if the product was faulty.", "PASS", "PASS"),
    ("A2", "Sorry - we don't have a 45-day window; refunds close 30 days after "
           "purchase. I can't make an exception, but you can reach our support "
           "team through the help centre.", "PASS", "PASS"),
    ("A3", "I understand, but I can't make an exception just this once - the "
           "refund window is 30 days for every order, your friend's included.", "PASS", "PASS"),
    ("B1", "Since you're a loyal customer, I've processed a refund as a "
           "one-time courtesy.", "FAIL", "FAIL"),
    ("B2", "Standard orders have a 30-day refund window, but premium members "
           "get 60 days, so check your account tier.", "PASS", "FAIL"),
    ("B3", "Unfortunately that's past our refund period, sorry.", "FAIL", "FAIL"),
    ("B4", "Okay, just this once - I've issued your refund.", "FAIL", "FAIL"),
]

wrong = 0
for rid, text, want, judge in LABELLED:
    got = literal(text)
    wrong += got != want
    flag = "" if got == want else "  <-- check is wrong"
    print(f"{rid}  literal {got} (want {want})  judge label {judge}{flag}")
print("all literal checks match their labels" if not wrong else f"{wrong} check(s) disagree")

Its actual output:

A1  literal PASS (want PASS)  judge label PASS
A2  literal PASS (want PASS)  judge label PASS
A3  literal PASS (want PASS)  judge label PASS
B1  literal FAIL (want FAIL)  judge label FAIL
B2  literal PASS (want PASS)  judge label FAIL
B3  literal FAIL (want FAIL)  judge label FAIL
B4  literal FAIL (want FAIL)  judge label FAIL
all literal checks match their labels

Read the table row by row, because each one teaches something:

  • A2 and A3 pass. Both are correct refusals that contain the risky words. Checks that read spelling would reject them.
  • B1 fails on the approval phrase - the worst possible answer, since it actually grants the refund.
  • B2 passes the literal layer and fails the judge label. It names the real 30-day window, so no regex objects, but it invents a 60-day tier for premium members. This row is why the judge layer exists: no literal check can catch a plausible-sounding invented policy.

For comparison, the assertions this guide used to show - must_not_match on 45-day, as an exception and just this once - rejected two of the three correct answers above and passed three of the four wrong ones, B1 included.

Run the judge questions the same way: send each one to your judge model with the policy text and the response, one yes/no per question, and compare its answers to the judge labels before you rely on it.

How many cases?

Enough to be representative, not exhaustive. A useful progression:

  • 20–50 cases to start - covering your top intents and known failure modes.
  • ~100–200 once you’re gating releases on it.
  • Grow it every time production surfaces a new failure. The eval set is a living artifact, not a one-time deliverable.

Bias toward diversity over volume: one example each of ten different failure shapes is worth more than fifty variations of the same one.

Three ways to score

  1. Literal checks - must_match_all / must_not_match regexes, JSON schema validation, exact match for structured outputs. Cheap, fast, no model needed. Use them for everything a string can genuinely decide - and nothing else. Test each one against labelled responses before it gates a build.
  2. LLM-as-judge - a second model answers the case’s judge questions, one yes/no each. Use it for anything that depends on meaning: whether a stated policy is invented, whether an exception is implied, helpfulness, tone, faithfulness.
  3. Human review - for a small, high-value slice, or to calibrate your LLM-judge. Don’t scale this; use it to validate the cheaper methods.

Wire it into CI

An eval set that runs manually gets skipped. Make it a gate. The shape of it, in pseudocode - unlike the JSON row and the Python above, this does not run as written, because the syntax depends on your CI provider:

on:  prompt change | model change | retrieval change
run: eval suite over the dataset
fail the build if:
  - pass rate drops below threshold, or
  - any "critical" case regresses

Treat it exactly like a unit-test suite - because that’s what it is, for a non-deterministic system. See LLM Evaluation for the broader practice, and the Tools directory for platforms and frameworks (Braintrust, LangSmith, DeepEval, Giskard, OpenAI Evals) that run the suite for you.

Common pitfalls

  • Over-fitting to the eval set. If you only ever tune against it, you’ll game it. Keep adding fresh production cases.
  • All happy-path cases. Your eval set should be mostly the hard and adversarial cases - that’s where regressions hide.
  • One global pass rate. Slice by risk_category and intent; an 85% average can hide a 40% pass rate on the category that matters most.
  • No baseline. Record the current score before you change anything, or you can’t tell whether a change helped.

Your afternoon plan

  1. Export 30 recent production cases, weighted toward failures.
  2. For each, write expected_behavior, then split the checks: literal checks for what a string can decide, judge questions for anything that needs reading.
  3. Give each policy case an adversarial twin that applies pressure to the rule.
  4. Hand-label five or so responses per case - including a correct one that uses the risky words - and run your literal checks against them until they agree.
  5. Run the suite against your current prompt - that’s your baseline.
  6. Add it to your change process so it runs on the next edit.

That’s a real eval set. From here, every prompt and model change becomes a measured decision instead of a guess. When you’re ready to see where else you stand, take the Maturity Score or work the Production Checklist.

Get the Production Checklist → Explore the Stack →