RAGOps: How to monitor retrieval quality
Most "the model is wrong" bugs are retrieval bugs. A production guide to measuring retrieval quality, building a retrieval eval set, hybrid search with rank fusion, citation checks, chunking experiments and permission-aware retrieval.
Retrieval-augmented generation grounds a model’s answer in documents you retrieve at query time - the pattern introduced by Lewis et al. (cited below). In production it fails in a specific, recurring way: the model gets blamed for what is really a retrieval problem. RAGOps is the practice of measuring and maintaining the retrieval half of the system.
Everything before Generate is retrieval - and that is where most RAG bugs live. The rest of this guide measures and hardens each stage.
The two questions
Prototype: “Does RAG work on the query I just tried?” Production: “Is the right context retrieved - measured - and is the index still true?”
A demo retrieves for one question. Production has to retrieve the right context across every query, and keep the corpus current as reality changes underneath it.
What breaks in production
- Retrieval bugs blamed on the model. The answer is wrong because the wrong chunks were fetched (or the right ones were missing) - no model swap fixes that.
- Stale index, confident answers. The corpus moved on; the system serves outdated policy or pricing with full confidence.
- Bad chunking. Chunks split a key fact across two pieces, so neither is retrieved whole.
- Embedding drift. A new embedding model or shifting data changes what’s “similar,” silently degrading recall.
- No fallback. Nothing relevant matches, so the model hallucinates instead of saying “I don’t have that.”
Measure retrieval, not just answers
Separate two questions and measure them independently:
- Did we retrieve the right context? - retrieval quality.
- Did the model use it well? - generation quality (faithfulness).
For retrieval, score against a small labelled set (query → the doc ids that should be retrieved):
def retrieval_scores(eval_set, retrieve, k=5):
hit, recall = 0, 0.0
for case in eval_set: # case = {"query", "relevant_ids"}
got = {d.id for d in retrieve(case["query"], k=k)}
relevant = set(case["relevant_ids"])
found = got & relevant
hit += bool(found) # at least one relevant doc came back
recall += len(found) / len(relevant) # share of the relevant docs that came back
n = len(eval_set)
return {"hit_rate@k": hit / n, "recall@k": recall / n}
The two numbers answer different questions. Hit rate asks whether the
answer had anything to work with; recall asks whether it had
everything, which matters as soon as an answer needs facts from more than one
document. Now chunking strategy, embedding model and k become evaluated
choices, not guesses - compare them by their effect on both before you touch
the prompt.
Build the retrieval set in an afternoon
You need query-to-document labels, and you don’t need thousands. Fifty to a hundred queries are enough to compare two configurations. Three sources, from best to fastest:
- Production queries with a known answer source. Pull real questions where the answer cited a document and the user was satisfied (no follow-up, no thumbs-down, no escalation). The cited document is your label - check a sample by hand.
- Your support team’s known pairs. “When people ask X, the answer is in article Y” is exactly the label you need, and support staff can list dozens from memory.
- Synthetic questions generated from chunks. Fast, but biased: a question written by looking at a chunk tends to reuse its wording, so it overstates how well retrieval will do on real, messier phrasing. Use it to fill gaps, not as the whole set.
Keep the questions that failed in production in the set permanently. They are the ones most likely to regress.
Instrument it: hybrid search and citation checks
Vector search finds paraphrases; keyword search finds exact tokens - error
codes, SKUs, product names - that embeddings tend to blur. Hybrid search runs
both and merges the ranked lists. Reciprocal rank fusion (RRF) does the merge
without comparing the two engines’ incompatible scores: each document earns
1 / (k + rank) from every list it appears in. Elasticsearch, for example,
defaults k to 60. The same file shows a literal check you should run on every
answer: the model may only cite documents it was actually given. This runs
as-is:
def rrf(rankings: list[list[str]], k: int = 60) -> list[tuple[str, float]]:
"""Reciprocal rank fusion: merge ranked lists without comparing their scores."""
scores: dict[str, float] = {}
for ranking in rankings:
for rank, doc_id in enumerate(ranking, start=1):
scores[doc_id] = scores.get(doc_id, 0.0) + 1.0 / (k + rank)
return sorted(scores.items(), key=lambda kv: kv[1], reverse=True)
# "error E-4012 on checkout" - keyword search nails the error code,
# vector search finds the paraphrased troubleshooting guide
keyword = ["kb/e-4012", "kb/checkout-errors", "kb/payment-limits"]
vector = ["kb/checkout-troubleshooting", "kb/checkout-errors", "kb/e-4012", "kb/cart-timeouts"]
for doc_id, score in rrf([keyword, vector])[:4]:
print(f"{doc_id:<28} {score:.4f}")
def citation_check(answer_citations: list[str], retrieved_ids: list[str]) -> list[str]:
"""Literal check: the answer may only cite what was actually retrieved."""
return sorted(set(answer_citations) - set(retrieved_ids))
print("\ncited but not retrieved:", citation_check(["kb/e-4012", "kb/refund-policy"], ["kb/e-4012", "kb/checkout-errors"]))
Its actual output:
kb/e-4012 0.0323
kb/checkout-errors 0.0323
kb/checkout-troubleshooting 0.0164
kb/payment-limits 0.0159
cited but not retrieved: ['kb/refund-policy']
Documents that both engines agree on rise to the top, and the exact-match hit for the error code stays there even though vector search ranked it third. A reranker then reorders the fused shortlist with a model that reads the query and each chunk together. The citation check is the cheapest hallucination detector you will ever run: a citation to a document that wasn’t retrieved is a fabrication by definition, and a string comparison finds it.
Treat chunking as an experiment
Chunk size is a trade-off with a measurable answer. Pick three or four candidates and run them against the same retrieval set:
| Variant | Hit rate@5 | Recall@5 | Tokens per request | Faithfulness |
|---|---|---|---|---|
| 300-token fixed chunks | ||||
| 800-token fixed chunks | ||||
| Split on headings | ||||
| Headings + 15% overlap |
Fill it in from your own data - the right answer depends on your documents. Then take the smallest context that holds recall: every extra token of context is paid for on every request, and more context is not automatically better context.
Permissions belong in the query
If different users may see different documents, filter by permission inside the retrieval query, as a metadata filter on the vector or keyword search. Filtering after retrieval is too late twice over: restricted text may already have reached the model, and dropping results after top-k leaves some users with too few relevant chunks. Test it with personas - the same question asked as three users with different access should retrieve three different sets. The internal enterprise assistant shows the pattern, including a canary document that must never appear for the wrong user.
Fail safe when nothing matches
Most “confident hallucination” in RAG is an empty-retrieval problem. This is the grounding gate in the diagram above - gate on the retrieval score:
docs = retrieve(query, k=5)
if not docs or docs[0].score < MIN_SCORE:
return fallback() # ask to rephrase, or hand off - don't invent
context = format_context(docs)
Log top_score, retrieved ids and an index_lag signal on every request so you
can see retrieval health, not just final answers.
Keep the index fresh
Retrieval quality asks “did we fetch the right context?”; freshness asks “is it still true?” Both need monitoring - a refresh schedule, index-lag and chunk-age signals, deletion propagation, and drift detection. The RAG freshness monitoring checklist covers this field by field.
Minimal vs mature
| Aspect | Minimal | Production-grade |
|---|---|---|
| Retrieval | Top-k vector search | Hybrid search (RRF) + reranker |
| Quality | Judge final answers | Hit rate and recall measured separately |
| Citations | Trusted | Checked against retrieved ids on every answer |
| Chunking | Fixed guess | Chosen by experiment against the retrieval set |
| Permissions | Filtered after retrieval | Filtered inside the query, tested with personas |
| Freshness | Manual re-index | Scheduled + lag/drift alerts |
| No match | Model guesses | Safe fallback on low score |
Where this lives in a real system
See the RAG chatbot reference architecture for where the retriever, reranker and grounding check sit, the RAG / vector tools for the infrastructure, and the RAG items in the Production Checklist for the bar to clear.