retrieval · Level 4

Evaluate a RAG system with a grounded test set

Measure the evidence found, the claims supported and the questions a system should leave unanswered.

By Mickarle Wagstaff-Irons - Micky Irons

  • Level 4Frontier
  • 180 min
  • 6 chapters
  • Free PDF, no account
The Indexretrieval / 04

Start with the essentials

The short answer

Evaluate a RAG system at both the retrieval and answer stages. A relevant document in the context does not establish that the answer is supported. This course uses original fictional evidence, labelled claims and recorded A/B outputs to calculate retrieval, support, coverage and abstention results in standard-library Python. The grader checks annotation arithmetic; it does not determine whether arbitrary free text is true.

What you will learn

  • Separate retrieval quality, answer support, completeness and abstention.
  • Define question, evidence and claim labels against a frozen corpus.
  • Calculate fixed-cutoff precision, recall and reciprocal rank.
  • Implement and test an explicit citation-support policy.
  • Explain per-question regressions in paired system outputs.
  • Prepare a reproducible report with uncertainty and integration limits.

Who it is for

RAG builders who can read Python and want a transparent evaluation lab before connecting an automated judge or a live model.

Before you start

  • Understand RAG, document chunks and retrieved context.
  • Read Python dictionaries, tuples, sets and functions.
  • Run local scripts and unit tests.

Read a sample · Chapter 01 of 06

01

Define the question and its evidence

A test set needs an answer policy and trustworthy labels before it needs a dashboard.

A retrieval-augmented answer has at least two opportunities to fail. The retriever can miss useful evidence, and the answer stage can misuse evidence it received. The Ragas paper describes this need to evaluate different parts of a RAG pipeline. This workbook makes those parts inspectable with a small, original Python exercise. It does not install Ragas, reproduce its metrics or claim a benchmark against any deployed system.

Our corpus contains five fictional notes and four questions. Each question has a tuple of relevant document IDs and a map from expected claim IDs to the documents that support them. A recorded run contains ranked retrieved IDs and an ordered tuple of claim/citation pairs. The IDs stand for human-reviewed propositions. They are not generated or extracted by this program. The grader never reads a sentence and decides whether it follows logically from another sentence.

For example, monday represents the proposition that the fictional release window starts on Monday. The launch note supports it. switch and smoke represent two separate rollback actions, each with its own supporting note. The region question is unanswerable from this frozen corpus: there is no evidence naming a backup region. Unanswerable means absent from the selected evidence, rather than impossible to know anywhere in the world.

The SQuAD 2.0 paper motivates testing questions for which a passage contains no answer. Here we use an original question and explicit empty labels, not examples copied from that dataset. Treat abstention as a behaviour to inspect. A system that declines every question could avoid unsupported statements while failing to help with questions that the corpus can answer. Answerable-case coverage exposes that failure.

Create a local folder containing evaluate.py, fixtures.py, demo.py and test_evaluate.py. Python 3.12 is sufficient; no install, account, GPU or private document is needed. Save fixtures.py below. Systems A and B are names for hand-authored output fixtures. They are not model versions, measured retrievers or actual product results. Their differences are deliberately chosen to make the scoring rules visible.

python · 56 lines
"""Original fictional sources, reviewed labels and invented system outputs."""
CORPUS = {
    "launch": "The fictional release window starts on Monday.",
    "rollback": "Switch traffic back to the previous version.",
    "checklist": "After rollback, run the smoke test.",
    "owners": "The fictional Aurora team owns the alert.",
    "archive": "An old planning note discusses a retired project.",
}
CASES = {
    "launch": {
        "question": "When does the release window start?",
        "relevant": ("launch",),
        "support": {"monday": ("launch",)},
    },
    "rollback": {
        "question": "What are the two rollback actions?",
        "relevant": ("rollback", "checklist"),
        "support": {"switch": ("rollback",), "smoke": ("checklist",)},
    },
    "owner": {
        "question": "Who owns the alert?",
        "relevant": ("owners",),
        "support": {"aurora": ("owners",)},
    },
    "region": {
        "question": "Which region hosts the backup?",
        "relevant": (), "support": {},
    },
}
RUNS = {
    "launch": {
        "A": {"retrieved": ("launch", "archive"),
              "claims": (("monday", ("launch",)),)},
        "B": {"retrieved": ("launch", "archive"),
              "claims": (("monday", ("launch",)),
                         ("bonus", ("launch",)))},
    },
    "rollback": {
        "A": {"retrieved": ("rollback",),
              "claims": (("switch", ("rollback",)),)},
        "B": {"retrieved": ("rollback", "checklist"),
              "claims": (("switch", ("rollback",)),
                         ("smoke", ("checklist",)))},
    },
    "owner": {
        "A": {"retrieved": ("archive", "owners"),
              "claims": (("aurora", ("archive",)),)},
        "B": {"retrieved": ("owners", "archive"),
              "claims": (("aurora", ("owners",)),)},
    },
    "region": {
        "A": {"retrieved": ("archive",),
              "claims": (("region", ("archive",)),)},
        "B": {"retrieved": (), "claims": ()},
    },
}

Before adapting this structure, select representative questions and freeze the source-document versions. Keep development examples separate from a held-out evaluation set. Tune on development data, then record the held-out result before investigating failures. If you repeatedly tune to those failures, that set has become development data; obtain fresh evaluation questions. A label author should record uncertainty and disagreements rather than silently treating an incomplete reference as universal truth.

Try it yourself · Activity 01

15 min

Label one answerable and one unanswerable question

Inspect the rollback and region cases in fixtures.py.

  1. Write the exact propositions needed for a complete rollback answer.
  2. Match each proposition to its source note.
  3. Explain why an empty region label differs from a failed retrieval.

Who would resolve a disagreement about whether a source supports a claim?

Worked answer

Rollback needs switch, supported by rollback, and smoke, supported by checklist. Region has no supporting document in this frozen corpus. A retriever returning no documents for rollback is a retrieval miss; no evidence for region is instead part of the question label. In a real collection, a reviewer must check that the supposedly absent evidence was not merely overlooked.

Read a sample · Chapter 02 of 06

02

Measure retrieval at a fixed cutoff

Declare denominators, ordering and missing-value rules before comparing numbers.

Stanford’s information-retrieval text defines precision using retrieved relevant items and recall using all relevant items. For this lab we evaluate the first k ranked documents. Precision at k is relevant hits in those positions divided by k; recall at k is those hits divided by the number of labelled relevant documents. A short list keeps the k denominator. This fixed-slot convention must be disclosed because precision over returned items uses a different denominator.

With k=2, rollback A returns only rollback. One of two target slots is a relevant hit, so P@2 is 0.50. It retrieves one of the two relevant documents, so R@2 is also 0.50. Precision over the single returned item would be 1.00, but it is not the metric this script calls P@2. Never compare two dashboards that silently use these different definitions as though their precision values were equivalent.

Reciprocal rank at k is one divided by the first relevant document’s one-based rank within that cutoff. It is zero when an answerable question has no hit. owner A retrieves archive first and owners second, so RR@2 is 0.50 while recall is 1.00. owner B moves owners first, making RR@2 equal to 1.00 without changing recall. This small diagnostic says where the first hit appeared; it does not measure whether the full answer is present.

For a question with no labelled relevant documents, this lab records all three retrieval metrics as None. That is a declared reporting policy, not a mathematical requirement for every retrieval evaluator. The display uses n/a, keeping these cases out of answerable-only aggregates. An empty retrieval on an answerable question instead has precision, recall and reciprocal rank equal to zero. Missing values must not become successful zeros or perfect scores by accident.

Save the first part of evaluate.py below. It rejects duplicate or malformed IDs, unknown retrieved documents, empty corpus text and inconsistent labels. Expected claims need labelled sources from the relevant set. Unknown answer citations remain valid input so the scorer can mark them unsupported. The next two code blocks are consecutive parts of the same file. Append the second after the first, keeping each function definition at the left margin.

python · 13 lines
"""Score a bounded, pre-annotated teaching fixture, not free-text truth."""
import re


def ids(value):
    if type(value) is not tuple or len(value) > 100:
        raise ValueError("Expected at most 100 IDs in a tuple")
    if any(type(x) is not str or not re.fullmatch(
            r"[a-z][a-z0-9-]{0,31}", x) for x in value):
        raise ValueError("Invalid ID")
    if len(value) != len(set(value)):
        raise ValueError("Duplicate ID")
    return set(value)

Continue evaluate.py with the complete validate function below. Keep its definition at the left margin.

python · 34 lines
def validate(corpus, case, run, k):
    if type(k) is not int or not 1 <= k <= 100:
        raise ValueError("Invalid cutoff")
    if type(corpus) is not dict or not corpus:
        raise ValueError("Expected a non-empty corpus")
    known = ids(tuple(corpus))
    if any(type(text) is not str or not text.strip()
           for text in corpus.values()):
        raise ValueError("Empty source text")
    if type(case) is not dict or type(run) is not dict:
        raise ValueError("Expected case and run objects")
    relevant = ids(case.get("relevant"))
    support = case.get("support")
    if type(support) is not dict:
        raise ValueError("Expected claim support labels")
    ids(tuple(support))
    if not relevant <= known or bool(relevant) != bool(support):
        raise ValueError("Inconsistent relevance labels")
    for sources in support.values():
        labelled = ids(sources)
        if not labelled or not labelled <= relevant:
            raise ValueError("Invalid supporting sources")
    retrieved = run.get("retrieved")
    if not ids(retrieved) <= known:
        raise ValueError("Unknown retrieved document")
    claims = run.get("claims")
    if type(claims) is not tuple or len(claims) > 100:
        raise ValueError("Expected a bounded claim tuple")
    if any(type(c) is not tuple or len(c) != 2 for c in claims):
        raise ValueError("Expected claim and citations")
    ids(tuple(claim for claim, _ in claims))
    for _, citations in claims:
        ids(citations)
    return relevant, support, retrieved, claims

These checks establish structural consistency only. A well-formed claim label can still be factually wrong. Relevance labels can omit a useful document, and the corpus may contain contradictory or stale evidence. Keep label review separate from input validation. This tiny binary-relevance lab also omits graded relevance, multiple passages required jointly for one claim, document-length effects and ranking metrics such as nDCG.

Try it yourself · Activity 02

20 min

Calculate the retrieval rows by hand

Work at k=2 before running the demo.

  1. Calculate P@2, R@2 and RR@2 for rollback A and B.
  2. Calculate those metrics for owner A and B.
  3. State what changes if you replace the fixed denominator with the returned list length.

Would a metric that only checks the first relevant hit reveal the missing second rollback action?

Worked answer

Rollback A is 0.50, 0.50, 1.00; B is 1.00, 1.00, 1.00. Owner A is 0.50, 1.00, 0.50; B is 0.50, 1.00, 1.00. Dividing rollback A’s single hit by its returned list length gives 1.00 instead of P@2=0.50. That is a different precision definition and must have a different documented label.

Read a sample · Chapter 03 of 06

03

Score supported claims and answer coverage

A source must support the specific proposition, and that source must be in the visible context.

Append the final part below to evaluate.py at the left margin. It implements a deliberately strict citation policy. A generated claim is supported only when it has at least one citation and every cited document is both labelled as supporting that claim and present in the first k retrieved documents. An existing document ID is not enough. One correct citation does not cancel an incorrect extra citation.

python · 22 lines
def evaluate(corpus, case, run, *, k=2):
    relevant, support, retrieved, claims = validate(
        corpus, case, run, k)
    top = retrieved[:k]
    visible = set(top)
    hits = len(visible & relevant)
    precision = hits / k if relevant else None
    recall = hits / len(relevant) if relevant else None
    rr = (next((1 / rank for rank, doc in enumerate(top, 1)
                if doc in relevant), 0.0) if relevant else None)
    supported = set()
    for claim, citations in claims:
        cited = set(citations)
        permitted = set(support.get(claim, ())) & visible
        if cited and cited <= permitted:
            supported.add(claim)
    return {
        "precision": precision, "recall": recall, "rr": rr,
        "support": len(supported) / len(claims) if claims else None,
        "coverage": len(supported) / len(support) if support else None,
        "abstention_correct": not claims if not support else None,
    }

Support is the number of supported generated claims divided by the number of generated claims. Coverage is the number of supported expected claims divided by the number of expected claims. Duplicate claim IDs are rejected, so repeating a correct claim cannot inflate the numerator. Under this fixture’s binary policy, a claim counts in coverage only after meeting the same citation requirement. These are custom annotation-based measures, not interchangeable implementations of every tool’s faithfulness or completeness score.

launch B includes monday with a valid citation and bonus with a citation to launch. The second claim has no supporting label, so it is unsupported even though the document exists and is relevant to the question. Support is one of two, or 0.50. The one expected claim is covered, so coverage is 1.00. A completeness-only report would miss the unsupported addition; a retrieval-only report would miss it too.

owner A retrieves the correct owners note but cites archive for aurora. Retrieval recall is 1.00 and answer support is zero. Conversely, rollback A cites its one returned action correctly, giving support 1.00 and coverage 0.50. These cases distinguish evidence availability, evidence use and answer completeness. A production review also needs to inspect clarity, relevance, contradictions and usefulness; no single field here represents all of answer quality.

An answer with no claims has undefined support, displayed as n/a. If the question is answerable, coverage is zero. If it is labelled unanswerable, coverage is n/a and abstention_correct is true. Any claim on an unanswerable case makes abstention_correct false. This representation models abstention as an empty claim tuple; it does not recognise a natural-language refusal or verify that its wording is helpful.

A genuine free-text evaluator would first identify propositions, align them to the rubric and inspect the cited passages. A human or model judge can make mistakes in that process. Preserve the original answer, the exact visible context and the review decision so disagreements can be audited. Do not pass arbitrary model text into this toy grader and claim it has checked factuality. Its corpus text is validated for shape but is never used for semantic judgement.

Try it yourself · Activity 03

20 min

Explain a good retrieval with a bad answer

Inspect launch B and owner A at k=2.

  1. Find the unsupported claim or citation in each case.
  2. Calculate support and coverage.
  3. Try rollback B at k=1 and explain why a citation stops counting.

When might a real claim need two documents jointly, rather than either document independently?

Worked answer

Launch B has unsupported bonus: support 0.50, coverage 1.00. Owner A cites the wrong source for aurora: both support and coverage are zero. With k=1, rollback B’s checklist citation is outside the visible context, so only switch counts and both answer measures become 0.50. The source still exists in the corpus, but existence does not prove it was supplied as evidence.

Read a sample · Chapter 04 of 06

04

Compare paired outputs without hiding regressions

Keep every question’s result visible before reaching for an average.

Evidence is only half the story

Fictional launch example: both outputs retrieve the same two documents. A supports its one claim; B supports one of two claims. Recall and coverage stay 1.00, but B's support falls to 0.50.
The same relevant evidence can accompany different answer quality. Claim labels and fractions carry the meaning alongside the gold highlights. Open the full-size evidence diagram.

The question is “When does the release window start?” The fictional launch note says the release window starts on Monday. The only expected claim is monday, supported by that note. Both A and B retrieve launch first and archive second, at k=2. Only launch is labelled relevant. Both therefore have precision at two of 1/2 = 0.50, recall of 1/1 = 1.00, and reciprocal rank of 1/1 = 1.00.

A emits monday citing launch. B emits that same pair and an extra bonus claim, also citing launch. The extra claim has no support label. A supports one of one generated claims; B supports one of two. Both cover the one expected claim, so coverage remains 1.00 while B's support falls to 0.50. An existing or relevant citation is not automatically evidence for every claim.

All four fictional questions: retrieval at cutoff two
Case / outputP@2 R@2RR@2
launch A0.501.001.00
launch B0.501.001.00
rollback A0.500.501.00
rollback B1.001.001.00
owner A0.501.000.50
owner B0.501.001.00
region An/an/an/a
region Bn/an/an/a
All four fictional questions: answer support, coverage and abstention
Case / outputSupport CoverageCorrect abstention
launch A1.001.00n/a
launch B0.501.00n/a
rollback A1.000.50n/a
rollback B1.001.00n/a
owner A0.000.00n/a
owner B1.001.00n/a
region A0.00n/aNo
region Bn/an/aYes

Score contract: P@2 uses two as its denominator even for a short list. Recall divides retrieved relevant documents by all labelled relevant documents. RR@2 uses the first relevant hit's one-based rank. A claim counts as supported only when it has a citation and every cited document is both labelled as supporting that claim and present in the visible top-two context. Coverage divides supported expected claims by expected claims. These are the course's declared annotation rules.

Missing values: n/a is not zero. Region is labelled unanswerable, so its retrieval metrics and coverage are n/a. Region B has no claims, making support n/a and correct abstention Yes. Region A asserts an unsupported claim, making support zero and correct abstention No. Correct abstention is n/a for the three answerable cases. An empty answer to an answerable question would have zero coverage.

B improves the rollback and owner examples and abstains on region, but introduces the launch regression. Do not replace that per-question result with a single favourable average. These are hand-authored output fixtures, not measured model performance. The grader uses pre-annotated claim IDs; it does not extract claims or judge free-text truth. The complete source and expected output remain in the lesson above.

Save demo.py beside the other files and run python demo.py. It prints the four questions for both recorded systems using the same labels and cutoff. This pairing keeps the question set constant. No network request or generated answer happens during the command; you are evaluating the fixed RUNS data.

python · 18 lines
from evaluate import evaluate
from fixtures import CORPUS, CASES, RUNS


def display(value):
    if value is None:
        return "n/a"
    if type(value) is bool:
        return "yes" if value else "no"
    return f"{value:.2f}"


print("case system | P@2 R@2 RR@2 support coverage abstention")
for name, case in CASES.items():
    for system, run in RUNS[name].items():
        score = evaluate(CORPUS, case, run, k=2)
        values = " ".join(display(value) for value in score.values())
        print(f"{name} {system} | {values}")

Expected output from the delivered implementation:

text · 9 lines
case system | P@2 R@2 RR@2 support coverage abstention
launch A | 0.50 1.00 1.00 1.00 1.00 n/a
launch B | 0.50 1.00 1.00 0.50 1.00 n/a
rollback A | 0.50 0.50 1.00 1.00 0.50 n/a
rollback B | 1.00 1.00 1.00 1.00 1.00 n/a
owner A | 0.50 1.00 0.50 0.00 0.00 n/a
owner B | 0.50 1.00 1.00 1.00 1.00 n/a
region A | n/a n/a n/a 0.00 n/a no
region B | n/a n/a n/a n/a n/a yes

B retrieves the second rollback action and corrects the owner citation, while also abstaining on region. However, B adds an unsupported claim on launch. The sensible conclusion is that the fixture contains improvements and a regression. Calling B the winner without an acceptance policy would hide a decision. For a release decision, name the affected question types, failure severity and required follow-up.

For the three answerable questions, mean recall rises from (1 + 0.5 + 1)/3 = 0.8333 to 1.00. Mean support across those same three nonempty answers rises from (1 + 1 + 0)/3 = 0.6667 to (0.5 + 1 + 1)/3 = 0.8333. Both averages improve while launch becomes worse. Always record the sample counts and per-case deltas. These tiny numbers demonstrate arithmetic; four deliberately constructed cases cannot estimate general performance.

Macro averaging gives each question equal weight. Pooling claim counts instead gives longer answers more influence. For A’s answerable cases the pooled support is two supported claims out of three; for B it is four out of five. That differs from B’s macro result. A report must name its aggregation, denominator and treatment of n/a. Keep unanswerable abstention results in a separate slice instead of converting them into retrieval successes.

If you later compare actual systems, hold the corpus, questions, permissions, cutoff and rubric constant, and record the configuration of each run. For stochastic generation, retain repeated outcomes and uncertainty rather than one convenient answer. A change to document splitting may require reviewing relevance labels and source identifiers; if the benchmark itself changes, document that before comparing with older scores.

Try it yourself · Activity 04

15 min

Write a balanced comparison note

Summarise the paired fixture for someone deciding whether to continue an experiment.

  1. Name two improvements and the launch regression.
  2. Report one average together with its denominator.
  3. Specify the case you would investigate before accepting B.

Which failures would your application treat as release blockers even if an average improved?

Worked answer

B improves rollback coverage and the owner answer, and it abstains on region. Launch support drops from 1.00 to 0.50 despite unchanged retrieval. Mean answerable recall is 0.8333 for A and 1.00 for B across three questions. Investigate the unsupported bonus claim and preserve that case as a regression test. This is a fixture comparison, not evidence that a real model or retrieval configuration improved.

Read a sample · Chapter 05 of 06

05

Test the evaluator before trusting its report

A scoring defect can make a weak system look strong.

Save test_evaluate.py and run python -m unittest -v test_evaluate. Python’s unittest runner discovers the TestCase methods and reports failed assertions separately from errors. The ten tests check cutoff denominators, unsupported extra claims, correct answers, wrong citations, context visibility, empty answers, duplicate identities and invalid labels. They test the arithmetic contract that the course has explicitly defined.

python · 75 lines
from copy import deepcopy
import unittest
from evaluate import evaluate
from fixtures import CORPUS, CASES, RUNS


class EvaluationTests(unittest.TestCase):
    def score(self, name, system, **kwargs):
        return evaluate(CORPUS, CASES[name], RUNS[name][system], **kwargs)

    def test_short_list_keeps_cutoff_denominator(self):
        score = self.score("rollback", "A")
        self.assertEqual((score["precision"], score["recall"]), (.5, .5))

    def test_good_retrieval_does_not_hide_extra_claim(self):
        score = self.score("launch", "B")
        self.assertEqual(score["recall"], 1)
        self.assertEqual(score["support"], .5)

    def test_complete_supported_answer(self):
        score = self.score("rollback", "B")
        self.assertEqual(score["support"], 1)
        self.assertEqual(score["coverage"], 1)

    def test_existing_but_wrong_citation_is_not_support(self):
        score = self.score("owner", "A")
        self.assertEqual((score["recall"], score["rr"]), (1, .5))
        self.assertEqual((score["support"], score["coverage"]), (0, 0))

    def test_citation_outside_visible_context_is_unsupported(self):
        score = self.score("rollback", "B", k=1)
        self.assertEqual((score["support"], score["coverage"]), (.5, .5))

    def test_missing_or_unknown_citations_do_not_count(self):
        for citations in ((), ("unknown",), ("owners", "archive")):
            with self.subTest(citations=citations):
                run = {"retrieved": ("owners", "archive"),
                       "claims": (("aurora", citations),)}
                score = evaluate(CORPUS, CASES["owner"], run)
                self.assertEqual(score["support"], 0)

    def test_abstention_is_separate_from_missing_metrics(self):
        score = self.score("region", "B")
        self.assertIsNone(score["recall"])
        self.assertIsNone(score["support"])
        self.assertTrue(score["abstention_correct"])
        self.assertFalse(self.score("region", "A")["abstention_correct"])

    def test_answerable_abstention_has_zero_coverage(self):
        score = evaluate(CORPUS, CASES["owner"],
                         {"retrieved": (), "claims": ()})
        self.assertEqual(score["coverage"], 0)
        self.assertIsNone(score["support"])
        self.assertIsNone(score["abstention_correct"])

    def test_duplicate_documents_and_claims_are_rejected(self):
        runs = [{"retrieved": ("owners", "owners"), "claims": ()},
                {"retrieved": ("owners",),
                 "claims": (("aurora", ("owners",)),)*2}]
        for run in runs:
            with self.assertRaises(ValueError):
                evaluate(CORPUS, CASES["owner"], run)

    def test_invalid_labels_and_cutoffs_are_rejected(self):
        case = deepcopy(CASES["owner"])
        case["support"]["aurora"] = ("archive",)
        with self.assertRaises(ValueError):
            evaluate(CORPUS, case, RUNS["owner"]["B"])
        for k in (0, True, 1.5, 101):
            with self.assertRaises(ValueError):
                self.score("owner", "B", k=k)


if __name__ == "__main__":
    unittest.main()

A useful test should be able to expose a relevant defect. In a separate copy, replace hits / k with hits / max(1, len(top)). The short-list test should fail. Restore the source, then try accepting any supporting citation instead of requiring all citations to be permitted. The mixed valid-and-invalid citation case should fail. Restore again, then remove the visible-context intersection; the k=1 citation test should fail.

Record the actual assertion failure and restore the passing baseline after each experiment. An import error is not evidence that the intended scoring rule was tested. The delivered verification ran these three isolated mutations and observed failed assertions. It also compared the original scorer with separate arithmetic over 384 small retrieval configurations and ten answer patterns, making 3,840 evaluations. This broadens local arithmetic coverage without turning the synthetic fixture into a real-world RAG benchmark.

Tests cannot repair a wrong annotation policy. If reviewers disagree about whether a passage supports a proposition, capture both judgements and resolve or retain the ambiguity. Consider paraphrases, partial support, conflicting versions, multi-step inferences and claims with several clauses. Do not force every complex answer into one convenient claim ID. Document the unit of evaluation and review a sample independently when a human review process is available.

Try it yourself · Activity 05

15 min

Make one score test fail for the right reason

Use a copy of the lab folder for a deliberate defect.

  1. Run the original ten tests and keep the passing output.
  2. Change the short-list denominator and identify the failed assertion.
  3. Restore evaluate.py and rerun the suite.

What kind of incorrect source label could pass every arithmetic test in this workbook?

Worked answer

The incorrect returned-length denominator gives rollback A precision 1.00 instead of 0.50. The short-list test fails on that difference. After restoring hits / k, all ten methods pass again. If Python cannot import the file, repair the syntax before interpreting the experiment; a syntax failure does not validate the precision test.

Read a sample · Chapter 06 of 06

06

Prepare an evaluation record for a real pipeline

Carry the transparent contract forward, then test the parts that this fixture does not implement.

Make a run manifest containing a run ID, date, corpus and question-set versions, code revision, retrieval settings, context cutoff, prompt revision, model identifier where applicable, generation settings and rubric version. Store the actual ordered context and answer, not only their scores. Distinguish source-document IDs from chunk IDs and record which source version each chunk represents. Preserve a route back to the exact evidence used for the decision.

Begin with one controlled pipeline change, such as chunk size or reranking. Capture both configurations against the same frozen questions, and evaluate retrieval and answer stages separately. If retrieval improves but answer support falls, inspect the changed context and generated claims before changing several more settings. A top-k retrieval record must match the context actually delivered to the answer stage; otherwise the visible-context rule evaluates the wrong input.

A practical test collection should contain straightforward lookups, questions requiring several pieces of evidence, near-duplicate notes, outdated information, conflicting passages, unanswerable questions and permission-sensitive cases. The permissions slice checks which evidence may be used for the caller; answer accuracy cannot excuse disclosure of a restricted document. Use synthetic or appropriately authorised material, with redacted diagnostics and defined retention, before considering private production traces.

Measure latency, cost and failures alongside answer quality if those matter to the application. Record the measurement environment and request count, including retries and failures. This lab makes no latency, token-cost or GPU requirement claim because it calls no real model or retriever. A hosted judge adds its own cost, variability and error; validate its decisions against a reviewed sample rather than treating its score as unquestionable ground truth.

Define acceptance before looking at a favourite configuration: required evidence coverage on critical questions, tolerable unsupported-claim failures, expected abstention behaviour and a plan for unresolved labels. Keep the full per-question table, slice counts, regression examples and exceptions. For larger samples, report appropriate uncertainty and avoid treating repeated runs of the same few questions as independent new questions. Monitor later changes to corpus, prompts and providers with versioned evaluations rather than assuming a passing run lasts forever.

Your deliverable is an executable local scoring contract plus a concrete integration plan. The four files can be reproduced without an account, and the fixture’s mixed results can be explained by hand. Real retrieval, claim extraction, semantic judging, human review and production monitoring are separate work. Preserve those distinctions in the report so someone else can tell which conclusions were measured and which still need evidence.

Try it yourself · Activity 06

15 min

Draft the next experiment manifest

Prepare a short record for evaluating one future retrieval change.

  1. Choose one changed setting and list the settings held constant.
  2. Name the stored evidence, answer and annotation versions.
  3. Define a regression rule and an unanswerable-question slice.

Could another person reproduce the result from your saved record without asking what changed?

Worked answer

A usable record might compare two chunk-size settings with one frozen corpus, question set, prompt and model configuration. Retain ranked chunk IDs, exact supplied text, answers, annotations and errors for both runs. Report per-question changes and a separate unanswerable slice. Investigate any new unsupported claim on a designated critical question, even if average recall rises. State that the experiment remains planned until actual pipeline runs and review exist.

Keep learning

The complete workbook

Create an evidence-labelled test set, implement precision at a fixed cutoff, recall and reciprocal rank, then score supported claims and answer coverage. Run paired fixtures and ten tests, expose scoring defects, and design a versioned evaluation record for a real retrieval-augmented application.

  1. 01
    Define the question and its evidence

    A test set needs an answer policy and trustworthy labels before it needs a dashboard.

    Read here · 1 exercise
  2. 02
    Measure retrieval at a fixed cutoff

    Declare denominators, ordering and missing-value rules before comparing numbers.

    Read here · 1 exercise
  3. 03
    Score supported claims and answer coverage

    A source must support the specific proposition, and that source must be in the visible context.

    Read here · 1 exercise
  4. 04
    Compare paired outputs without hiding regressions

    Keep every question’s result visible before reaching for an average.

    Read here · 1 exercise
  5. 05
    Test the evaluator before trusting its report

    A scoring defect can make a weak system look strong.

    Read here · 1 exercise
  6. 06
    Prepare an evaluation record for a real pipeline

    Carry the transparent contract forward, then test the parts that this fixture does not implement.

    Read here · 1 exercise

Also inside: a 8-point checklist, a glossary of 8 terms and 10 questions and answers to test yourself. 6 hands-on exercises, each with a worked answer at the back where the workbook gives one.

No login, no card, no account. Before the download we ask you to follow Mickai (two quick links). Free to download and use for personal learning, study groups and inside your own team. Please do not resell the workbooks or republish them as your own. Link people to trust-agent.ai instead.

Test yourself

Questions and answers

Does this grader judge arbitrary answer text?

No. It calculates scores from pre-annotated claim IDs and supporting document IDs. It does not extract claims, read passages semantically or decide whether free text is true.

What does unanswerable mean in the fixture?

The frozen corpus has no labelled evidence answering that question. This is relative to that corpus and label review, not a claim that the answer cannot exist elsewhere.

Why is rollback A precision 0.50?

There is one relevant hit and k is two. This lab keeps the fixed cutoff denominator even when only one document is returned. Precision over returned items would be a different metric.

Can recall be perfect while reciprocal rank is lower?

Yes. Owner A retrieves its one relevant document in second position, giving recall 1.00 and reciprocal rank 0.50 at k=2.

Does one good citation cancel a bad extra citation?

No. Under this lab’s strict policy, every citation for a claim must be labelled as supporting it and be present in the visible top-k context.

Why can support be perfect while coverage is incomplete?

All generated claims may be supported even though an expected claim is missing. Rollback A supports its one action but covers only one of two expected actions.

Is an empty answer assigned perfect support?

No. Support is n/a because there are no generated claims. Coverage is zero on an answerable case. Correct abstention is recorded separately for unanswerable cases.

Why not choose B from its improved average recall?

B also introduces an unsupported launch claim. Per-question and severity-based review is needed; a higher retrieval average does not eliminate an answer regression.

Does a syntax error validate a mutation test?

No. A meaningful mutation should execute and trigger an assertion about the intended defect. Record that failed assertion, then restore and rerun the passing source.

What remains before this becomes a live RAG evaluation?

Capture real ranked context and answers, establish claim extraction and semantic review, version the corpus and rubric, test representative slices and report uncertainty, cost and failures. The workbook executes no live pipeline.

When you have finished

Get your certificate of completion

Type your name and download a certificate for this workbook as a PDF, ready to print or to add to LinkedIn. It is made on your own device, so your name is never sent to us. It is a self-declared certificate, not an accredited qualification.

Learn the language

Key terms

Grounded answer
An answer whose claims are supported by the evidence supplied for the question.
Relevance label
A recorded judgement that a document is useful evidence for a particular question.
Precision at k
Here, relevant hits in the first k positions divided by k, including unused slots.
Recall at k
Relevant hits in the first k positions divided by all labelled relevant documents.
Reciprocal rank
One divided by the first relevant hit’s rank; zero when an answerable case has no hit.
Claim support
Here, the fraction of generated claims meeting the declared citation policy.

6 of the workbook's 8 terms. The complete glossary is in the workbook.

Follow the evidence

Sources and checks

Facts last checked: .

Examples in this workbook were run on: Python 3.12.10 on Windows 11. Four local standard-library files, original fictional documents and manually specified outputs. No retriever, language model, API, network or external judge is called. (2026-09-27).

These workbooks use AI assistance. See how the workbooks are made.

  1. Evaluation of unranked retrieval setsStanford Introduction to Information Retrieval
  2. Evaluation of ranked retrieval resultsStanford Introduction to Information Retrieval
  3. Ragas: Automated Evaluation of Retrieval Augmented GenerationEs et al., arXiv abstract
  4. Know What You Don't Know: Unanswerable Questions for SQuADRajpurkar et al., arXiv abstract
  5. Python 3.12 unittestPython Software Foundation

Created by Mickarle Wagstaff-Irons - Micky Irons with the Mickai team. Published by Mickai LTD. Last updated 27 September 2026.

NextKeep going

Where to go next