retrieval · Level 4

Hybrid search and reranking for better retrieval

Combine rankings, inspect candidate limits and measure when another stage helps.

By Mickarle Wagstaff-Irons - Micky Irons

  • Level 4Frontier
  • 180 min
  • 6 chapters
  • Free PDF, no account
The Indexretrieval / 04

Start with the essentials

The short answer

Hybrid search combines complementary retrieval rankings. Rank fusion merges their candidates, while a reranker scores a selected shortlist again. Neither stage can recover an absent document without another retrieval step. This course builds and tests the ranking plumbing with explicit fictional fixtures, then explains how to evaluate real retrievers without confusing a toy result with measured production quality.

What you will learn

  • Distinguish candidate retrieval, rank fusion and reranking.
  • Implement one-based reciprocal rank fusion with reproducible ties.
  • Validate a reranking score map without adding or silently losing candidates.
  • Calculate recall and reciprocal rank at an explicit cutoff.
  • Test ranking boundaries, empty results and deliberate implementation defects.
  • Design a representative evaluation and integration record for real retrievers.

Who it is for

Builders who understand embeddings and can run small Python programs, and want to reason about hybrid retrieval before adding a model or service.

Before you start

  • Understand document embeddings and keyword search.
  • Read Python lists, dictionaries, sets and functions.
  • Run a local Python script and its unit tests.

Read a sample · Chapter 01 of 06

01

Separate the stages and state the limits

A longer retrieval pipeline is useful only when its additional work improves the result you need. Name the role of each stage before comparing scores.

Imagine a document collection containing invoice instructions, error-code notes and camera guides. A query may contain an exact identifier, a broad description or both. The first-stage retrievers return ordered document IDs. A fusion stage combines their lists. A later reranking stage changes the order of a bounded candidate set. The final context selector may shorten that list again before an answer is composed.

The named primary references at the end explain reciprocal rank fusion and the retrieve-then-rerank pattern. A bi-encoder can represent queries and documents separately; a cross-encoder can inspect a query and candidate text together. That describes possible model architecture, not the implementation in this workbook. Here we isolate the ranking plumbing using supplied ordered IDs and supplied scores. We do not download a model, build a keyword index or run semantic inference.

The three queries below are invented teaching fixtures. The keyword and dense lists are hand-authored possible outputs, and the scores are a scripted reranker response. They are neither human study data nor measurements from Mickai infrastructure. The relevant set is the chosen answer key for each toy query. It stays separate from the reranker scores so the program cannot use the evaluation answer key to choose an order.

Save the following block as fixtures.py in a new local folder. You will create ranking.py, demo.py and test_ranking.py alongside it. Use Python 3.12 for the tested path. No package install, account, GPU or paid model call is needed. These files contain only synthetic IDs; do not replace them with confidential queries while experimenting with an external service.

python · 24 lines
"""Invented rankings and scores, not results from a trained model."""
CASES = [
    {
        "query": "read scanned invoice",
        "keyword": ["invoice-layout", "scan-settings", "ocr-check"],
        "dense": ["ocr-check", "invoice-layout", "camera-guide"],
        "relevant": {"ocr-check"},
        "scores": {"ocr-check": .9, "invoice-layout": .4,
                   "scan-settings": .55, "camera-guide": .2},
    },
    {
        "query": "reset error E17",
        "keyword": ["e17-reset", "generic-restart", "e71-reset"],
        "dense": ["generic-restart", "e71-reset", "e17-reset"],
        "relevant": {"e17-reset"},
        "scores": {"generic-restart": .9, "e71-reset": .8,
                   "e17-reset": .3},
    },
    {
        "query": "volcano map",
        "keyword": [], "dense": [], "scores": {},
        "relevant": {"volcano-atlas"},
    },
]

The volcano query is intentionally empty at both retrieval stages even though the answer key names a relevant document. This is a candidate-generation miss. Reordering an empty list cannot fix it. In a real system, distinguish no matching records, an upstream error, an expired permission scope and an intentionally empty collection; those states need different diagnostics even if the user-facing list is empty.

Try it yourself · Activity 01

15 min

Describe the information flow

Write the inputs and outputs of each stage for the invoice query.

  1. List the IDs returned by each retriever and their union.
  2. Identify the supplied values that are evaluation labels rather than ranking inputs.
  3. Explain which component would have to change to retrieve volcano-atlas.

What evidence would you require before describing the dense list as the output of a real semantic retriever?

Worked answer

The invoice union is invoice-layout, scan-settings, ocr-check and camera-guide. Fusion reads the two ordered lists. Reranking reads candidate IDs and their supplied scores. Evaluation reads the separate relevant set. Retrieving volcano-atlas requires changing retrieval, the query, the collection or its eligible scope; changing a reranking score is insufficient.

Read a sample · Chapter 02 of 06

02

Fuse ranks rather than incompatible raw scores

Reciprocal rank fusion gives each appearance a contribution based on its position. Keep the rank offset separate from the number of results you retrieve or return.

For a document at one-based rank r, this lab adds 1 / (offset + r). Its fused score is the sum over the lists containing that document. An absent document contributes nothing for that list. With offset 60, the first position contributes 1/61. The offset controls the difference between positions; it is not the nearest-neighbour count, candidate window, evaluation cutoff or final context length.

This approach does not add a keyword engine score directly to a vector similarity. Rank fusion uses order and discards the magnitude of the original score gaps. That is a deliberate trade-off: a large gap and a tiny gap become indistinguishable if they produce the same ordering. Preserve raw scores separately in a diagnostic record if you need to inspect why a retriever produced its list.

In the invoice fixture, invoice-layout appears at keyword rank 1 and dense rank 2. Its score is 1/61 + 1/62, about 0.03252. ocr-check appears at keyword rank 3 and dense rank 1, giving 1/63 + 1/61, about 0.03227. Fusion therefore places invoice-layout first. Agreement between retrievers is a ranking signal here, not proof that the top document answers the question.

Save this block as the first part of ranking.py. The next two chapters append the rest. IDs are short lowercase local keys. Each source ranking may contain an ID once; the same ID in different lists is expected. Duplicate IDs within one list are rejected instead of counting as extra votes. List sizes and offsets are bounded for this teaching tool, not as a production capacity claim.

python · 32 lines
"""Small local ranking lab. Inputs are ordered IDs, not raw engine scores."""
from collections import defaultdict
from math import isfinite, fsum
import re


def ids(value):
    if type(value) is not list or len(value) > 1000:
        raise ValueError("Expected a list of at most 1000 IDs")
    if any(type(x) is not str or
           re.fullmatch(r"[a-z][a-z0-9-]{0,63}", x) is None
           for x in value):
        raise ValueError("Invalid document ID")
    if len(set(value)) != len(value):
        raise ValueError("Duplicate ID within a ranking")
    return value


def rrf(rankings, *, offset=60, limit=10):
    if type(rankings) is not list or not 1 <= len(rankings) <= 8:
        raise ValueError("Expected 1 to 8 rankings")
    if type(offset) is not int or not 1 <= offset <= 10000:
        raise ValueError("Invalid rank offset")
    if type(limit) is not int or not 1 <= limit <= 1000:
        raise ValueError("Invalid result limit")
    terms = defaultdict(list)
    for ranking in rankings:
        for rank, doc_id in enumerate(ids(ranking), start=1):
            terms[doc_id].append(1 / (offset + rank))
    scored = [(doc_id, fsum(parts))
              for doc_id, parts in terms.items()]
    return sorted(scored, key=lambda row: (-row[1], row[0]))[:limit]

The tie policy is explicit: equal fused scores are ordered by document ID. This makes repeated runs reproducible and independent of which retriever list is supplied first. The function builds its own contribution lists and does not mutate the supplied rankings. It uses fsum for the floating-point sum; printed scores may be rounded for display, but sorting uses the unrounded values.

Eight source lists are permitted by this local contract. They need not be independent. If you submit the same ranking through several duplicated query channels, its candidates receive several contributions and may dominate. Channel selection and weighting are modelling decisions. This implementation deliberately has equal channel weights and rejects malformed data; it does not infer which channels deserve more influence.

Try it yourself · Activity 02

20 min

Calculate and perturb a fusion

Verify the invoice ranking with arithmetic before changing a parameter.

  1. Calculate both two-list scores using one-based ranks.
  2. Repeat with offset 1 and note the larger gap between nearby ranks.
  3. Try one empty list and then a duplicate ID inside one list.

Could two near-duplicate retrieval channels distort your result even though the input passes validation?

Worked answer

At offset 60, invoice-layout is about 0.03252 and ocr-check about 0.03227. At offset 1 they score 1/2 + 1/3 = 0.83333 and 1/4 + 1/2 = 0.75. The order stays the same for this fixture; the change is not evidence of a better setting generally. An empty list contributes nothing. A repeated ID within a list raises ValueError.

Read a sample · Chapter 03 of 06

03

Rerank a bounded candidate set

Reranking revisits candidates already retrieved. Treat its score response as a contract you can validate, and retain a defined tie policy.

Watch the order change

Fictional invoice query. ocr-check moves from keyword rank 3 to fusion rank 2, then rerank position 1; the four candidates stay the same.
The query is “read scanned invoice”. The lab labels only ocr-check relevant. The graphic follows its position using a gold border and a written marker. Open the full-size ranking diagram.

Keyword ranking: invoice-layout → scan-settings → ocr-check. Dense ranking: ocr-check → invoice-layout → camera-guide.

Fuse the two lists with equal weights and one-based ranks: add 1 / (60 + rank) for every list containing the document. Missing documents contribute zero. Keep all four unique candidates, then sort by the invented reranker scores. Those scores are not probabilities. The two score columns use different scales and must not be compared directly.

Invoice query: rows in fused order, scores rounded for display
Document IDRRF score Fixture scoreFinal position
invoice-layout0.0325220.403
ocr-check0.0322660.901
scan-settings0.0161290.552
camera-guide0.0158730.204

For ocr-check, 1/63 + 1/61 = 0.032266 (rounded). For invoice-layout, 1/61 + 1/62 = 0.032522 (rounded), so it ranks first after fusion. Reranking orders the same candidates as ocr-check → scan-settings → invoice-layout → camera-guide.

All three fictional queries: reciprocal rank at cutoff 3 (RR@3)
QueryKeyword FusionRerank
read scanned invoice1/31/21
reset error E1711/21/3
volcano map000

The invoice example improves, the error-code example regresses, and the volcano query retrieves no candidates at any stage. Its relevant volcano-atlas item stays outside the candidate set. A reranker cannot recover it. All rankings, scores and relevance judgements here are invented teaching fixtures, not measured search-engine or model performance.

Append the next block to ranking.py. It accepts a list of candidate IDs and exactly one numeric score per candidate. Higher means earlier in this example. Real model score ranges and interpretation vary; a large raw value is not automatically a probability or confidence estimate. This lab permits finite negative scores and does not impose an invented zero-to-one range.

The keys must match the candidate set exactly. A missing score may indicate truncation or a failed request; an extra ID may indicate a stale or mismatched response. Silently filling missing scores with zero would obscure that distinction. Booleans, strings, infinities and NaN are rejected. An empty shortlist with an empty score map remains a valid empty result.

python · 17 lines
def rerank(candidates, scores, *, limit=10):
    ids(candidates)
    if type(limit) is not int or not 1 <= limit <= 1000:
        raise ValueError("Invalid result limit")
    if type(scores) is not dict or set(scores) != set(candidates):
        raise ValueError("Need exactly one score per candidate")
    for score in scores.values():
        if type(score) not in (int, float):
            raise ValueError("Expected real finite scores")
        try:
            finite = isfinite(score)
        except OverflowError:
            finite = False
        if not finite:
            raise ValueError("Expected real finite scores")
    # Stable sorting preserves candidate order when scores tie.
    return sorted(candidates, key=lambda x: -scores[x])[:limit]

Sorting is stable: tied reranker scores retain the incoming fused order. That differs from fusion, where document ID breaks ties. Both policies are legitimate local choices because they are explicit and tested. The function does not use relevance labels, document titles or the original raw retriever scores. It cannot add a new document, regardless of how plausible another title would look.

The fusion limit determines which candidates reach the reranker. If the invoice shortlist is cut to one item, it contains only invoice-layout. ocr-check is already gone and cannot become the final result. A larger window can admit useful candidates, but requires more text processing or model pairs in a real system. Measure that cost instead of assuming a fixed window works for every collection and query type.

Save demo.py as follows. The score-map comprehension deliberately selects only the IDs in the current candidate window. This keeps a stored fixture reusable across different window sizes without weakening the strict rerank function. In a real adapter, send the selected document texts, keep their IDs with them, and map returned scores back using the documented response contract. Do not zip unrelated result orders by accident.

python · 17 lines
from fixtures import CASES
from ranking import rrf, rerank, metrics

for case in CASES:
    fused = rrf([case["keyword"], case["dense"]], limit=4)
    candidates = [doc_id for doc_id, score in fused]
    # Subset recorded scores after the candidate window is chosen.
    scores = {doc_id: case["scores"][doc_id] for doc_id in candidates}
    final = rerank(candidates, scores, limit=4)
    print(case["query"])
    stages = [("keyword", case["keyword"]),
              ("fusion", candidates), ("rerank", final)]
    for name, ranking in stages:
        result = metrics(ranking, case["relevant"], cutoff=3)
        print(f"  {name}: {','.join(ranking) or '(empty)'}")
        print(f"    R@3={result['recall']:.3f} "
              f"RR@3={result['rr']:.3f}")

Do not run demo.py until the metrics function from the next chapter has been appended. The query text is present for interpretation, but our replay does not compute a score from it. Replacing the fixtures with a cross-encoder or another learned scorer is a separate integration task. Record the model revision, maximum input length and any document truncation so later quality changes can be investigated.

Try it yourself · Activity 03

15 min

Expose the candidate ceiling

Change only the fusion window for the invoice case.

  1. Predict the candidate IDs when the fusion limit is one.
  2. Compare the result with the four-candidate window.
  3. Explain why a higher reranker score cannot rescue an omitted ID.

Would increasing the final displayed result count recover a document excluded before reranking?

Worked answer

The one-candidate window contains invoice-layout, so its final result remains invoice-layout. With four candidates the scripted reranker places ocr-check first. A score is only consulted for an ID in the shortlist; the strict score-map contract also rejects an attempt to add an outside ID. The experiment illustrates a window limit, not a measured improvement from a real model.

Read a sample · Chapter 04 of 06

04

Measure individual queries before an average

A single relevance number can hide a serious regression. Use an explicit answer key, cutoff and missing-judgement policy, then inspect every teaching case.

Append metrics to ranking.py. Recall at a cutoff is the number of judged relevant IDs found in the returned prefix divided by the size of the relevant set. Reciprocal rank at that cutoff is one divided by the position of the first relevant result, or zero if none appears in the prefix. They answer different questions: coverage of known relevant items versus how soon the first useful result appears.

python · 12 lines
def metrics(ranking, relevant, *, cutoff=3):
    ids(ranking)
    if type(relevant) is not set or not relevant:
        raise ValueError("Need at least one relevant judgement")
    ids(list(relevant))
    if type(cutoff) is not int or not 1 <= cutoff <= 1000:
        raise ValueError("Invalid metric cutoff")
    head = ranking[:cutoff]
    hits = sum(doc_id in relevant for doc_id in head)
    reciprocal = next((1 / rank for rank, doc_id in
                       enumerate(head, 1) if doc_id in relevant), 0.0)
    return {"recall": hits / len(relevant), "rr": reciprocal}

This function requires at least one relevant judgement. An unjudged query is not assigned a zero-quality score, because that would mix absent evaluation evidence with a measured miss. A deliberately judged no-answer query needs a separate measure, such as whether the system abstains appropriately. For this exercise, volcano-atlas is explicitly relevant and missing, so the volcano query has valid zero recall and zero reciprocal rank.

Run python demo.py from the folder containing all four files. The output below is from the tested program. At cutoff three every invoice and E17 ranking still contains the one relevant item, so recall stays at one. Reciprocal rank exposes the movement: invoice improves from one-third to one-half to one, while E17 drops from one to one-half to one-third.

text · 21 lines
read scanned invoice
  keyword: invoice-layout,scan-settings,ocr-check
    R@3=1.000 RR@3=0.333
  fusion: invoice-layout,ocr-check,scan-settings,camera-guide
    R@3=1.000 RR@3=0.500
  rerank: ocr-check,scan-settings,invoice-layout,camera-guide
    R@3=1.000 RR@3=1.000
reset error E17
  keyword: e17-reset,generic-restart,e71-reset
    R@3=1.000 RR@3=1.000
  fusion: generic-restart,e17-reset,e71-reset
    R@3=1.000 RR@3=0.500
  rerank: generic-restart,e71-reset,e17-reset
    R@3=1.000 RR@3=0.333
volcano map
  keyword: (empty)
    R@3=0.000 RR@3=0.000
  fusion: (empty)
    R@3=0.000 RR@3=0.000
  rerank: (empty)
    R@3=0.000 RR@3=0.000

These opposing movements are intentional. The E17 query needs an exact identifier; the invented dense list and scorer prefer the generic restart note. The keyword and final-rerank reciprocal-rank averages across the first two queries are both two-thirds, despite substantial per-query changes. Including the missed volcano query lowers both means to four-ninths. Neither equality establishes that the systems are interchangeable for users.

The answer keys are complete only by definition of this small fixture. In a real collection, unjudged results may contain useful evidence that the annotation process missed. Define relevance with examples, record who or what supplied the judgements, and review disagreements. Keep evaluation labels separate from training, query tuning and reranker prompts so a good score does not merely reflect access to the answers.

Try it yourself · Activity 04

20 min

Write an honest result statement

Compare the three stages without claiming an overall quality improvement.

  1. Calculate reciprocal-rank means for the first two queries.
  2. Include the third query and recompute the means.
  3. State a limitation of using one relevant document per query.

Which failure would matter more for your users: a generic answer preceding an exact error code, or a useful paraphrase appearing one place later?

Worked answer

For the first two queries, keyword and rerank means are 2/3; fusion is 1/2. Across all three they are 4/9, 4/9 and 1/3 respectively. Invoice improves while E17 worsens. These are arithmetic results for three invented fixtures, not an empirical search benchmark. With one relevant item, recall is only zero or one at a given cutoff and tells us little about broad evidence coverage.

Read a sample · Chapter 05 of 06

05

Test invariants and deliberately break them

Tests should detect a wrong ranking, not merely confirm that the program returns a list. Small counterexamples make ranking mistakes easier to diagnose.

Save the following as test_ranking.py. Run python -m unittest -v test_ranking. The tested version passes ten test methods. The tests separately check one-based contributions, absent-list behaviour, both tie policies, duplicates, score alignment, invalid numbers, descending reranker order, metric cutoffs, missing judgements and result limits.

python · 57 lines
import unittest
from ranking import rrf, rerank, metrics


class RankingTests(unittest.TestCase):
    def test_one_based_rank_sum(self):
        result = dict(rrf([["a", "b"], ["b"]], offset=10))
        self.assertAlmostEqual(result["a"], 1/11)
        self.assertAlmostEqual(result["b"], 1/12 + 1/11)

    def test_absence_adds_nothing(self):
        self.assertEqual(rrf([["a"], []]), [("a", 1/61)])
        self.assertEqual(rrf([[], []]), [])

    def test_ties_are_reproducible(self):
        self.assertEqual([x[0] for x in rrf([["b"], ["a"]])],
                         ["a", "b"])
        self.assertEqual(rerank(["b", "a"], {"a": 2, "b": 2}),
                         ["b", "a"])

    def test_duplicates_are_not_extra_votes(self):
        with self.assertRaises(ValueError):
            rrf([["a", "a"]])

    def test_reranker_cannot_invent_or_lose_scores(self):
        for scores in ({}, {"a": 1, "b": 2}):
            with self.assertRaises(ValueError):
                rerank(["a"], scores)

    def test_bad_scores_fail(self):
        for bad in (float("nan"), float("inf"), True, "0.8"):
            with self.assertRaises(ValueError):
                rerank(["a"], {"a": bad})

    def test_highest_score_first(self):
        self.assertEqual(rerank(["a", "b"], {"a": -2, "b": 3}),
                         ["b", "a"])

    def test_metric_cutoff_and_denominator(self):
        self.assertEqual(metrics(["x", "a", "b"], {"a", "b"},
                                 cutoff=2), {"recall": .5, "rr": .5})
        self.assertEqual(metrics([], {"a"}),
                         {"recall": 0.0, "rr": 0.0})

    def test_unjudged_query_is_not_zero_quality(self):
        with self.assertRaises(ValueError):
            metrics(["a"], set())

    def test_window_bounds(self):
        self.assertEqual(len(rrf([["a", "b"]], limit=1)), 1)
        for bad in (0, -1, True, 1.5):
            with self.assertRaises(ValueError):
                rrf([["a"]], limit=bad)


if __name__ == "__main__":
    unittest.main()

A useful additional test uses a two-document universe where you can calculate every expected order by hand. The release verification goes further: it enumerates all partial rankings over four IDs, combines every pair at offsets 1, 10 and 60 and limits 1 and 4, and compares the implementation with an independent rational-arithmetic formula. That produces 25,350 configurations. It also swaps channel order, checks malformed inputs and exercises metric cutoffs.

The 76,817 recorded assertions include ranking orders, numeric comparisons, invariance and boundary checks. They demonstrate the stated behaviour on a small finite domain, not correctness for every possible production input. Long lists are bounded here, but the implementation is an in-memory teaching tool. It does not implement network timeouts, request cancellation, byte limits or an external service boundary.

Three isolated mutations were also rejected by genuine failed expectations. Starting ranks at zero changes the first contribution. Sorting reranker scores ascending prefers worse scores under this contract. Ignoring the metric cutoff counts results that should not be evaluated. A test suite that survives all three changes would not be checking the central promises of this lab.

Try it yourself · Activity 05

20 min

Prove that a test can fail

Make one deliberate defect in a copy of the lab, observe the failure, then restore it.

  1. Change start=1 in the fusion loop to start=0.
  2. Run the tests and explain the changed arithmetic.
  3. Restore it, then remove the minus sign from the reranking sort key and repeat.

Which existing test would detect a stale score map from a previous query?

Worked answer

The zero-based version gives the first document 1/60 instead of 1/61 and fails the one-based contribution test. Removing the minus sign sorts low scores first and fails test_highest_score_first. Restore both edits and rerun all ten tests before using the file. These controlled defects belong in a disposable exercise copy, not in a published service.

Read a sample · Chapter 06 of 06

06

Replace the fixtures with evidence, one stage at a time

The next step is a measured adapter to a real retrieval system. Preserve the stage boundaries so an upstream improvement or failure can be isolated.

Start with a small, permission-appropriate corpus and a versioned query set. Connect one keyword retriever and one dense retriever, both returning the same stable document or chunk IDs. Preserve the original text and source reference behind each ID. Two chunks from one document are distinct candidates only if your identity and evaluation policies say so; otherwise duplicated passages can occupy the context budget and distort apparent coverage.

Apply the required eligibility restrictions before content is sent to a scorer or answer model. Recheck access at delivery where the application requires it. Ranking is not an authorisation mechanism: a low score does not protect a restricted record. Do not infer permissions from relevance labels, query wording or a generated instruction inside a retrieved passage. This course does not implement identity, permission filtering or prompt-injection protection.

Keep an execution record for each evaluated query. Include the corpus snapshot, retriever configuration, candidate depths, ordered IDs, fusion offset, reranker revision, input truncation, selected context and observed timings. Record hashes or suitable references rather than indiscriminately logging sensitive source text. The list below is an experiment worksheet, not a requirement to send private data to an analytics vendor.

Scroll sideways to see every column.

Replace the fixtures with evidence, one stage at a time · Table 1
RecordQuestion it answers
Corpus and judgement versionDid the document collection or answer key change?
Each retriever list and depthWas the useful document ever a candidate?
Fusion offset and shortlistWas it removed before scoring?
Reranker input and revisionWas key text truncated or the scoring model changed?
Per-query quality and timingsWhich query groups gained, lost or slowed down?

Compare keyword-only, dense-only, fusion and fusion-plus-reranking on the same held-out queries. Include exact identifiers, paraphrases, ambiguous wording, stale documents and known no-answer cases. Break down results by query type and inspect regressions before choosing a configuration. The supplied toy demo is deliberately too small for an inference about expected quality, statistical significance, cost or user satisfaction.

For latency, measure the real stage boundaries. Two retrievers running concurrently need not cost the sum of their separate durations, yet shared capacity, queueing and network overhead may dominate. Candidate text length and batching affect the reranker. Report observed distributions and the test conditions, not only one warm local request. Do not turn the timing of this rank-only Python script into a claim about an embedding model or GPU.

Specify fallback behaviour before an outage. If one retriever fails, either return a clearly recorded degraded result from the other under an agreed policy, or fail the request; do not quietly report an empty successful list. If the reranker fails, a validated fused shortlist may be an acceptable fallback after the same permission checks. Record the chosen path so offline evaluation can distinguish normal and degraded runs.

Finally, evaluate the composed answer separately from retrieval. A relevant passage can be misquoted or ignored, and a fluent answer can be unsupported. Preserve source references through the final context selection and check whether the answer actually follows from them. Hybrid retrieval and reranking are ways to select material; they do not certify that a generated conclusion is correct.

Try it yourself · Activity 06

15 min

Plan a bounded real experiment

Write a one-page experiment plan before connecting an external retriever or scorer.

  1. Choose a small allowed corpus and three meaningful query categories.
  2. Specify stable IDs, candidate depths, the ranking stages and an evaluation cutoff.
  3. Define a quality check, a latency check and the fallback decision for each failed stage.

Which observation would make you keep a simpler keyword-only path for a subset of queries?

Worked answer

A workable plan compares four retrieval configurations on one frozen corpus and a separate judged query set, inspects exact-ID and paraphrase regressions, records candidate recall and reciprocal rank with a stated cutoff, and measures real request latency under declared load. It defines permitted data, score alignment, empty/error distinctions and whether validated fusion can be used after a reranker failure. The plan does not claim a winner until those measurements exist.

Keep learning

The complete workbook

Write a standard-library Python rank-fusion and reranking lab. Preserve document identities, reject malformed rankings, inspect candidate windows and compare recall and reciprocal rank. The example deliberately includes an improvement, a regression and an unretrieved relevant document.

  1. 01
    Separate the stages and state the limits

    A longer retrieval pipeline is useful only when its additional work improves the result you need. Name the role of each stage before comparing scores.

    Read here · 1 exercise
  2. 02
    Fuse ranks rather than incompatible raw scores

    Reciprocal rank fusion gives each appearance a contribution based on its position. Keep the rank offset separate from the number of results you retrieve or return.

    Read here · 1 exercise
  3. 03
    Rerank a bounded candidate set

    Reranking revisits candidates already retrieved. Treat its score response as a contract you can validate, and retain a defined tie policy.

    Read here · 1 exercise
  4. 04
    Measure individual queries before an average

    A single relevance number can hide a serious regression. Use an explicit answer key, cutoff and missing-judgement policy, then inspect every teaching case.

    Read here · 1 exercise
  5. 05
    Test invariants and deliberately break them

    Tests should detect a wrong ranking, not merely confirm that the program returns a list. Small counterexamples make ranking mistakes easier to diagnose.

    Read here · 1 exercise
  6. 06
    Replace the fixtures with evidence, one stage at a time

    The next step is a measured adapter to a real retrieval system. Preserve the stage boundaries so an upstream improvement or failure can be isolated.

    Read here · 1 exercise

Also inside: a 8-point checklist, a glossary of 8 terms and 10 questions and answers to test yourself. 6 hands-on exercises, each with a worked answer at the back where the workbook gives one.

No login, no card, no account. Before the download we ask you to follow Mickai (two quick links). Free to download and use for personal learning, study groups and inside your own team. Please do not resell the workbooks or republish them as your own. Link people to trust-agent.ai instead.

Test yourself

Questions and answers

Does this lab run a semantic model?

No. It replays explicitly invented retriever rankings and reranker scores. It implements and tests rank fusion, score-map validation and evaluation arithmetic, not embedding generation or model inference.

Why combine ranks instead of raw engine scores?

Different retrieval methods can produce scores on different scales. This lab uses positions to avoid directly adding those values, while deliberately discarding information about the size of score gaps.

Is the fusion offset a candidate count?

No. The offset moderates reciprocal-rank contributions. Retriever depth, fused shortlist length, reranker limit and metric cutoff are separate parameters.

Can one ranking repeat a document ID?

No. This lab rejects duplicates within a source list so one document cannot receive extra votes from the same ranking. The same ID may appear in separate source lists.

Can reranking recover an omitted document?

Not in this contract. Reranking only sorts the supplied candidates. A new retrieval step or a larger earlier candidate window is needed to admit an omitted item.

Why reject NaN or a missing candidate score?

Neither represents a complete usable ordering under the lab contract. Rejecting a malformed score map exposes a bad or mismatched response instead of silently inventing a fallback score.

What happens when reranker scores tie?

Stable sorting preserves the incoming candidate order. Fusion has a separate tie rule: document ID sorts equal fused scores reproducibly.

Does recall alone expose the E17 regression?

At cutoff three it does not: the relevant document remains in all three prefixes. Reciprocal rank changes from one to one-half to one-third, exposing its worsening position.

Why not score an unjudged query as zero?

Missing evaluation labels are different from a measured miss. This metric requires a non-empty relevant set. Judged no-answer queries need a separate abstention measure.

Do passing tests prove production retrieval quality?

No. They check the lab implementation and its finite fixtures. Real quality requires an appropriate corpus, held-out judgements, actual retrieval and reranking runs, error analysis and measured operating conditions.

When you have finished

Get your certificate of completion

Type your name and download a certificate for this workbook as a PDF, ready to print or to add to LinkedIn. It is made on your own device, so your name is never sent to us. It is a self-declared certificate, not an accredited qualification.

Learn the language

Key terms

Candidate
A retrieved item considered by later stages.
Rank fusion
Combining ordered lists into a new ranking.
Rank offset
A constant moderating reciprocal-rank contributions.
Reranker
A stage that scores an existing candidate set again.
Recall at cutoff
Fraction of judged relevant items found in a prefix.
Reciprocal rank
Inverse position of the first relevant result.

6 of the workbook's 8 terms. The complete glossary is in the workbook.

Follow the evidence

Sources and checks

Facts last checked: .

Examples in this workbook were run on: Python 3.12.10 on Windows 11; ten learner tests, 25,350 oracle configurations, 76,817 additional assertions and three caught mutations. The lab makes no network or model calls. (2026-09-27).

These workbooks use AI assistance. See how the workbooks are made.

  1. Hybrid search scoring with reciprocal rank fusionMicrosoft Learn
  2. Retrieve and rerankSentence Transformers documentation
  3. Python 3.12 mathematical functionsPython Software Foundation
  4. Evaluation of ranked retrieval resultsIntroduction to Information Retrieval, Stanford

Created by Mickarle Wagstaff-Irons - Micky Irons with the Mickai team. Published by Mickai LTD. Last updated 27 September 2026.

NextKeep going

Where to go next