retrieval · Level 3

Document pipelines: parsing, cleaning and chunking

Build traceable passages and test what survives each boundary.

By Mickarle Wagstaff-Irons - Micky Irons

  • Level 3Building
  • 120 min
  • 6 chapters
  • Free PDF, no account
The Indexretrieval / 03

Start with the essentials

The short answer

A document pipeline turns source material into traceable retrieval passages. Separate extraction, careful normalisation, chunking and indexing, preserving revisions and useful locations throughout. This workbook builds a small Python text pipeline, tests its boundaries and compares overlap on a labelled fixture. It also explains why OCR, layout, permissions and retrieval quality require separate checks.

What you will learn

  • Separate extraction, normalisation, chunking and indexing.
  • Preserve input and cleaned-text revisions with clear coordinate systems.
  • Implement bounded whitespace windows and repeatable chunk identities.
  • Test malformed input, exact endings and source-span integrity.
  • Compare overlap using explicit evidence strings and repetition counts.
  • Write a handover that identifies unsupported formats and operational limits.

Who it is for

Learners who can run small Python scripts and want to prepare traceable text for a retrieval system.

Before you start

  • Complete the RAG introduction and understand why retrieved evidence matters.
  • Be able to save Python files and run functions and unit tests locally.

Read a sample · Chapter 02 of 06

02

Normalise carefully and keep the original revision

Cleaning is a policy decision. It should remove known representational noise without silently changing the evidence. Preserve source bytes or an authorised immutable reference, and label the coordinate system used by the cleaned text.

Create an empty folder for this exercise and save the following code as pipeline.py. Use Python 3.12 or a compatible version; the recorded execution for this course uses Python 3.12.10 on Windows 11. No additional packages, model accounts or network calls are required. The file is deliberately built in two parts: the function continues in the next chapter. Do not try to run a half-copied function as the final lab.

python · 31 lines
import hashlib
import json
import re
import unicodedata

POLICY = "word-window-v1"


def normalise(raw):
    if not isinstance(raw, bytes):
        raise TypeError("raw must be bytes")
    text = raw.decode("utf-8").removeprefix("\ufeff")
    if "\x00" in text:
        raise ValueError("NUL is outside this text contract")
    text = text.replace("\r\n", "\n").replace("\r", "\n")
    return unicodedata.normalize("NFC", text)


def chunk_document(doc_id, raw, max_words=8, overlap=2):
    if not isinstance(doc_id, str) or not doc_id.strip():
        raise ValueError("doc_id must be a non-empty string")
    if type(max_words) is not int or max_words < 1:
        raise ValueError("max_words must be a positive integer")
    if type(overlap) is not int or not 0 <= overlap < max_words:
        raise ValueError("overlap must be an integer below max_words")
    text = normalise(raw)
    raw_hash = hashlib.sha256(raw).hexdigest()
    text_hash = hashlib.sha256(text.encode("utf-8")).hexdigest()
    spans = list(re.finditer(r"\S+", text))
    chunks = []
    step = max_words - overlap

The input is decoded strictly. Malformed UTF-8 raises an exception rather than substituting a replacement character that might disguise a damaged identifier or amount. The policy removes one leading byte-order mark, converts common newline encodings to LF and applies Unicode NFC. Python documents NFC as canonical normalisation; compatibility normalisation is a different choice. This course does not silently lowercase text, remove punctuation or merge every line.

The raw hash identifies these exact input bytes. The text hash identifies the normalised UTF-8 text. Two files can have the same cleaned text and different raw hashes, for example after a newline change. Keeping both helps explain that distinction. A hash is not a signature, an access-control decision or proof that a source is truthful. The lab uses SHA-256 for repeatable content fingerprints, as provided by Python hashlib.

Offsets in later chunks refer to the normalised Python string. They are not original file byte positions, page coordinates, model tokens or JavaScript string offsets. Normalisation can change string length. Save the cleaned revision with its hash if you need to reproduce a slice later. To highlight the original PDF you need a separate mapping supplied by the extraction stage; counting characters in its text is not a substitute.

Avoid universal cleaning rules. Joining all line breaks could damage a poem, a table or code indentation. Removing every repeated line could erase a repeated safety warning. Deleting punctuation could alter a decimal number or a command. For the workshop policy, inspect the three sentences and retain their boundaries. If you later add a transformation, give the pipeline a new policy version and test representative material before replacing the index.

Try it yourself · Activity 02

10 min

Explain two hashes

Compare b"Desk\r\nClosed" with b"Desk\nClosed" using normalise and hashlib.sha256.

  1. Predict which hash changes.
  2. Confirm the two normalised strings are equal.
  3. Explain why a raw-file offset cannot be inferred from a cleaned offset.

Which revision would you retain for an auditable citation?

Worked answer

The raw hashes differ because the bytes differ. Normalisation produces the same LF string, so the text hashes match. The removed carriage return shifts later byte positions. Retain the source revision and the cleaned revision, plus an explicit mapping if original-format highlighting is required.

Read a sample · Chapter 03 of 06

03

Build bounded windows with reproducible identities

A chunk needs both content and a way back to its source. Finish the function by constructing windows over explicit word positions, recording character spans and including the processing policy in each identifier.

Append this indented continuation directly after step = max_words - overlap in pipeline.py. The regular expression finds runs of non-whitespace characters. We call those runs words for the exercise; they are neither linguistic word segmentation nor a model tokenizer. The step is positive because overlap must be smaller than the window. The final break prevents an unnecessary extra tail once a window has already reached the last word.

python · 21 lines
    for start in range(0, len(spans), step):
        end = min(start + max_words, len(spans))
        lo = spans[start].start()
        hi = spans[end - 1].end()
        identity = [POLICY, doc_id, raw_hash, text_hash,
                    max_words, overlap, start, end]
        encoded = json.dumps(identity, separators=(",", ":"))
        chunk_id = hashlib.sha256(encoded.encode("utf-8")).hexdigest()
        chunks.append({
            "id": chunk_id, "doc_id": doc_id,
            "raw_sha256": raw_hash, "text_sha256": text_hash,
            "policy": POLICY,
            "max_words": max_words, "overlap": overlap,
            "word_start": start, "word_end": end,
            "char_start": lo, "char_end": hi,
            "text": text[lo:hi],
        })
        if end == len(spans):
            break
    return {"text": text, "chunks": chunks,
            "status": "ready" if chunks else "empty"}

Word ranges use a zero-based start and an exclusive end. A range from 6 to 14 contains eight whitespace units. Character ranges use the same half-open convention and preserve the whitespace between the selected first and last units. Leading and trailing whitespace outside those units is not part of any chunk. That is an intentional contract, so a coverage test should require every word, not every whitespace character, to be covered.

The ID includes the document identity, both hashes, policy name, sizes and word range. Repeating the same inputs produces the same records. Changing the document identity, raw revision or window policy produces a different identity even when a visible passage looks similar. This prevents this exercise from collapsing two distinct source records just because their text matches. Your document identifiers must themselves come from a trustworthy, stable upstream scheme.

Version the implementation honestly. The policy string names this algorithm, but it does not automatically detect a code change. If you alter splitting or normalisation, update the policy and record the tested implementation revision. Downstream, use that revision to rebuild and reconcile old chunks. Do not leave both old and new windows indefinitely searchable merely because their IDs differ. Plan an atomic index generation or another tested replacement process.

This implementation materialises the full text, a match list and every result in memory. It has no file-size ceiling or streaming adapter. Use the small supplied fixtures only. Before connecting uploads, add format and resource bounds appropriate to the actual ingestion service. Store source permissions with a real authorisation design outside this example; an ID or hash alone must never confer access to a chunk.

Try it yourself · Activity 03

15 min

Trace the final window

For sixteen words, use max_words=8 and overlap=2.

  1. Write each start and exclusive end.
  2. Count repeated word occurrences.
  3. Predict the output for exactly eight words and explain the final break.

Would a chunk-size limit stated in words fit a model context limit?

Worked answer

The ranges are [0,8), [6,14) and [12,16). They contain 20 occurrences of 16 words, with four repeated occurrences. Exactly eight words produce one window. Without the final break, a smaller redundant tail could appear. Model contexts use their own token counts and also need space for instructions and answers.

Read a sample · Chapter 04 of 06

04

Inspect the passages before measuring retrieval

Run a transparent fixture and look at what each window actually contains. The first useful review is whether a learner can trace a passage and whether the evidence needed for an answer remains together.

Save the following file as demo.py beside pipeline.py. Run python demo.py in that folder. The main guard keeps importing RAW into another script quiet. All text is invented. The blank line and newline are retained within windows, so repr displays them as escapes instead of creating misleading extra output lines.

python · 11 lines
from pipeline import chunk_document

RAW = (b"Workshop loan policy.\n\n"
       b"Borrowers may keep a kit for seven days.\n"
       b"Return kits to the desk.")

if __name__ == "__main__":
    result = chunk_document("workshop-loans-v1", RAW)
    for row in result["chunks"]:
        print(row["word_start"], row["word_end"],
              repr(row["text"]))

text · 3 lines
0 8 'Workshop loan policy.\n\nBorrowers may keep a kit'
6 14 'a kit for seven days.\nReturn kits to'
12 16 'kits to the desk.'

The borrowing duration is complete in the second window. The return instruction is split: one window stops after to and the next starts after Return. Overlap repeats words but does not necessarily preserve every useful statement. A window can also combine the end of one section with the start of another. Examine boundaries rather than treating the existence of overlap as proof of semantic completeness.

For a structured policy collection, a later design could use headings and paragraphs as primary boundaries, splitting an oversized section only when necessary. Keep heading context as explicit metadata or clearly labelled context text rather than silently adding it to the quoted passage. Record any such added text separately so a citation slice still reproduces the source span. Lists, tables and code need their own fixtures; the whitespace splitter is a baseline, not a universal document strategy.

Check provenance by slicing result["text"] with the recorded character bounds and comparing it with each row. Keep an ID-to-source-revision mapping when storing records. If a citation points to a later edited document, the quoted passage may no longer exist at those offsets. Retaining the cleaned revision permits local verification, while retaining an original reference allows a reader to understand where that text came from.

Try it yourself · Activity 04

10 min

Locate a broken answer

A learner asks where to return a kit. Inspect the default three chunks.

  1. Find every occurrence of Return or desk.
  2. Identify whether one chunk contains the full instruction.
  3. Propose a different boundary policy and a fixture to test it.

Could a model plausibly fill the gap with an unsupported location?

Worked answer

The second chunk contains Return kits to; the third contains kits to the desk. No single default chunk contains the whole return sentence. Try a sentence/paragraph-aware boundary or a larger measured window. Test that the full instruction is present, and still check retrieval and answer behaviour separately.

Read a sample · Chapter 05 of 06

05

Test coverage and measure the cost of overlap

Use two kinds of evidence: deterministic checks of the pipeline contract and a small content check of what survives together. Neither one measures a deployed search engine or the quality of a model answer.

Save the next block as compare.py and run python compare.py. Each fixture fact is an exact, manually selected evidence string. A hit means one chunk contains that entire string. This is a deliberately narrow ingestion check. It is not retrieval recall, semantic equivalence or evidence that a model answered correctly. The facts, sizes and results below describe only the sixteen-word policy.

python · 15 lines
from pipeline import chunk_document
from demo import RAW

FACTS = ("Workshop loan policy.",
         "a kit for seven days.",
         "Return kits to the desk.")

for size, overlap in [(8, 0), (8, 2), (10, 4)]:
    rows = chunk_document("workshop-loans-v1", RAW,
                          size, overlap)["chunks"]
    covered = sum(any(fact in r["text"] for r in rows)
                  for fact in FACTS)
    occurrences = sum(r["word_end"] - r["word_start"]
                      for r in rows)
    print(size, overlap, len(rows), covered, occurrences)

Scroll sideways to see every column.

Test coverage and measure the cost of overlap · Table 1
Window / overlapChunksComplete fixture facts / 3Word occurrences
8 / 02216
8 / 23220
10 / 42320

Overlap changes which evidence survives; the middle setting does not improve the total. More repeated words also consume storage and potentially context. Counts here exclude metadata and tokenizer behaviour, so they cannot be converted directly into embedding cost. For a real decision, keep a varied labelled set, retrieve against fixed queries and inspect ranked results and grounded answers separately. Use held-out cases to reduce tuning to a few convenient examples.

Save the next code as test_pipeline.py. Run python -m unittest -v test_pipeline. These tests specify observable requirements: known ranges, exact slices, explicit empty status, invalid input rejection, no redundant tail and repeatable identities. They run locally with the standard library. The course release also checks a larger matrix of window/overlap lengths and malformed inputs; those additional checks do not replace representative document review.

python · 44 lines
import unittest
from pipeline import chunk_document, normalise
from demo import RAW


class PipelineTests(unittest.TestCase):
    def test_windows_and_slice_provenance(self):
        result = chunk_document("fixture", RAW)
        rows = result["chunks"]
        self.assertEqual([(r["word_start"], r["word_end"])
                          for r in rows], [(0, 8), (6, 14), (12, 16)])
        for r in rows:
            self.assertEqual(r["text"], result["text"][
                r["char_start"]:r["char_end"]])

    def test_empty_is_not_ready(self):
        result = chunk_document("empty", b" \n\t")
        self.assertEqual(result["status"], "empty")
        self.assertEqual(result["chunks"], [])

    def test_invalid_utf8_fails(self):
        with self.assertRaises(UnicodeDecodeError):
            normalise(b"\xff")

    def test_invalid_overlap_fails(self):
        for bad in [-1, 8, 9, True, 1.5]:
            with self.subTest(overlap=bad):
                with self.assertRaises(ValueError):
                    chunk_document("fixture", RAW, 8, bad)

    def test_exact_end_has_no_redundant_tail(self):
        rows = chunk_document("small", b"a b c d", 4, 2)["chunks"]
        self.assertEqual(len(rows), 1)

    def test_ids_repeat_and_change_with_source(self):
        first = chunk_document("fixture", RAW)["chunks"]
        again = chunk_document("fixture", RAW)["chunks"]
        other = chunk_document("different", RAW)["chunks"]
        self.assertEqual(first, again)
        self.assertNotEqual(first[0]["id"], other[0]["id"])


if __name__ == "__main__":
    unittest.main()

Challenge a test with a deliberate defect in a disposable copy. Remove the break after end == len(spans), leaving a pass so the Python syntax stays valid. The exact-end test should now fail for four words with overlap two because a redundant tail appears. Restore the correct code and rerun. This demonstrates detection of that particular bug, not that the tests discover every possible pipeline defect.

Try it yourself · Activity 05

15 min

Write a release decision from the fixture

Compare the three settings without calling any model or vector database.

  1. Reproduce the table.
  2. Explain why 8/2 is not an improvement in complete fact count.
  3. Write one reason 10/4 still cannot be declared best for a real collection.

What extra labelled case would most challenge this choice?

Worked answer

The outputs are 8 0 2 2 16, 8 2 3 2 20 and 10 4 2 3 20. The middle setting exchanges one intact fact for another while increasing repetition. The largest setting fits this tiny fixture, but a long table, contradictory revisions or a multi-section document could expose a different failure.

Keep learning

The complete workbook

Build a reproducible text-to-chunks lab with the Python standard library. Preserve source fingerprints and exact cleaned-text spans, reject malformed input and measure a small evidence-coverage fixture. Parsing PDFs, OCR and production indexing are discussed as separate stages, not claimed as implemented features.

  1. 01
    Choose the source contract before splitting text

    A retrieval system can only find evidence that reached its index intact. This lab builds the small, inspectable stage between an extracted text document and a set of traceable passages. Start by deciding what the input represents and which failures must remain visible.

    In the workbook · 1 exercise
  2. 02
    Normalise carefully and keep the original revision

    Cleaning is a policy decision. It should remove known representational noise without silently changing the evidence. Preserve source bytes or an authorised immutable reference, and label the coordinate system used by the cleaned text.

    Read here · 1 exercise
  3. 03
    Build bounded windows with reproducible identities

    A chunk needs both content and a way back to its source. Finish the function by constructing windows over explicit word positions, recording character spans and including the processing policy in each identifier.

    Read here · 1 exercise
  4. 04
    Inspect the passages before measuring retrieval

    Run a transparent fixture and look at what each window actually contains. The first useful review is whether a learner can trace a passage and whether the evidence needed for an answer remains together.

    Read here · 1 exercise
  5. 05
    Test coverage and measure the cost of overlap

    Use two kinds of evidence: deterministic checks of the pipeline contract and a small content check of what survives together. Neither one measures a deployed search engine or the quality of a model answer.

    Read here · 1 exercise
  6. 06
    Publish an inspectable ingestion record

    A useful final deliverable explains which sources were processed, what changed and what remains untested. Treat the index as a derived release with its own replacement and deletion behaviour.

    In the workbook · 1 exercise

Also inside: a 8-point checklist, a glossary of 8 terms and 10 questions and answers to test yourself. 6 hands-on exercises, each with a worked answer at the back where the workbook gives one.

No login, no card, no account. Before the download we ask you to follow Mickai (two quick links). Free to download and use for personal learning, study groups and inside your own team. Please do not resell the workbooks or republish them as your own. Link people to trust-agent.ai instead.

Test yourself

Questions and answers

Does this lab parse PDF files?

No. It accepts UTF-8 text bytes. PDF extraction, layout reconstruction and OCR need separate adapters and checks before this stage.

Why preserve both raw and cleaned hashes?

They describe different byte sequences. A newline change can alter the raw hash while leaving the normalised text hash unchanged.

Do the character offsets point into the original file?

No. They are positions in the normalised Python string. Original bytes, page coordinates and other languages may use different coordinate systems.

Is a whitespace word the same as a model token?

No. This lab counts runs of non-whitespace characters. A model tokenizer and multilingual segmentation can produce very different units.

Why must overlap be smaller than the window?

The next starting position advances by window size minus overlap. A positive advance prevents a non-progressing window policy.

Why stop when the final word is reached?

The last window already covers the remaining words. Stopping avoids smaller redundant tails caused by later overlapping start positions.

Does overlap keep every sentence intact?

No. The default fixture still splits the return instruction across chunks. Inspect the actual boundaries and evaluate representative evidence.

What does the three-fact comparison measure?

It checks whether each exact evidence string appears wholly inside any chunk. It does not measure ranked retrieval or model answer quality.

Does an unchanged chunk ID prove a source is trustworthy?

No. Repeatable IDs support bookkeeping. They do not authenticate a source, enforce permissions or establish factual accuracy.

What happens to old chunks when the source changes?

The lab creates new IDs but has no index lifecycle. A real system must reconcile obsolete records and test replacement, withdrawal and cache behaviour.

When you have finished

Get your certificate of completion

Type your name and download a certificate for this workbook as a PDF, ready to print or to add to LinkedIn. It is made on your own device, so your name is never sent to us. It is a self-declared certificate, not an accredited qualification.

Learn the language

Key terms

Ingestion
The sequence that prepares source material for a retrieval system.
Extraction
Obtaining text and structure from an input format.
OCR
Recognition of text from image pixels, requiring its own verification.
Normalisation
An explicit transformation into a chosen representation.
Chunk
A bounded retrieval unit with content and source metadata.
Overlap
Input units repeated between adjacent windows.

6 of the workbook's 8 terms. The complete glossary is in the workbook.

Follow the evidence

Sources and checks

Facts last checked: .

Examples in this workbook were run on: Python 3.12.10 on Windows 11, standard library only. Six learner tests, 41 targeted checks, 2,574 boundary-matrix cases and a detected redundant-tail mutation; no model or network calls. (2026-09-27).

These workbooks use AI assistance. See how the workbooks are made.

  1. Unicode normalisation in Python 3.12Python Software Foundation
  2. Regular expression match iteration and spansPython Software Foundation
  3. SHA-256 and content digestsPython Software Foundation
  4. PDF text extraction and OCR limitationspypdf documentation

Created by Mickarle Wagstaff-Irons - Micky Irons with the Mickai team. Published by Mickai LTD. Last updated 27 September 2026.

NextKeep going

Where to go next