retrieval · Level 3
Document pipelines: parsing, cleaning and chunking
Build traceable passages and test what survives each boundary.

Start with the essentials
The short answer
A document pipeline turns source material into traceable retrieval passages. Separate extraction, careful normalisation, chunking and indexing, preserving revisions and useful locations throughout. This workbook builds a small Python text pipeline, tests its boundaries and compares overlap on a labelled fixture. It also explains why OCR, layout, permissions and retrieval quality require separate checks.
What you will learn
- Separate extraction, normalisation, chunking and indexing.
- Preserve input and cleaned-text revisions with clear coordinate systems.
- Implement bounded whitespace windows and repeatable chunk identities.
- Test malformed input, exact endings and source-span integrity.
- Compare overlap using explicit evidence strings and repetition counts.
- Write a handover that identifies unsupported formats and operational limits.
Who it is for
Learners who can run small Python scripts and want to prepare traceable text for a retrieval system.
Before you start
- Complete the RAG introduction and understand why retrieved evidence matters.
- Be able to save Python files and run functions and unit tests locally.
Read a sample · Chapter 02 of 06
Normalise carefully and keep the original revision
Cleaning is a policy decision. It should remove known representational noise without silently changing the evidence. Preserve source bytes or an authorised immutable reference, and label the coordinate system used by the cleaned text.
Create an empty folder for this exercise and save the following code as pipeline.py. Use Python 3.12 or a compatible version; the recorded execution for this course uses Python 3.12.10 on Windows 11. No additional packages, model accounts or network calls are required. The file is deliberately built in two parts: the function continues in the next chapter. Do not try to run a half-copied function as the final lab.
import hashlib
import json
import re
import unicodedata
POLICY = "word-window-v1"
def normalise(raw):
if not isinstance(raw, bytes):
raise TypeError("raw must be bytes")
text = raw.decode("utf-8").removeprefix("\ufeff")
if "\x00" in text:
raise ValueError("NUL is outside this text contract")
text = text.replace("\r\n", "\n").replace("\r", "\n")
return unicodedata.normalize("NFC", text)
def chunk_document(doc_id, raw, max_words=8, overlap=2):
if not isinstance(doc_id, str) or not doc_id.strip():
raise ValueError("doc_id must be a non-empty string")
if type(max_words) is not int or max_words < 1:
raise ValueError("max_words must be a positive integer")
if type(overlap) is not int or not 0 <= overlap < max_words:
raise ValueError("overlap must be an integer below max_words")
text = normalise(raw)
raw_hash = hashlib.sha256(raw).hexdigest()
text_hash = hashlib.sha256(text.encode("utf-8")).hexdigest()
spans = list(re.finditer(r"\S+", text))
chunks = []
step = max_words - overlapThe input is decoded strictly. Malformed UTF-8 raises an exception rather than substituting a replacement character that might disguise a damaged identifier or amount. The policy removes one leading byte-order mark, converts common newline encodings to LF and applies Unicode NFC. Python documents NFC as canonical normalisation; compatibility normalisation is a different choice. This course does not silently lowercase text, remove punctuation or merge every line.
The raw hash identifies these exact input bytes. The text hash identifies the normalised UTF-8 text. Two files can have the same cleaned text and different raw hashes, for example after a newline change. Keeping both helps explain that distinction. A hash is not a signature, an access-control decision or proof that a source is truthful. The lab uses SHA-256 for repeatable content fingerprints, as provided by Python hashlib.
Offsets in later chunks refer to the normalised Python string. They are not original file byte positions, page coordinates, model tokens or JavaScript string offsets. Normalisation can change string length. Save the cleaned revision with its hash if you need to reproduce a slice later. To highlight the original PDF you need a separate mapping supplied by the extraction stage; counting characters in its text is not a substitute.
Avoid universal cleaning rules. Joining all line breaks could damage a poem, a table or code indentation. Removing every repeated line could erase a repeated safety warning. Deleting punctuation could alter a decimal number or a command. For the workshop policy, inspect the three sentences and retain their boundaries. If you later add a transformation, give the pipeline a new policy version and test representative material before replacing the index.
Try it yourself · Activity 02
10 minExplain two hashes
Compare b"Desk\r\nClosed" with b"Desk\nClosed" using normalise and hashlib.sha256.
- Predict which hash changes.
- Confirm the two normalised strings are equal.
- Explain why a raw-file offset cannot be inferred from a cleaned offset.
Which revision would you retain for an auditable citation?
Worked answer
The raw hashes differ because the bytes differ. Normalisation produces the same LF string, so the text hashes match. The removed carriage return shifts later byte positions. Retain the source revision and the cleaned revision, plus an explicit mapping if original-format highlighting is required.
Read a sample · Chapter 03 of 06
Build bounded windows with reproducible identities
A chunk needs both content and a way back to its source. Finish the function by constructing windows over explicit word positions, recording character spans and including the processing policy in each identifier.
Append this indented continuation directly after step = max_words - overlap in pipeline.py. The regular expression finds runs of non-whitespace characters. We call those runs words for the exercise; they are neither linguistic word segmentation nor a model tokenizer. The step is positive because overlap must be smaller than the window. The final break prevents an unnecessary extra tail once a window has already reached the last word.
for start in range(0, len(spans), step):
end = min(start + max_words, len(spans))
lo = spans[start].start()
hi = spans[end - 1].end()
identity = [POLICY, doc_id, raw_hash, text_hash,
max_words, overlap, start, end]
encoded = json.dumps(identity, separators=(",", ":"))
chunk_id = hashlib.sha256(encoded.encode("utf-8")).hexdigest()
chunks.append({
"id": chunk_id, "doc_id": doc_id,
"raw_sha256": raw_hash, "text_sha256": text_hash,
"policy": POLICY,
"max_words": max_words, "overlap": overlap,
"word_start": start, "word_end": end,
"char_start": lo, "char_end": hi,
"text": text[lo:hi],
})
if end == len(spans):
break
return {"text": text, "chunks": chunks,
"status": "ready" if chunks else "empty"}Word ranges use a zero-based start and an exclusive end. A range from 6 to 14 contains eight whitespace units. Character ranges use the same half-open convention and preserve the whitespace between the selected first and last units. Leading and trailing whitespace outside those units is not part of any chunk. That is an intentional contract, so a coverage test should require every word, not every whitespace character, to be covered.
The ID includes the document identity, both hashes, policy name, sizes and word range. Repeating the same inputs produces the same records. Changing the document identity, raw revision or window policy produces a different identity even when a visible passage looks similar. This prevents this exercise from collapsing two distinct source records just because their text matches. Your document identifiers must themselves come from a trustworthy, stable upstream scheme.
Version the implementation honestly. The policy string names this algorithm, but it does not automatically detect a code change. If you alter splitting or normalisation, update the policy and record the tested implementation revision. Downstream, use that revision to rebuild and reconcile old chunks. Do not leave both old and new windows indefinitely searchable merely because their IDs differ. Plan an atomic index generation or another tested replacement process.
This implementation materialises the full text, a match list and every result in memory. It has no file-size ceiling or streaming adapter. Use the small supplied fixtures only. Before connecting uploads, add format and resource bounds appropriate to the actual ingestion service. Store source permissions with a real authorisation design outside this example; an ID or hash alone must never confer access to a chunk.
Try it yourself · Activity 03
15 minTrace the final window
For sixteen words, use max_words=8 and overlap=2.
- Write each start and exclusive end.
- Count repeated word occurrences.
- Predict the output for exactly eight words and explain the final break.
Would a chunk-size limit stated in words fit a model context limit?
Worked answer
The ranges are [0,8), [6,14) and [12,16). They contain 20 occurrences of 16 words, with four repeated occurrences. Exactly eight words produce one window. Without the final break, a smaller redundant tail could appear. Model contexts use their own token counts and also need space for instructions and answers.
Read a sample · Chapter 04 of 06
Inspect the passages before measuring retrieval
Run a transparent fixture and look at what each window actually contains. The first useful review is whether a learner can trace a passage and whether the evidence needed for an answer remains together.
Save the following file as demo.py beside pipeline.py. Run python demo.py in that folder. The main guard keeps importing RAW into another script quiet. All text is invented. The blank line and newline are retained within windows, so repr displays them as escapes instead of creating misleading extra output lines.
from pipeline import chunk_document
RAW = (b"Workshop loan policy.\n\n"
b"Borrowers may keep a kit for seven days.\n"
b"Return kits to the desk.")
if __name__ == "__main__":
result = chunk_document("workshop-loans-v1", RAW)
for row in result["chunks"]:
print(row["word_start"], row["word_end"],
repr(row["text"]))0 8 'Workshop loan policy.\n\nBorrowers may keep a kit'
6 14 'a kit for seven days.\nReturn kits to'
12 16 'kits to the desk.'The borrowing duration is complete in the second window. The return instruction is split: one window stops after to and the next starts after Return. Overlap repeats words but does not necessarily preserve every useful statement. A window can also combine the end of one section with the start of another. Examine boundaries rather than treating the existence of overlap as proof of semantic completeness.
For a structured policy collection, a later design could use headings and paragraphs as primary boundaries, splitting an oversized section only when necessary. Keep heading context as explicit metadata or clearly labelled context text rather than silently adding it to the quoted passage. Record any such added text separately so a citation slice still reproduces the source span. Lists, tables and code need their own fixtures; the whitespace splitter is a baseline, not a universal document strategy.
Check provenance by slicing result["text"] with the recorded character bounds and comparing it with each row. Keep an ID-to-source-revision mapping when storing records. If a citation points to a later edited document, the quoted passage may no longer exist at those offsets. Retaining the cleaned revision permits local verification, while retaining an original reference allows a reader to understand where that text came from.
Try it yourself · Activity 04
10 minLocate a broken answer
A learner asks where to return a kit. Inspect the default three chunks.
- Find every occurrence of Return or desk.
- Identify whether one chunk contains the full instruction.
- Propose a different boundary policy and a fixture to test it.
Could a model plausibly fill the gap with an unsupported location?
Worked answer
The second chunk contains Return kits to; the third contains kits to the desk. No single default chunk contains the whole return sentence. Try a sentence/paragraph-aware boundary or a larger measured window. Test that the full instruction is present, and still check retrieval and answer behaviour separately.
Read a sample · Chapter 05 of 06
Test coverage and measure the cost of overlap
Use two kinds of evidence: deterministic checks of the pipeline contract and a small content check of what survives together. Neither one measures a deployed search engine or the quality of a model answer.
Save the next block as compare.py and run python compare.py. Each fixture fact is an exact, manually selected evidence string. A hit means one chunk contains that entire string. This is a deliberately narrow ingestion check. It is not retrieval recall, semantic equivalence or evidence that a model answered correctly. The facts, sizes and results below describe only the sixteen-word policy.
from pipeline import chunk_document
from demo import RAW
FACTS = ("Workshop loan policy.",
"a kit for seven days.",
"Return kits to the desk.")
for size, overlap in [(8, 0), (8, 2), (10, 4)]:
rows = chunk_document("workshop-loans-v1", RAW,
size, overlap)["chunks"]
covered = sum(any(fact in r["text"] for r in rows)
for fact in FACTS)
occurrences = sum(r["word_end"] - r["word_start"]
for r in rows)
print(size, overlap, len(rows), covered, occurrences)Scroll sideways to see every column.
| Window / overlap | Chunks | Complete fixture facts / 3 | Word occurrences |
|---|---|---|---|
| 8 / 0 | 2 | 2 | 16 |
| 8 / 2 | 3 | 2 | 20 |
| 10 / 4 | 2 | 3 | 20 |
Overlap changes which evidence survives; the middle setting does not improve the total. More repeated words also consume storage and potentially context. Counts here exclude metadata and tokenizer behaviour, so they cannot be converted directly into embedding cost. For a real decision, keep a varied labelled set, retrieve against fixed queries and inspect ranked results and grounded answers separately. Use held-out cases to reduce tuning to a few convenient examples.
Save the next code as test_pipeline.py. Run python -m unittest -v test_pipeline. These tests specify observable requirements: known ranges, exact slices, explicit empty status, invalid input rejection, no redundant tail and repeatable identities. They run locally with the standard library. The course release also checks a larger matrix of window/overlap lengths and malformed inputs; those additional checks do not replace representative document review.
import unittest
from pipeline import chunk_document, normalise
from demo import RAW
class PipelineTests(unittest.TestCase):
def test_windows_and_slice_provenance(self):
result = chunk_document("fixture", RAW)
rows = result["chunks"]
self.assertEqual([(r["word_start"], r["word_end"])
for r in rows], [(0, 8), (6, 14), (12, 16)])
for r in rows:
self.assertEqual(r["text"], result["text"][
r["char_start"]:r["char_end"]])
def test_empty_is_not_ready(self):
result = chunk_document("empty", b" \n\t")
self.assertEqual(result["status"], "empty")
self.assertEqual(result["chunks"], [])
def test_invalid_utf8_fails(self):
with self.assertRaises(UnicodeDecodeError):
normalise(b"\xff")
def test_invalid_overlap_fails(self):
for bad in [-1, 8, 9, True, 1.5]:
with self.subTest(overlap=bad):
with self.assertRaises(ValueError):
chunk_document("fixture", RAW, 8, bad)
def test_exact_end_has_no_redundant_tail(self):
rows = chunk_document("small", b"a b c d", 4, 2)["chunks"]
self.assertEqual(len(rows), 1)
def test_ids_repeat_and_change_with_source(self):
first = chunk_document("fixture", RAW)["chunks"]
again = chunk_document("fixture", RAW)["chunks"]
other = chunk_document("different", RAW)["chunks"]
self.assertEqual(first, again)
self.assertNotEqual(first[0]["id"], other[0]["id"])
if __name__ == "__main__":
unittest.main()Challenge a test with a deliberate defect in a disposable copy. Remove the break after end == len(spans), leaving a pass so the Python syntax stays valid. The exact-end test should now fail for four words with overlap two because a redundant tail appears. Restore the correct code and rerun. This demonstrates detection of that particular bug, not that the tests discover every possible pipeline defect.
Try it yourself · Activity 05
15 minWrite a release decision from the fixture
Compare the three settings without calling any model or vector database.
- Reproduce the table.
- Explain why 8/2 is not an improvement in complete fact count.
- Write one reason 10/4 still cannot be declared best for a real collection.
What extra labelled case would most challenge this choice?
Worked answer
The outputs are 8 0 2 2 16, 8 2 3 2 20 and 10 4 2 3 20. The middle setting exchanges one intact fact for another while increasing repetition. The largest setting fits this tiny fixture, but a long table, contradictory revisions or a multi-section document could expose a different failure.
Keep learning
The complete workbook
Build a reproducible text-to-chunks lab with the Python standard library. Preserve source fingerprints and exact cleaned-text spans, reject malformed input and measure a small evidence-coverage fixture. Parsing PDFs, OCR and production indexing are discussed as separate stages, not claimed as implemented features.
- 01Choose the source contract before splitting textIn the workbook · 1 exercise
A retrieval system can only find evidence that reached its index intact. This lab builds the small, inspectable stage between an extracted text document and a set of traceable passages. Start by deciding what the input represents and which failures must remain visible.
- 02Normalise carefully and keep the original revisionRead here · 1 exercise
Cleaning is a policy decision. It should remove known representational noise without silently changing the evidence. Preserve source bytes or an authorised immutable reference, and label the coordinate system used by the cleaned text.
- 03Build bounded windows with reproducible identitiesRead here · 1 exercise
A chunk needs both content and a way back to its source. Finish the function by constructing windows over explicit word positions, recording character spans and including the processing policy in each identifier.
- 04Inspect the passages before measuring retrievalRead here · 1 exercise
Run a transparent fixture and look at what each window actually contains. The first useful review is whether a learner can trace a passage and whether the evidence needed for an answer remains together.
- 05Test coverage and measure the cost of overlapRead here · 1 exercise
Use two kinds of evidence: deterministic checks of the pipeline contract and a small content check of what survives together. Neither one measures a deployed search engine or the quality of a model answer.
- 06Publish an inspectable ingestion recordIn the workbook · 1 exercise
A useful final deliverable explains which sources were processed, what changed and what remains untested. Treat the index as a derived release with its own replacement and deletion behaviour.
Also inside: a 8-point checklist, a glossary of 8 terms and 10 questions and answers to test yourself. 6 hands-on exercises, each with a worked answer at the back where the workbook gives one.
No login, no card, no account. Before the download we ask you to follow Mickai (two quick links). Free to download and use for personal learning, study groups and inside your own team. Please do not resell the workbooks or republish them as your own. Link people to trust-agent.ai instead.
Test yourself
Questions and answers
Does this lab parse PDF files?
No. It accepts UTF-8 text bytes. PDF extraction, layout reconstruction and OCR need separate adapters and checks before this stage.
Why preserve both raw and cleaned hashes?
They describe different byte sequences. A newline change can alter the raw hash while leaving the normalised text hash unchanged.
Do the character offsets point into the original file?
No. They are positions in the normalised Python string. Original bytes, page coordinates and other languages may use different coordinate systems.
Is a whitespace word the same as a model token?
No. This lab counts runs of non-whitespace characters. A model tokenizer and multilingual segmentation can produce very different units.
Why must overlap be smaller than the window?
The next starting position advances by window size minus overlap. A positive advance prevents a non-progressing window policy.
Why stop when the final word is reached?
The last window already covers the remaining words. Stopping avoids smaller redundant tails caused by later overlapping start positions.
Does overlap keep every sentence intact?
No. The default fixture still splits the return instruction across chunks. Inspect the actual boundaries and evaluate representative evidence.
What does the three-fact comparison measure?
It checks whether each exact evidence string appears wholly inside any chunk. It does not measure ranked retrieval or model answer quality.
Does an unchanged chunk ID prove a source is trustworthy?
No. Repeatable IDs support bookkeeping. They do not authenticate a source, enforce permissions or establish factual accuracy.
What happens to old chunks when the source changes?
The lab creates new IDs but has no index lifecycle. A real system must reconcile obsolete records and test replacement, withdrawal and cache behaviour.
When you have finished
Get your certificate of completion
Type your name and download a certificate for this workbook as a PDF, ready to print or to add to LinkedIn. It is made on your own device, so your name is never sent to us. It is a self-declared certificate, not an accredited qualification.
Learn the language
Key terms
- Ingestion
- The sequence that prepares source material for a retrieval system.
- Extraction
- Obtaining text and structure from an input format.
- OCR
- Recognition of text from image pixels, requiring its own verification.
- Normalisation
- An explicit transformation into a chosen representation.
- Chunk
- A bounded retrieval unit with content and source metadata.
- Overlap
- Input units repeated between adjacent windows.
6 of the workbook's 8 terms. The complete glossary is in the workbook.
Follow the evidence
Sources and checks
Facts last checked: .
Examples in this workbook were run on: Python 3.12.10 on Windows 11, standard library only. Six learner tests, 41 targeted checks, 2,574 boundary-matrix cases and a detected redundant-tail mutation; no model or network calls. (2026-09-27).
These workbooks use AI assistance. See how the workbooks are made.
- Unicode normalisation in Python 3.12Python Software Foundation
- Regular expression match iteration and spansPython Software Foundation
- SHA-256 and content digestsPython Software Foundation
- PDF text extraction and OCR limitationspypdf documentation
Created by Mickarle Wagstaff-Irons - Micky Irons with the Mickai team. Published by Mickai LTD. Last updated 27 September 2026.
NextKeep going
Where to go next
Recommended for you
Embeddings and vector search from first principles
Build vector search from scratch: turn text into vectors, compare them three ways, search by brute force, score it with recall@k and MRR, and see where it fails.
Recommended for you
Design a repeatable evaluation suite for an AI application
Build an offline AI evaluation harness with test cases, a rubric, regression comparisons, uncertainty checks and a reproducible release record.