retrieval · Level 4
Knowledge graphs and structured retrieval
Follow explicit relationships and keep the evidence for every step.

Start with the essentials
The short answer
A knowledge graph represents entities and their explicit relationships so retrieval can follow a defined pattern instead of relying only on similar wording. A useful answer retains the evidence for each relationship and states the query limits. This workbook builds a small local Python graph, checks its structure and retrieves reproducible paths through fictional project records.
What you will learn
- Separate stable entity identifiers, relation meaning and source records.
- Validate directed relationships against an explicit local schema.
- Implement bounded traversal that terminates on cycles.
- Return a reproducible witness path and source identifiers for each result.
- Test missing edges, input errors, budgets and deliberately broken code.
- Explain how structured and text retrieval can complement one another.
Who it is for
Builders comfortable with Python collections and retrieval pipelines who want inspectable relationship-based answers.
Before you start
- Read Python functions, dictionaries, loops and exceptions.
- Run local Python scripts and basic unit tests.
- Understand how a RAG pipeline retrieves supporting material before composing an answer.
Read a sample · Chapter 01 of 06
Model a question before drawing a graph
Start with the question you need to answer and the evidence you can inspect. A diagram can make relationships visible, but a readable picture does not establish that its edges are true.
Trace the evidence
| Edge | Relationship | Source |
|---|---|---|
| e1 | Atlas uses Parser | s1 |
| e2 | Parser depends on Index | s2 |
| e3 | Index depends on Parser | s2 |
| e4 | Parser maintained by North | s3 |
| e5 | Index maintained by North | s3 |
- s1: Fictional Atlas manifest, revision 1.
- s2: Fictional package dependency note, revision 1.
- s3: Fictional maintenance register, revision 1.
Parser and Index form a directed dependency cycle (e2 and e3). Index also points to North (e5). Archive is isolated in this fixture. No recorded path to Archive does not prove that no real-world relationship exists. These source labels identify fictional records, not external evidence.
Our fictional Atlas project uses a Parser package. Parser depends on an Index package, and Index depends on Parser. Both packages are maintained by the North team. An Archive package has no recorded relationships. These invented records let us practise direction, cycles and missing information without making claims about Mickai systems or a real software supply chain.
The question is deliberately narrow: which entities can we reach from Atlas by following selected outgoing relationships for at most two steps? This is a reachability query, not a dependency-risk score or a judgement about the team. The answer must state the selected relation types, hop limit and graph snapshot. Asking which packages are safe would require a different definition and additional evidence.
Give each entity a stable identifier distinct from its display name. A package renamed in documentation can keep its identifier; two teams with the same display name must not be silently merged. In this lab, lowercase identifiers are local keys. In a shared system, decide the namespace and entity-resolution policy before combining sources. A similar name is a candidate match, not proof that two records describe the same entity.
Scroll sideways to see every column.
| Relation | Allowed direction | Meaning in this fixture |
|---|---|---|
| uses | project to package | The manifest lists a package used by the project. |
| depends_on | package to package | The dependency note names an outgoing dependency. |
| maintained_by | package to team | The register names the package maintainer. |
A source identifier attaches a relationship to a record. Here the source descriptions are explicitly fictional revision labels, not fabricated public URLs. In a real collection, retain a retrievable document version and the exact supporting passage. The graph may record what a source asserted without establishing that the assertion is current or correct.
RDF describes statements using subject, predicate and object. The W3C primer explains that this relationship is directed. Our JSON lab borrows that useful way to think, but is not RDF, JSON-LD or a SPARQL implementation. Its local type checks are application validation rules. Do not treat them as an implementation of RDF inference or assume a file becomes standards-compatible because it contains triples.
Create a new local folder for the four files in this workbook. Save the following fixture as graph.json. It has five nodes, three source descriptions and five uniquely identified edges. The short keys s, p and o mean subject, predicate and object. Keep these exact identifiers for the first run so you can compare your output with the worked result.
{
"nodes": {
"atlas": "project",
"parser": "package",
"index": "package",
"north": "team",
"archive": "package"
},
"sources": {
"s1": "Fictional Atlas manifest, revision 1",
"s2": "Fictional package dependency note, revision 1",
"s3": "Fictional maintenance register, revision 1"
},
"edges": [
{"id": "e1", "s": "atlas", "p": "uses", "o": "parser", "source": "s1"},
{"id": "e2", "s": "parser", "p": "depends_on", "o": "index", "source": "s2"},
{"id": "e3", "s": "index", "p": "depends_on", "o": "parser", "source": "s2"},
{"id": "e4", "s": "parser", "p": "maintained_by", "o": "north", "source": "s3"},
{"id": "e5", "s": "index", "p": "maintained_by", "o": "north", "source": "s3"}
]
}Try it yourself · Activity 01
15 minTrace the fixture by hand
Draw or write a small adjacency table for the five nodes before running any code.
- List the outgoing edges from Atlas, Parser and Index, including their identifiers.
- Mark the dependency cycle and the isolated Archive node.
- Write the expected one-step and two-step answers from Atlas.
Which part of your answer is explicitly recorded, and which part describes a path you computed?
Worked answer
Atlas has e1 to Parser. Parser has e2 to Index and e4 to North. Index has e3 to Parser and e5 to North. One step reaches Parser; two steps reach Parser, Index and North. Archive is absent from these answers. The e2/e3 cycle does not create a new fact that Atlas directly depends on Index.
Read a sample · Chapter 02 of 06
Validate relationships before retrieval
A malformed graph should fail clearly before producing an answer. Validation checks the contract of this small application; it does not verify a source document or solve entity resolution.
Save the next block as the first part of graph_lab.py. The following chapter supplies the rest of the same file. The schema allows projects, packages and teams, plus only the three named relations. Every endpoint must exist and have the right kind. For example, an edge from Atlas directly to North labelled uses is invalid because uses must end at a package.
Each edge has an identifier of its own. Repeating an edge identifier is rejected because a witness path must refer to one unambiguous record. Different identifiers may describe the same relationship with different supporting sources. The lab permits that duplication, then chooses one witness per reachable node. It does not merge evidence, reconcile disagreement or list every possible explanation.
The graph object is capped at 256 nodes, 128 sources and 1,024 edges. Those are teaching limits chosen for a small local exercise, not measured production capacity. The query also has a traversal budget. Neither limit protects the earlier JSON parsing step from an arbitrarily large file: use only the supplied local fixture here. A public ingest service needs byte limits and a separate parsing boundary.
from collections import deque
import re
KINDS = {"project", "package", "team"}
RELATIONS = {
"uses": ("project", "package"),
"depends_on": ("package", "package"),
"maintained_by": ("package", "team"),
}
def valid_id(value):
return (type(value) is str and
re.fullmatch(r"[a-z][a-z0-9-]{0,39}", value) is not None)
def validate(data):
if type(data) is not dict or set(data) != {"nodes", "sources", "edges"}:
raise ValueError("Expected nodes, sources and edges")
nodes, sources, edges = data["nodes"], data["sources"], data["edges"]
if type(nodes) is not dict or not 1 <= len(nodes) <= 256:
raise ValueError("Expected 1 to 256 nodes")
if any(not valid_id(k) or type(v) is not str or v not in KINDS
for k, v in nodes.items()):
raise ValueError("Invalid node identifier or kind")
if type(sources) is not dict or not 1 <= len(sources) <= 128:
raise ValueError("Expected 1 to 128 sources")
if any(not valid_id(k) or type(v) is not str or not v.strip()
for k, v in sources.items()):
raise ValueError("Invalid source identifier or description")
if type(edges) is not list or len(edges) > 1024:
raise ValueError("Expected up to 1024 edges")
seen, adjacency = set(), {node: [] for node in nodes}
for edge in edges:
if type(edge) is not dict or set(edge) != {"id", "s", "p", "o", "source"}:
raise ValueError("Invalid edge fields")
if any(type(v) is not str for v in edge.values()):
raise ValueError("Edge values must be strings")
eid, s, p, o = edge["id"], edge["s"], edge["p"], edge["o"]
if not valid_id(eid) or eid in seen:
raise ValueError("Invalid or repeated edge identifier")
if s not in nodes or o not in nodes or p not in RELATIONS:
raise ValueError("Unknown endpoint or relation")
if (nodes[s], nodes[o]) != RELATIONS[p]:
raise ValueError("Relation endpoints have the wrong kinds")
if edge["source"] not in sources:
raise ValueError("Unknown source")
seen.add(eid)
adjacency[s].append(edge)
for outgoing in adjacency.values():
outgoing.sort(key=lambda edge: edge["id"])
return nodes, adjacencyValidation builds an adjacency list for each node. The list contains only outgoing edges, so a relationship never becomes bidirectional accidentally. Sorting each list by edge identifier makes traversal reproducible when the input edge rows are rearranged. The function sorts its own lists, not the original edge list; it does not rewrite the supplied graph.
The source dictionary requires a non-empty description and a known key for each edge. That is a structural check only. An attacker or a mistaken editor could still invent a convincing description. Before accepting real data, independently inspect the document, version and passage, record extraction confidence and arrange review of ambiguous relationships. Keep an unsupported candidate separate from accepted records.
Standard JSON decoding creates an in-memory object. Its default treatment of duplicate object keys is another reason this exercise is not a hostile-file ingestion boundary. The duplicate-edge test later checks explicit edge IDs in the edge list; it does not certify every property of an arbitrary JSON document.
Try it yourself · Activity 02
15 minReject a plausible-looking bad graph
Use a copy of the fixture to distinguish structural validity from factual support.
- Change e1 object to north and predict the validation result.
- Restore e1, then change its source to an unknown source identifier.
- Restore the fixture and write one false statement that would still satisfy the schema.
What additional record would a reviewer need to accept a real relationship?
Worked answer
The first change fails the project-to-package endpoint rule. The second fails the known-source rule. A uses edge from Atlas to Archive with a valid source identifier would satisfy the structural schema even if the named document never asserted it. Structural validation therefore cannot replace source review.
Read a sample · Chapter 03 of 06
Follow bounded paths and handle cycles
A traversal needs a starting entity, allowed relation types and a stopping rule. Return evidence with the result instead of asking a language model to reconstruct the route afterwards.
Append this block directly after the first part of graph_lab.py. The queue holds a node, the edge IDs used to reach it and their source IDs in the same order. Python deque provides the queue operations used here. Taking from the front and adding at the back visits shallower paths before deeper ones.
def traverse(
data, start, relations, *, max_hops=2, max_edges=100
):
nodes, adjacency = validate(data)
if type(start) is not str or start not in nodes:
raise ValueError("Unknown start node")
if (type(relations) not in (set, frozenset) or not relations or
not relations <= RELATIONS.keys()):
raise ValueError("Choose a non-empty set of known relations")
if type(max_hops) is not int or not 0 <= max_hops <= 5:
raise ValueError("max_hops must be an integer from 0 to 5")
if type(max_edges) is not int or not 1 <= max_edges <= 1024:
raise ValueError("max_edges must be an integer from 1 to 1024")
queue = deque([(start, [], [])])
seen, results, examined = {start}, [], 0
while queue:
node, path, sources = queue.popleft()
if len(path) == max_hops:
continue
for edge in adjacency[node]:
if edge["p"] not in relations:
continue
examined += 1
if examined > max_edges:
raise ValueError("Traversal edge budget exceeded")
target = edge["o"]
if target in seen:
continue
seen.add(target)
next_path = path + [edge["id"]]
next_sources = sources + [edge["source"]]
results.append({
"node": target, "edges": next_path,
"sources": next_sources,
})
queue.append((target, next_path, next_sources))
return resultsThe seen set starts with the starting node and is updated when a target is queued. Parser and Index can point back to one another, but neither is returned twice. The start node is not included in its own results. The hop limit counts edges in a witness path. With zero hops the answer is empty; with one hop from Atlas only Parser is reachable.
This breadth-first traversal returns one shortest witness, measured by the number of edges, for each reachable node. Sorted edge identifiers break ties deterministically. This is not a ranking by source quality, recency or probability. A shorter route can still rely on a weak source. If the task requires all supporting routes or a specific sequence of relations, this function is not a complete solution.
The relation selector is a set, so any chosen relation can be followed at any step. It does not express a pattern such as exactly one uses edge followed by exactly one maintained_by edge. Our two-hop Atlas result includes both Index and North. Filter or design an explicit query pattern when the question asks specifically for maintainers; do not call every reachable node a maintainer.
The edge budget counts eligible outgoing edges examined, including edges pointing to a node already seen. It does not count rejected relation types or all validation work. If another eligible edge would exceed the budget, the function raises an error instead of returning a partial answer as though it were complete. A production caller should turn that into an explicit incomplete-query state.
Save demo.py beside the other files. Its path is anchored to the script directory, so the fixture is found even if the terminal starts elsewhere. The file is local and no model, database server, account or paid tool is required.
import json
from pathlib import Path
from graph_lab import traverse
data = json.loads(Path(__file__).with_name("graph.json").read_text(encoding="utf-8"))
relations = {"uses", "depends_on", "maintained_by"}
for row in traverse(data, "atlas", relations, max_hops=2):
print(row["node"], ">".join(row["edges"]), ",".join(row["sources"]))python demo.pyparser e1 s1
index e1>e2 s1,s2
north e1>e4 s1,s3Try it yourself · Activity 03
20 minMeasure the stopping rules
Run the demo, then change only its query arguments in a disposable copy.
- Compare max_hops values zero, one and two.
- At two hops, try max_edges values two and three.
- Start at north and explain why the reverse route to Atlas is absent.
How would you phrase an error caused by a budget without saying that no relationship exists?
Worked answer
Zero hops returns no rows; one returns Parser; two returns Parser, Index and North. Budget two raises an error because e1, e2 and e4 must be examined. Budget three succeeds for the two-hop query. North has no outgoing edges, so starting there returns no rows. Following an incoming maintained_by edge would be a different query.
Read a sample · Chapter 04 of 06
Turn witness paths into careful answers
A retrieved path is an explanation of how the query reached a node. Preserve its edges and sources, then decide whether they support the exact sentence you intend to publish.
The North row contains e1 followed by e4. In the fixture, e1 says that Atlas uses Parser and e4 says that Parser is maintained by North. A careful answer is: the fictional register names North as maintainer of Parser, which the Atlas manifest lists as a used package. Cite both source records. Do not shorten this into North maintains Atlas, because the graph contains no such relationship.
Scroll sideways to see every column.
| Result | Witness | What the path supports |
|---|---|---|
| Parser | e1; s1 | Atlas uses Parser in the fixture. |
| Index | e1 then e2; s1 then s2 | Atlas uses a package that depends on Index. |
| North | e1 then e4; s1 then s3 | Atlas uses a package maintained by North. |
The edge IDs and source IDs are parallel lists. Reading each edge in order should begin at the query start and end at the returned node. Every relation must belong to the allowed set, and every cited source must be the source recorded on that edge. This is more inspectable than a list of unexplained entity names, but the validity of the underlying claims still depends on the sources.
The W3C PROV overview treats provenance as information about the entities, activities and people involved in producing something. Our source labels are a much smaller teaching device, not a PROV representation. A real extraction pipeline should also record what transformation created an edge, which document version it used and who or what reviewed it. A label alone is insufficient to replay that process.
Missing edges require careful language. No path to Archive means that the selected graph and query limits produced no path. It does not establish that Atlas never used Archive. The relationship might be absent from the source, missed by extraction, excluded by a filter or beyond the chosen hop bound. Record which of those possibilities you checked rather than inventing a negative fact.
Conflicting sources need a policy before answer generation. If two accepted records name different maintainers at different dates, keep the dates and versions, decide the relevant time and show disagreement when it cannot be resolved. Our fixture has no temporal schema and the traversal chooses just one path per node. Do not extend its claims to current ownership, trust or authorisation.
If you later add a language model, supply the exact permitted witness bundle and require claim-by-claim attribution. Treat retrieved source text as data, not as instructions to run tools or change the query. Check the resulting sentence against the bundle. Merely including citations in a prompt does not demonstrate that the answer faithfully uses them.
Try it yourself · Activity 04
15 minRewrite an overconfident answer
Review this proposed answer: North owns Atlas, and Atlas has no connection with Archive.
- Identify the relationship word that is not in the schema.
- Rewrite the North claim using the two recorded edges.
- Rewrite the Archive claim with the graph and query limits stated.
Would a reader be able to check every clause against the named records?
Worked answer
Owns is not a permitted relation, and there is no direct North-to-Atlas edge. A supported statement names North as maintainer of Parser, which Atlas uses, citing s1 and s3 in the fictional snapshot. For Archive: this query found no outgoing path from Atlas to Archive in the supplied graph within two hops. That is not proof of no real-world connection.
Read a sample · Chapter 05 of 06
Test the query and its evidence
Choose expected answers from the fixture before trusting the implementation. Test both which entities are returned and whether their explanation paths actually connect the right records.
Save the following as test_graph.py. It uses the standard-library test runner and the exact local fixture. Run the command below from the folder containing all four files. If an import fails or the fixture is missing, fix that environment error before interpreting anything as a graph test result.
import copy
import json
from pathlib import Path
import unittest
from graph_lab import traverse
DATA = json.loads(Path(__file__).with_name("graph.json").read_text(encoding="utf-8"))
ALL = {"uses", "depends_on", "maintained_by"}
class GraphTests(unittest.TestCase):
def test_expected_witnesses(self):
self.assertEqual(traverse(DATA, "atlas", ALL), [
{"node": "parser", "edges": ["e1"], "sources": ["s1"]},
{"node": "index", "edges": ["e1", "e2"], "sources": ["s1", "s2"]},
{"node": "north", "edges": ["e1", "e4"], "sources": ["s1", "s3"]},
])
def test_direction_and_filter(self):
self.assertEqual(traverse(DATA, "north", ALL), [])
self.assertEqual(traverse(DATA, "atlas", {"depends_on"}), [])
self.assertEqual(traverse(DATA, "archive", ALL), [])
def test_hops_and_cycles(self):
self.assertEqual(traverse(DATA, "atlas", ALL, max_hops=0), [])
rows = traverse(DATA, "atlas", ALL, max_hops=1)
self.assertEqual([r["node"] for r in rows], ["parser"])
rows = traverse(DATA, "parser", ALL, max_hops=5)
self.assertEqual([r["node"] for r in rows], ["index", "north"])
def test_order_and_input_are_preserved(self):
snapshot = copy.deepcopy(DATA)
reversed_rows = copy.deepcopy(DATA)
reversed_rows["edges"].reverse()
self.assertEqual(traverse(reversed_rows, "atlas", ALL), traverse(DATA, "atlas", ALL))
self.assertEqual(DATA, snapshot)
def test_invalid_edge(self):
for field, value in [("o", "absent"), ("p", "owns"), ("source", "absent"), ("o", "north")]:
with self.subTest(field=field, value=value):
broken = copy.deepcopy(DATA)
broken["edges"][0][field] = value
with self.assertRaises(ValueError):
traverse(broken, "atlas", ALL)
def test_duplicate_id(self):
broken = copy.deepcopy(DATA)
broken["edges"][1]["id"] = "e1"
with self.assertRaises(ValueError):
traverse(broken, "atlas", ALL)
def test_query_arguments(self):
for args in [{"max_hops": True}, {"max_hops": -1}, {"max_hops": 6}, {"max_edges": 0}]:
with self.subTest(args=args), self.assertRaises(ValueError):
traverse(DATA, "atlas", ALL, **args)
with self.assertRaises(ValueError):
traverse(DATA, "missing", ALL)
with self.assertRaises(ValueError):
traverse(DATA, "atlas", "uses")
def test_budget_rejects_incomplete_answer(self):
with self.assertRaisesRegex(ValueError, "budget"):
traverse(DATA, "atlas", ALL, max_edges=2)
self.assertEqual(len(traverse(DATA, "atlas", ALL, max_edges=3)), 3)
if __name__ == "__main__":
unittest.main()python -m unittest -v test_graphThe eight test methods cover exact witnesses, direction and filters, hop limits and cycles, row order and input preservation, invalid edges, duplicate identifiers, query arguments and the budget boundary. Assertions about sources matter as much as assertions about returned nodes. A graph could retrieve North correctly while attaching the wrong source to the explanation.
The release also ran an independent recursive reachability oracle across 210 combinations of start node, relation subset and hop bound. It checked every returned path and tried all 120 permutations of the five input edges. Together with rejection checks, this produced 950 additional assertions. These are recorded results for this tiny fixture, not measurements of general retrieval accuracy or production scale.
Three isolated mutations were tested: following the subject as the target, inventing the source ID s1 for every hop, and allowing an extra hop. All were caught. The extra-hop defect also exhausted an edge budget that the correct implementation satisfies. This illustrates why a test suite should check limits and evidence, not just one successful demo.
Try it yourself · Activity 05
20 minMake a test fail on purpose
Use a disposable copy to check that the learner suite detects a real defect.
- Run the unchanged suite and record all eight methods passing.
- Replace the expression sources + [edge["source"]] with sources + ["s1"] in the copied implementation.
- Run the suite again, identify the failing expectation, then restore the original file and rerun.
What defect would escape a test that checks only the set of returned node names?
Worked answer
The exact-witness test fails because the Index and North paths require s2 and s3 respectively for their second edge. The returned node names can still look correct. Restoring the implementation restores the pass. A useful record includes the changed line, test result and restored result, without claiming that one mutation proves all cases.
Read a sample · Chapter 06 of 06
Connect graphs to a retrieval workflow
Use structure where relationships are explicit and valuable. Keep documents available for details that the graph does not represent, and evaluate the combined system against questions with known evidence.
A practical workflow separates ingestion, candidate extraction, review, graph validation, query planning, retrieval and answer checking. Each boundary has a different job. An extractor proposes entities and relationships from documents; a reviewer accepts or rejects them; validation checks the graph shape; retrieval follows permitted edges; answer checking compares each sentence with the retrieved evidence.
Text retrieval and graph retrieval solve different parts of the problem. A graph query can enforce a relation pattern when entity IDs and edges are reliable. Text search can locate supporting passages, definitions or information that was never represented as an edge. Combining them can help, but adds work: the entity mapping, source version and passage must all refer to the same thing.
For the Atlas question, first resolve Atlas to its exact project identifier. Retrieve a bounded path, collect its source IDs, then fetch only the corresponding reviewed passages from the same snapshot. If the passage no longer supports the edge, stop or mark the answer unresolved. Do not silently replace it with a similar passage from another revision.
Build a small evaluation sheet before adding more automation. Include a direct relationship, a two-hop relationship, a reverse-direction trap, a cycle, an isolated node, ambiguous entity names, conflicting revisions and a budget failure. Record expected entities, required evidence and the allowed abstention wording. Report retrieval correctness separately from whether a generated answer faithfully states it.
This lab has no persistence layer, authorisation model, source-fetching code or language-model call. The JSON is loaded into memory and structurally validated on every query. Before real deployment, define access restrictions, snapshot/version handling, byte limits, logging that avoids sensitive source text, update and withdrawal procedures, and performance measurements on representative data. Do not present this small script as a complete private knowledge-base service.
Finish with a handover another person can reproduce. Include the four files, interpreter version, exact query, expected output and tests actually run. Name the fictional nature of the fixture and the one-witness limitation. Record deferred decisions separately from completed checks. A graph drawing can support the explanation, but retain the adjacency table and evidence paths as accessible text.
Try it yourself · Activity 06
20 minPrepare a retrieval acceptance sheet
Draft a small acceptance sheet and handover for the lab, then propose one carefully bounded extension.
- Write one positive, one missing-evidence and one invalid-input case, with expected output or error.
- Record the exact four files, Python version and commands used.
- Propose an exact uses-then-maintained_by query and explain how it differs from the current relation set.
Which part of the next feature can you validate without introducing a model or a database?
Worked answer
A positive case is Atlas at two hops returning Parser, Index and North with their recorded witnesses. Archive from Atlas produces no path within the chosen graph and bound, without proving a universal negative. An e1 edge ending at North is structurally invalid. The extension must constrain the first edge to uses and the second to maintained_by; allowing both types at every step is a different query. Include tests for order, missing second edges and evidence for both steps.
Keep learning
The complete workbook
Model a fictional project and its package dependencies using stable identifiers and sourced relationships. Implement a bounded traversal, inspect its evidence paths and test failure cases before considering a language-model answer layer. The complete lab runs locally with Python standard-library modules.
- 01Model a question before drawing a graphRead here · 1 exercise
Start with the question you need to answer and the evidence you can inspect. A diagram can make relationships visible, but a readable picture does not establish that its edges are true.
- 02Validate relationships before retrievalRead here · 1 exercise
A malformed graph should fail clearly before producing an answer. Validation checks the contract of this small application; it does not verify a source document or solve entity resolution.
- 03Follow bounded paths and handle cyclesRead here · 1 exercise
A traversal needs a starting entity, allowed relation types and a stopping rule. Return evidence with the result instead of asking a language model to reconstruct the route afterwards.
- 04Turn witness paths into careful answersRead here · 1 exercise
A retrieved path is an explanation of how the query reached a node. Preserve its edges and sources, then decide whether they support the exact sentence you intend to publish.
- 05Test the query and its evidenceRead here · 1 exercise
Choose expected answers from the fixture before trusting the implementation. Test both which entities are returned and whether their explanation paths actually connect the right records.
- 06Connect graphs to a retrieval workflowRead here · 1 exercise
Use structure where relationships are explicit and valuable. Keep documents available for details that the graph does not represent, and evaluate the combined system against questions with known evidence.
Also inside: a 8-point checklist, a glossary of 8 terms and 10 questions and answers to test yourself. 6 hands-on exercises, each with a worked answer at the back where the workbook gives one.
No login, no card, no account. Before the download we ask you to follow Mickai (two quick links). Free to download and use for personal learning, study groups and inside your own team. Please do not resell the workbooks or republish them as your own. Link people to trust-agent.ai instead.
Test yourself
Questions and answers
What does a graph add to retrieval?
It represents explicit relationships between identified entities, allowing a query to follow defined directions and patterns. The benefit depends on correct entity mapping and supported edges; a graph does not make an assertion true.
Is the fixture an RDF implementation?
No. It is an application-specific JSON structure with local validation rules. RDF, JSON-LD and SPARQL have their own data and query contracts. Similar-looking triples do not make this lab standards-compatible.
Why reject repeated edge identifiers?
A witness must identify one edge record unambiguously. Different IDs may record the same relationship with different sources, but the lab returns only one witness per reached node.
Does a known source ID prove the edge is correct?
No. It proves only that the key exists in the source dictionary. A real reviewer must inspect the document version and supporting passage, including ambiguity or disagreement.
How does the traversal stop a cycle?
The seen set includes the start node and every queued target. A target already seen is not queued again. Hop and eligible-edge budgets supply separate limits.
Does the relation set express an exact sequence?
No. Any relation in the set can be followed at any step. A uses-then-maintained_by pattern needs separate logic that constrains the relation at each position.
Why must North not be described as owning Atlas?
The fixture has Atlas uses Parser and Parser maintained_by North. It has neither an owns relation nor a direct North-to-Atlas claim. The answer must preserve the meaning of both recorded edges.
What does no path to Archive mean?
The selected graph and query limits yielded no path. It does not establish that no real relationship exists; evidence might be missing, excluded or beyond the hop limit.
Why mutate a source identifier in a test?
The returned node names may stay correct while their evidence becomes wrong. An exact witness test detects that error; a node-set-only test can miss it.
What must be added before using private real-world documents?
Define source versions, access restrictions, ingestion byte limits, review and withdrawal procedures, and evaluation on representative questions. This local lab implements none of those service boundaries.
When you have finished
Get your certificate of completion
Type your name and download a certificate for this workbook as a PDF, ready to print or to add to LinkedIn. It is made on your own device, so your name is never sent to us. It is a self-declared certificate, not an accredited qualification.
Learn the language
Key terms
- Entity
- A represented thing with a stable identifier.
- Relation
- A directed connection with defined meaning and kinds.
- Adjacency list
- Each node mapped to its outgoing edges.
- Witness path
- Ordered edge records explaining a reached node.
- Provenance
- Information about a record or result's origin.
- Hop
- One traversed edge in a path.
6 of the workbook's 8 terms. The complete glossary is in the workbook.
Follow the evidence
Sources and checks
Facts last checked: .
Examples in this workbook were run on: Python 3.12.10 on Windows 11; eight learner tests, 950 additional assertions, 210 query cases, 120 edge-order permutations and three caught mutations. No network or model calls. (2026-09-27).
These workbooks use AI assistance. See how the workbooks are made.
- RDF 1.1 PrimerW3C
- PROV OverviewW3C
- Python 3.12: collections and dequePython Software Foundation
- Python 3.12: JSON encoder and decoderPython Software Foundation
Created by Mickarle Wagstaff-Irons - Micky Irons with the Mickai team. Published by Mickai LTD. Last updated 27 September 2026.
NextKeep going
Where to go next
Recommended for you
Document pipelines: parsing, cleaning and chunking
Build a local document pipeline: clean UTF-8 text, create traceable chunks, test boundaries and measure which evidence survives together.
Recommended for you
Design a repeatable evaluation suite for an AI application
Build an offline AI evaluation harness with test cases, a rubric, regression comparisons, uncertainty checks and a reproducible release record.