agents · Level 4
Orchestrating multiple agents with explicit responsibilities
Make roles, handoffs, dependencies and stop conditions explicit before adding more autonomy.

Start with the essentials
The short answer
Orchestration assigns work and decides when a result is acceptable, what depends on it and when to stop. Separate roles only when a measured benefit justifies the extra coordination. This course implements the control layer with scripted workers: a fixed task graph, strict reply checks, task-and-attempt binding, reference lineage and explicit terminal states. It makes no model calls and demonstrates neither autonomous planning nor parallel execution.
What you will learn
- Decide whether a task needs multiple roles or a simpler baseline.
- Define task ownership, dependencies and an acyclic execution plan.
- Validate replies against task identity, attempt number and available references.
- Apply one shared call budget across successful and failed attempts.
- Test failed dependencies, malformed replies and bounded retries.
- Prepare an evidence-based handover without overstating agent independence or reliability.
Who it is for
Builders who understand agent loops and structured tool outputs and want a small orchestration contract they can inspect.
Before you start
- Read Python functions, dictionaries, dataclasses and unit tests.
- Understand an agent loop and the distinction between a model reply and permission to act.
- Use Python 3.12 in a fresh folder with fictional data and no external tools.
Read a sample · Chapter 01 of 06
Give each role a reason to exist
Start with a task and a comparison, not a collection of personalities.
Our fictional Orchard lab needs one checked sentence about its opening time. A research role reads an invented source, a draft role turns the finding into text, and a review role compares the draft with that source. The role names identify responsibilities. They are not people, independent witnesses or evidence that three model calls improve an answer. All three roles in this workbook are short scripted Python functions so that the coordinator can be tested without a model account.
Anthropic distinguishes workflows with prescribed control paths from agents that dynamically decide how to proceed. Its building-effective-agents article recommends starting simply and adding complexity when it improves results. Our implementation sits on the workflow side: the host supplies the complete graph, role assignment and limits before execution. Replacing a worker with a model adapter would change that worker, but would not make this fixed coordinator an autonomous planner.
A useful baseline for the fictional task is one function reading the note and returning its sentence. It has no handoff to corrupt and no reviewer call to pay for. Split the work only if you can name a benefit to measure: different source access, a specialist quality check, an independent calculation or separate permissions. If every worker sees the same context and repeats the same unsupported assertion, the added conversation may provide no useful evidence.
Scroll sideways to see every column.
| Role | Input, output and responsibility |
|---|---|
| Research | Read the fictional source. Return a finding and source IDs. No external retrieval. |
| Draft | Receive only accepted dependency results. Return a proposed sentence with inherited reference IDs. |
| Review | Receive the accepted draft. Compare it with the original fixture through trusted adapter code. |
| Coordinator | Own the plan, attempt numbers, budget, acceptance checks and final states. |
Microsoft's orchestration-pattern guidance compares sequential, concurrent and handoff designs. These are choices about control flow, not decorative role labels. We use a fixed sequence with dependency edges and execute one call at a time. A graph can describe independent tasks without running them concurrently. The course therefore claims no latency improvement, parallel speedup or distributed execution.
Define the acceptance question before choosing a framework. For the coordinator it is: does only a correctly addressed, structurally valid reply become a dependency result within the total call allowance? For content quality it is: does the sentence actually match the authoritative source and answer the user's question? The first question can be checked by generic code. The second needs task-specific evidence and, for open-ended text, a carefully designed evaluation.
Create a fresh folder and save coordinator.py, demo.py and test_coordinator.py as printed. Coordinator.py continues across three chapters; append its labelled blocks in order with one blank line between blocks and preserve indentation. This is standard-library Python, executed with 3.12.10 on Windows. The version identifies the tested environment, not the newest Python release. Nothing in the lab contacts a provider, installs an agent framework, retrieves a live page or sends an external message.
Try it yourself · Activity 01
15 minJustify one handoff
Compare a single function with the three-role Orchard workflow.
- State one possible benefit of a separate review role.
- State two extra failure modes caused by splitting the work.
- Choose one quality measure and one resource measure for a future comparison.
Would a deterministic check answer the review question more directly than another model?
Worked answer
A separate review could compare a proposed sentence with source evidence. Extra failures include a reply addressed to the wrong task and a broken dependency that still reaches the writer. A comparison could measure source-supported answers on a fixed held-out set and total provider calls per completed task. The scripted fixture demonstrates neither a quality uplift nor provider cost.
Read a sample · Chapter 02 of 06
Validate the plan before any worker runs
A dependency means accepted output is required, not merely that someone was asked.
Start coordinator.py with the following blocks. Task identifies the work, its role, direct dependencies and prompt. Result is an accepted body with reference IDs. Call carries one task, a host-assigned attempt number, accepted dependency results and the allowed reference set. Event records a controlled status code. Report contains immutable tuples of final states, accepted results and events, plus the number of attempted worker calls.
"""A fixed, serial workflow. Workers are trusted code returning untrusted JSON."""
from dataclasses import dataclass
from graphlib import CycleError, TopologicalSorter
import json
import re
class ContractError(ValueError):
pass
@dataclass(frozen=True)
class Task:
id: str
role: str
needs: tuple
prompt: str
@dataclass(frozen=True)
class Result:
task: str
body: str
refs: tuple@dataclass(frozen=True)
class Call:
task: Task
attempt: int
inputs: tuple
allowed_refs: tuple
@dataclass(frozen=True)
class Event:
task: str
attempt: int
code: str
@dataclass(frozen=True)
class Report:
states: tuple
results: tuple
events: tuple
calls: int
def identifier(value):
return type(value) is str and re.fullmatch(r"[a-z][a-z0-9-]{0,31}", value)def order_plan(tasks, workers, sources, budget, retries):
if type(tasks) is not tuple or not 1 <= len(tasks) <= 8:
raise ContractError("one to eight tasks required")
if type(budget) is not int or not 1 <= budget <= 16:
raise ContractError("invalid call budget")
if type(retries) is not int or not 0 <= retries <= 2:
raise ContractError("invalid retry limit")
if (type(sources) is not tuple or len(sources) > 8
or not all(identifier(s) for s in sources)
or len(set(sources)) != len(sources)):
raise ContractError("invalid sources")
by_id = {}
for task in tasks:
if type(task) is not Task or not identifier(task.id):
raise ContractError("invalid task")
if task.id in by_id:
raise ContractError("duplicate task")
if (task.role not in ("research", "draft", "review")
or not callable(workers.get(task.role))):
raise ContractError("unknown or missing worker")
if (type(task.prompt) is not str or not task.prompt.strip()
or len(task.prompt) > 500):
raise ContractError("invalid prompt")
if (type(task.needs) is not tuple or len(task.needs) > 8
or not all(identifier(n) for n in task.needs)
or len(set(task.needs)) != len(task.needs)):
raise ContractError("invalid dependencies")
by_id[task.id] = task
if any(n not in by_id for t in tasks for n in t.needs):
raise ContractError("unknown dependency")
try:
order = tuple(TopologicalSorter({t.id: t.needs for t in tasks}).static_order())
except CycleError as error:
raise ContractError("cyclic plan") from error
return tuple(by_id[task_id] for task_id in order)The host supplies a tuple containing one to eight Task values. IDs use a small lowercase grammar and must be unique. Roles must be research, draft or review and must have a callable adapter. Prompts are non-empty and at most 500 Python characters. Dependency names must be unique within a task and must name tasks in this plan. These modest limits make the exercise inspectable; they are not a full admission-control policy for a service.
A missing dependency is rejected explicitly before graph sorting. That matters because Python's TopologicalSorter can add a mentioned predecessor that was not separately defined as a node. In our application, such an implied node would have no role or prompt. The application contract is stricter: every dependency must have an explicit Task before any calls occur.
Topological order places each predecessor before its dependent. We materialise the complete order as a tuple, so a cycle raises ContractError before executing even an otherwise independent root. Python's graphlib documents that the order of equally ready nodes can depend on insertion order. Our host-provided tuple is therefore part of reproducibility. The scheduler does not optimise priorities or fairness among independent roots.
Do not treat the graph as a permission system. A research label does not remove filesystem or network access from its Python function. All adapters are trusted code in the same process. The checked reply is untrusted data returned by that code. A real model adapter must expose only its authorised tools and keep identity, credentials and workflow mutations outside model-supplied text. A separate process or service has different failure and isolation properties.
The planning phase is deliberately fixed. A model could propose a graph for a larger application, but trusted code would need to validate allowed roles, reachable resources, graph size, repeated work and the total plan budget before accepting it. An accepted worker reply in this lab cannot add tasks, change roles or request a new branch. Such instructions can appear in its body as text; the coordinator never interprets them as scheduler commands.
Try it yourself · Activity 02
15 minFind the invalid plans
Sketch find -> write -> check, then introduce a cycle and a missing dependency.
- Make find depend on check and predict whether any worker runs.
- Replace write's dependency with an absent task ID.
- Reverse the valid plan's tuple order and predict the execution order.
What information is missing if a task names a predecessor that has no explicit definition?
Worked answer
The cycle and unknown dependency are rejected before any worker call. The reversed valid tuple still executes find, write, check because dependency order takes precedence. Independent tasks may retain insertion-related ordering, so reversing unrelated roots can change their call order without violating the graph.
Read a sample · Chapter 03 of 06
Make handoffs small and checkable
An accepted envelope establishes structure and lineage; it does not establish truth.
Continue coordinator.py with the reply parser below. A worker must return a Python string containing exactly one JSON object with task, attempt, body and refs. The task ID and integer attempt must match the Call supplied by the coordinator. A valid reply for a previous attempt is not accepted for a later one. Role and state are not reply fields; the worker cannot claim ownership of another role or mark a different task complete.
def unique_object(pairs):
result = {}
for key, value in pairs:
if key in result:
raise ContractError("duplicate reply field")
result[key] = value
return result
def reject_constant(value):
raise ContractError("non-JSON number")def accept_reply(raw, call):
try:
if type(raw) is not str or len(raw.encode("utf-8")) > 4096:
raise ContractError("reply size or type")
data = json.loads(raw, object_pairs_hook=unique_object,
parse_constant=reject_constant)
except (ValueError, RecursionError) as error:
raise ContractError("invalid JSON reply") from error
if type(data) is not dict or set(data) != {"task", "attempt", "body", "refs"}:
raise ContractError("reply fields")
if (data["task"] != call.task.id or type(data["attempt"]) is not int
or data["attempt"] != call.attempt):
raise ContractError("reply identity")
body, refs = data["body"], data["refs"]
try:
valid_body = (type(body) is str and bool(body.strip())
and len(body.encode("utf-8")) <= 800)
except UnicodeError:
valid_body = False
if not valid_body:
raise ContractError("reply body")
if (type(refs) is not list or len(refs) > 8
or not all(type(ref) is str and ref in call.allowed_refs for ref in refs)
or len(set(refs)) != len(refs)):
raise ContractError("reply references")
return Result(call.task.id, body, tuple(refs))The raw reply is limited to 4,096 UTF-8 bytes after it already exists as a string. Duplicate object keys and non-standard numeric constants are refused. Parsing successfully is only the first stage: exact field names, task identity, attempt type and body/reference limits still apply. Boolean True is refused as an attempt number even though Python treats bool as an integer subclass. The course checks this distinction explicitly.
The body must contain non-whitespace text and at most 800 UTF-8 bytes. Four hundred repetitions of the escaped character \u00e9 occupy 800 bytes and fit; 401 do not. The original string is retained rather than trimmed, so leading or trailing whitespace can remain within the byte bound. Invalid Unicode surrogate content is refused. The input cap does not bound the worker's memory use, transport receive size or time spent producing the string.
References must be a JSON list with at most eight unique strings. A root task may cite the host-provided source IDs. A dependent may cite only IDs carried by its direct accepted inputs. The writer cannot silently add s2 when research passed only s1. Empty reference lists are allowed by this structural contract; an application that requires every claim to have support needs a separate completeness policy.
An allowed reference ID is not evidence that a sentence follows from that source. A worker could cite s1 while saying the Orchard lab opens at midnight. The generic parser would accept the known ID and body shape. Our scripted reviewer checks an exact fictional sentence; a real reviewer would need access to the actual source content, a support rubric and a way to handle uncertainty or missing evidence. Structural success must not be displayed as a fact-check pass.
Results and Calls are frozen dataclasses containing strings, tuples and other frozen records. Ordinary assignment to their fields is refused, which helps prevent accidental handoff mutation. This is not process isolation or protection against malicious Python code. Type annotations also do not enforce runtime types by themselves; explicit checks at the plan and reply boundaries establish the types used here.
Task and attempt binding operate only inside one run. There is no run ID, cryptographic signature or persistent inbox. A future asynchronous adapter would need to correlate the workflow run as well as task and attempt, reject late messages and make acceptance conditional on the current state. Reusing task IDs after a process restart would require an explicit namespace and recovery policy.
Try it yourself · Activity 03
20 minSeparate a valid envelope from a supported claim
Imagine a reply for find, attempt 1, body “Orchard opens at midnight”, refs [s1].
- Explain which generic checks it passes.
- Compare it with the fictional source sentence.
- Design one task-specific acceptance check and one uncertainty outcome.
Would two reviewers citing the same mistaken intermediate result provide independent evidence?
Worked answer
The envelope can pass task, attempt, field, size and known-reference checks. It conflicts with s1, which says 09:00. A fixture-specific check can require the exact recorded opening time; an open-ended reviewer should instead compare the claim with source content and report unsupported or uncertain when evidence is insufficient. Known-reference membership alone is not factual validation.
Read a sample · Chapter 04 of 06
Schedule retries within one shared budget
Every attempt costs a call, including malformed replies and exceptions.
One shared budget. Every attempted call counts.
The task IDs are find, write and check, with
research, draft and review roles respectively. Write depends on find; check depends
on write. The host validates the whole fixed graph before running it serially. Each
task may retry once: at most two invocations per task, also subject to the shared
run budget. The first draft reply is deliberately not JSON.
| Step | Outcome | Calls used |
|---|---|---|
| find / attempt 1 | accepted | 1 |
| write / attempt 1 | invalid-reply | 2 |
| write / attempt 2 | accepted | 3 |
| check / attempt 1 | accepted | 4 |
A call is counted before invoking its worker. The invalid reply therefore costs
one call. The second draft attempt is accepted, and review receives that accepted
result. The draft is The fictional Orchard lab opens at 09:00., and the
review body is Exact fixture match.. This narrow review compares one
fictional sentence; it is not a general or independent factual evaluator.
| Step | Outcome | Calls used |
|---|---|---|
| find / attempt 1 | accepted | 1 |
| write / attempt 1 | invalid-reply | 2 |
| write / next attempt 2 | budget limit; not called | 2 |
| check / dependency | blocked; not called | 2 |
After the invalid draft reply, both calls are spent. The next attempt reaches the budget check and is never invoked. Check then has an unsuccessful dependency, so its worker is never invoked either. These two scheduler decisions cost no calls. Only find has an accepted result in this run; there is no accepted draft to review.
| Task | Budget 4 | Budget 2 |
|---|---|---|
| find | succeeded | succeeded |
| write | succeeded | budget_exhausted |
| check | succeeded | blocked |
The source ID s1 accompanies accepted handoffs. A separate parser
fixture accepts the false body The fictional Orchard lab opens at midnight.
with refs=[s1]: allowed lineage does not prove factual support. Changing
that reference to unknown s2 is refused with reply references.
Changing the task ID from find to write is refused with reply identity.
Those parser fixtures do not invoke workers or spend the scheduler's call budget.
This is a serial, in-memory workflow with scripted workers, not an autonomous planner or parallel-agent benchmark. Roles and frozen records do not sandbox trusted executable workers. A call cap is not a timeout, token allowance, provider-cost cap or fairness policy. A worker that never returns can stall the run. Persistent run identity, concurrent scheduling, deadlines, cancellation, restart recovery and external-effect idempotency need separate design. No production or browser acceptance is implied.
Complete coordinator.py with run below. The coordinator snapshots the adapter mapping and validates the entire plan. It walks a topological order serially. If any direct dependency did not succeed, the task becomes blocked without calling its adapter. Otherwise its inputs contain only accepted direct dependency results. Root reference scope comes from the supplied source IDs; dependent scope is the sorted union carried by its inputs.
def run(tasks, workers, sources=(), budget=8, retries=1):
workers = dict(workers)
order = order_plan(tasks, workers, sources, budget, retries)
states, results, events = {}, {}, []
calls = 0
for task in order:
if any(states[n] != "succeeded" for n in task.needs):
states[task.id] = "blocked"
events.append(Event(task.id, 0, "dependency"))
continue
inputs = tuple(results[n] for n in task.needs)
allowed = (tuple(sorted({ref for result in inputs for ref in result.refs}))
if task.needs else sources)
for attempt in range(1, retries + 2):
if calls >= budget:
states[task.id] = "budget_exhausted"
events.append(Event(task.id, attempt, "budget"))
break
calls += 1
call = Call(task, attempt, inputs, allowed)
events.append(Event(task.id, attempt, "started"))
try:
raw = workers[task.role](call)
except Exception:
events.append(Event(task.id, attempt, "worker-error"))
else:
try:
result = accept_reply(raw, call)
except ContractError:
events.append(Event(task.id, attempt, "invalid-reply"))
else:
results[task.id] = result
states[task.id] = "succeeded"
events.append(Event(task.id, attempt, "accepted"))
break
states[task.id] = "failed"
return Report(tuple(states.items()), tuple(results.values()), tuple(events), calls)The shared budget is an integer from 1 to 16. Retries is an integer from 0 to 2, so a task can make at most three attempts. The budget is checked before each call and incremented before invoking the worker. A malformed reply or ordinary adapter exception still consumes an attempt. Budget exhaustion ends that task as budget_exhausted; descendants later become blocked because they have no accepted input.
Each attempted call produces started followed by accepted, invalid-reply or worker-error. Blocked and budget-exhausted tasks produce dependency or budget events without a worker invocation. Failed means the permitted attempts ended without acceptance. Succeeded means a structurally valid Result was stored. These are workflow states, not claims about factual correctness, external delivery or user satisfaction.
The coordinator catches ordinary Exception values from adapters and records a fixed worker-error code rather than the exception text. KeyboardInterrupt is not swallowed, so an operator can stop the process. That distinction does not create a timeout: a worker that never returns can still hang the entire run. A real remote adapter needs transport deadlines, cancellation and resource supervision appropriate to its environment.
The next file is demo.py. Research returns a fictional opening-time note. Draft deliberately returns non-JSON on its first attempt and succeeds on the second. Review compares the accepted draft to the original source fixture available in its trusted adapter. Run python demo.py after saving both files. The five lines below are the observed output; the long events line may wrap visually while remaining one printed line.
"""Fictional scripted roles: no model, retrieval service or external action."""
import json
from coordinator import Task, run
PLAN = (
Task("find", "research", (), "Read the fictional lab note."),
Task("write", "draft", ("find",), "Draft one sentence from the finding."),
Task("check", "review", ("write",), "Check the supplied draft against its reference."),
)
SOURCES = {"s1": "The fictional Orchard lab opens at 09:00."}
def reply(call, body, refs):
return json.dumps({"task": call.task.id, "attempt": call.attempt,
"body": body, "refs": refs})
def research(call):
return reply(call, SOURCES["s1"], ["s1"])
def draft(call):
if call.attempt == 1:
return "not JSON"
return reply(call, call.inputs[0].body, list(call.inputs[0].refs))def review(call):
agrees = call.inputs[0].body == SOURCES["s1"]
verdict = "Exact fixture match." if agrees else "Needs review."
return reply(call, verdict, list(call.inputs[0].refs))
WORKERS = {"research": research, "draft": draft, "review": review}
def main():
report = run(PLAN, WORKERS, tuple(SOURCES), budget=4, retries=1)
print("states", dict(report.states))
print("calls", report.calls)
print("draft", next(r.body for r in report.results if r.task == "write"))
print("review", next(r.body for r in report.results if r.task == "check"))
print("events", " ".join(f"{e.task}:{e.attempt}:{e.code}" for e in report.events))
if __name__ == "__main__":
main()states {'find': 'succeeded', 'write': 'succeeded', 'check': 'succeeded'}
calls 4
draft The fictional Orchard lab opens at 09:00.
review Exact fixture match.
events find:1:started find:1:accepted write:1:started write:1:invalid-reply write:2:started write:2:accepted check:1:started check:1:acceptedFour calls are used: research once, draft twice, review once. A budget of two permits research and the malformed draft attempt, then stops the draft before its retry. The reviewer is blocked. When one branch fails, a separate root can still run if budget remains. The simple scheduler does not reserve capacity for later tasks, so an early flaky task can consume calls that another task would have used.
Retries are safe here only because the scripted workers have no consequential external effects. If an adapter sends a message and then loses its response, repeating it may send twice. An attempt number is not an external idempotency key. A real system needs effect-specific deduplication, an ambiguous-outcome state and recovery procedures before retrying writes, purchases or sends. This course performs none of those actions.
Try it yourself · Activity 04
20 minAccount for every call
Trace the demo with budget 4, then repeat it with budget 2.
- List each invocation and its attempt number.
- Distinguish the malformed reply from an unattempted retry.
- Record the final state of each task and the accepted results that remain.
What fairness policy would you add if one independent branch must not consume the whole allowance?
Worked answer
Budget 4 invokes find:1, write:1, write:2 and check:1, producing three accepted results. Budget 2 invokes only find:1 and write:1. Find succeeds, write becomes budget_exhausted before attempt 2 is called, and check is blocked. The report retains only the accepted find result. A failed reply still consumed its call.
Read a sample · Chapter 05 of 06
Test both progress and refusal
A coordinator that never runs anything is bounded but not useful.
Save test_coordinator.py below and run python -m unittest -v test_coordinator.py. The fourteen tests cover the successful retrying workflow, shared budget, blocked descendants, independent progress, invalid plans, integer limits, reversed plans, reply identity, JSON grammar, UTF-8 size, raw reply bounds, reference membership, dependency input scope, bounded exceptions and an operator stop. They require the intended successful results as well as refusal behaviour.
import json
from dataclasses import FrozenInstanceError
import unittest
from coordinator import Call, ContractError, Result, Task, accept_reply, run
from demo import PLAN, SOURCES, WORKERS, reply
class CoordinatorTests(unittest.TestCase):
def good(self, call):
return reply(call, "Fictional finding", list(call.allowed_refs))
def workers(self, worker=None):
return dict.fromkeys(("research", "draft", "review"), worker or self.good)
def test_complete_retrying_workflow(self):
report = run(PLAN, WORKERS, tuple(SOURCES), budget=4)
self.assertEqual(dict(report.states), dict.fromkeys(("find", "write", "check"), "succeeded"))
self.assertEqual(report.calls, 4)
self.assertEqual([e.code for e in report.events].count("invalid-reply"), 1)
self.assertEqual(report.results[-1].body, "Exact fixture match.")
def test_budget_is_shared_and_counts_bad_replies(self):
report = run(PLAN, WORKERS, tuple(SOURCES), budget=2)
self.assertEqual(report.calls, 2)
self.assertEqual(dict(report.states), {"find": "succeeded", "write": "budget_exhausted", "check": "blocked"})
self.assertEqual([r.task for r in report.results], ["find"]) def test_failed_dependency_blocks_only_descendants(self):
called = []
def worker(call):
called.append(call.task.id)
return "bad" if call.task.id == "find" else self.good(call)
plan = PLAN + (Task("other", "research", (), "Independent task"),)
report = run(plan, self.workers(worker), ("s1",))
self.assertEqual(called, ["find", "find", "other"])
self.assertEqual(dict(report.states)["write"], "blocked")
self.assertEqual(dict(report.states)["check"], "blocked")
self.assertEqual(dict(report.states)["other"], "succeeded")
def test_invalid_plan_calls_no_workers(self):
def forbidden(call):
self.fail("Invalid plan ran a worker")
bad_plans = [(), (PLAN[0], PLAN[0]),
(Task("a", "research", ("missing",), "x"),),
(Task("a", "research", ("a",), "x"),),
(Task("a", "research", ("b",), "x"), Task("b", "draft", ("a",), "x")),
(Task("a", "shell", (), "x"),),
(Task("a", "research", (), " "),),
(Task("a", "research", ["b"], "x"),),
(Task("a", "research", ("b", "b"), "x"), PLAN[0])]
for plan in bad_plans:
with self.subTest(plan=plan), self.assertRaises(ContractError):
run(plan, self.workers(forbidden)) def test_limits_are_exact_integers(self):
for options in ({"budget": True}, {"budget": 0}, {"budget": 17},
{"retries": True}, {"retries": -1}, {"retries": 3}):
with self.subTest(options=options), self.assertRaises(ContractError):
run(PLAN, self.workers(), **options)
with self.assertRaises(ContractError):
run(PLAN, self.workers(), ("s1", "s1"))
with self.assertRaises(ContractError):
run(PLAN, {})
def test_reversed_plan_obeys_dependencies(self):
report = run(tuple(reversed(PLAN)), self.workers(), ("s1",))
self.assertEqual([r.task for r in report.results], ["find", "write", "check"])
self.assertEqual(report.calls, 3)
def test_exact_reply_identity(self):
call = Call(PLAN[0], 2, (), ("s1",))
base = json.loads(reply(call, "hello", ["s1"]))
for changes in ({"task": "write"}, {"attempt": 1}, {"attempt": True}, {"role": "review"}):
with self.subTest(changes=changes), self.assertRaises(ContractError):
accept_reply(json.dumps(base | changes), call) def test_reply_grammar(self):
call = Call(PLAN[0], 1, (), ())
for raw in (None, "bad", "[]", "null", '{"task":1,"task":2}', '{"x":NaN}', "["*5000):
with self.subTest(raw=str(raw)[:30]), self.assertRaises(ContractError):
accept_reply(raw, call)
def test_body_utf8_and_nonempty_limits(self):
call = Call(PLAN[0], 1, (), ())
self.assertEqual(len(accept_reply(reply(call, "\u00e9"*400, []), call).body), 400)
for body in ("\u00e9"*401, "", " ", "\ud800", 7):
with self.subTest(body=str(body)[:20]), self.assertRaises(ContractError):
accept_reply(reply(call, body, []), call)
def test_raw_reply_limit(self):
call = Call(PLAN[0], 1, (), ())
raw = reply(call, "ok", [])
padded = raw + " "*(4096-len(raw.encode("utf-8")))
self.assertEqual(accept_reply(padded, call).body, "ok")
with self.assertRaises(ContractError):
accept_reply(padded+" ", call)
def test_references_are_subset_unique_and_typed(self):
call = Call(PLAN[0], 1, (), ("s1",))
for refs in (["s2"], ["s1", "s1"], [1], "s1"):
with self.subTest(refs=refs), self.assertRaises(ContractError):
accept_reply(reply(call, "ok", refs), call) def test_only_dependency_results_reach_each_worker(self):
observed = []
def worker(call):
observed.append((call.task.id, tuple(r.task for r in call.inputs), call.allowed_refs))
return reply(call, "x", ["s1"] if call.task.id == "find" else [])
run(PLAN, self.workers(worker), ("s1", "s2"))
self.assertEqual(observed, [("find", (), ("s1", "s2")),
("write", ("find",), ("s1",)), ("check", ("write",), ())])
def test_worker_error_is_bounded_and_not_logged_raw(self):
def fail(call):
raise RuntimeError("fictional private note text")
report = run((PLAN[0],), self.workers(fail), retries=2)
self.assertEqual(report.calls, 3)
self.assertEqual(dict(report.states), {"find": "failed"})
self.assertEqual([e.code for e in report.events].count("worker-error"), 3)
self.assertNotIn("private note", repr(report.events))
def test_frozen_handoff_and_stop_signal(self):
result = Result("find", "x", ("s1",))
with self.assertRaises(FrozenInstanceError):
result.body = "changed"
def stop(call):
raise KeyboardInterrupt
with self.assertRaises(KeyboardInterrupt):
run(PLAN, self.workers(stop))if __name__ == "__main__":
unittest.main()The budget test is important because retry limits alone are local. Eight tasks with three attempts each could otherwise invoke workers 24 times. The shared allowance is a second, run-wide constraint. It counts all attempted calls but says nothing about tokens, money, generated bytes, wall-clock time or subcalls hidden inside a trusted adapter. If one adapter internally contacts three models, the coordinator still observes one Python invocation.
The invalid-plan test installs a worker that immediately fails the test if called. That distinguishes preflight rejection from detecting a cycle after work has already started. The dependency test makes research return invalid data twice, then requires write and check to remain blocked while an independent root succeeds. Passing only an always-fail test would not establish this combination of isolation and progress.
Three deliberately defective copies were also executed in temporary folders. One removed the task-ID comparison; one ignored the shared-budget condition; one accepted arbitrary reference IDs. Each remained syntactically valid but failed the relevant learner test. The original files were unchanged. This is evidence that these three protections are exercised, not a measure of all possible failures or an independent security audit.
There is no model-quality benchmark in this release. A future evaluation should hold the task set and source snapshots fixed, compare the single-worker baseline with the proposed orchestration, and report accepted-result quality alongside completion rate, calls, latency and human correction effort. Include malformed replies, conflicting sources and failed adapters. Keep tuning data separate from the final comparison so the reported gain is not simply a result of repeatedly testing the same examples.
A second agent is not automatically an independent reviewer. It may share a model, prompt template, retrieved source or erroneous draft with the first. Record what evidence it actually inspected. Our review role has the original fixture through its trusted adapter and performs an exact equality check; that check is transparent but far narrower than judging a general document. Do not label the result “human approved” or “independently verified”.
Try it yourself · Activity 05
20 minChallenge a protection without changing the original
Copy the three files to a disposable folder and replace calls >= budget with False.
- Run all fourteen tests in that copy.
- Locate the shared-budget failure and inspect the reported call count.
- Restore the original and rerun before using it as your reference.
Which observable result would distinguish an unknown-reference bug from a wrong-task bug?
Worked answer
The modified scheduler makes four calls in the budget-2 demo because it ignores the global stop condition. The budget test fails. Restoring the check returns the suite to fourteen passing tests. The local retry cap still limits attempts per task, which explains why two different limits must be tested separately.
Read a sample · Chapter 06 of 06
Hand over responsibilities, not just a transcript
Make unresolved conditions visible before introducing models or effects.
A useful handover records the task purpose, role contracts, graph, source snapshot, adapter identities, limits, source revision, test commands and final state semantics. Include the accepted results and controlled event sequence, but avoid turning every prompt and source document into routine logs. Full bodies can create another confidential-data store. This lab uses only fictional text, and its event records contain task IDs, attempt numbers and fixed codes.
Scroll sideways to see every column.
| Current evidence | What a real integration still needs |
|---|---|
| Fixed graph validated before calls | Review policy for any model-proposed plan or dynamic branch. |
| Task and attempt reply binding | Workflow-run identity, late-message handling and authenticated transport. |
| Reference IDs follow accepted inputs | Source access controls, content snapshots and claim-support evaluation. |
| Shared invocation count and retry cap | Token/cost limits, deadlines, cancellation and admission control. |
| Serial in-memory final states | Durable state transitions, restart reconciliation and concurrency design. |
| No external effects in workers | Idempotency or recovery policy for each consequential effect. |
Adding parallel workers is a design change, not a switch that this code exposes. Ready tasks would need admission limits; completed replies would need atomic state transitions; cancellation would have to account for still-running work. The graph alone does not define how simultaneous results, late responses or retries interact. Start by naming those states and tests instead of claiming that the serial examples establish concurrent correctness.
Adding a model adapter creates a different boundary. Its prompt must distinguish task instructions from untrusted evidence. Its response still passes through accept_reply, and a valid body still needs task-specific quality checks. Tool permissions must be enforced outside the generated text. An output saying “run the next task now” remains body text; the coordinator's plan and state own scheduling authority.
Failure reporting should help the operator choose a safe next action. A blocked task needs a successful predecessor or a revised plan; a budget-exhausted task needs a new deliberate allowance; a failed task needs diagnosis of its rejected replies or adapter errors. Blindly replaying the whole run can duplicate successful work in a real integration. This report has no restart command or durable ledger, so it makes no resumability claim.
Acceptance for this course is a tested local control-layer example. The code and printout are reproducible, and their limits are explicit. Acceptance for a production multi-agent service would require observed model behaviour, real identity and tool boundaries, operator controls, recovery tests, accessibility where humans interact, and an owner decision about unresolved conditions. The educational byline does not imply that Micky personally ran those future checks.
Try it yourself · Activity 06
15 minWrite a reviewable handover
Draft a short release note for the scripted Orchard workflow.
- Name two checked contracts and the tests that support them.
- Record the observed successful call count and the budget-2 outcome.
- Name three missing integration properties and responsible roles.
Can a new builder tell exactly which results were executed and which remain proposed work?
Worked answer
The reply-identity and shared-budget tests support task/attempt binding and run-wide invocation limits. The successful demo uses four calls; budget 2 retains only the research result, exhausts the draft budget and blocks review. Adapter owners must supply deadlines, the storage owner must design durable transitions, and the evaluation owner must measure factual quality against real sources. No model-quality, concurrency or production recovery claim is made.
Keep learning
The complete workbook
Six chapters and six worked activities teach a complete three-file standard-library Python workflow. Run a deliberately retried handoff, inspect failures and call limits, and separate structural acceptance from factual quality. The full implementation, tests, observed output and production limitations are included.
- 01Give each role a reason to existRead here · 1 exercise
Start with a task and a comparison, not a collection of personalities.
- 02Validate the plan before any worker runsRead here · 1 exercise
A dependency means accepted output is required, not merely that someone was asked.
- 03Make handoffs small and checkableRead here · 1 exercise
An accepted envelope establishes structure and lineage; it does not establish truth.
- 04Schedule retries within one shared budgetRead here · 1 exercise
Every attempt costs a call, including malformed replies and exceptions.
- 05Test both progress and refusalRead here · 1 exercise
A coordinator that never runs anything is bounded but not useful.
- 06Hand over responsibilities, not just a transcriptRead here · 1 exercise
Make unresolved conditions visible before introducing models or effects.
Also inside: a 8-point checklist, a glossary of 10 terms and 10 questions and answers to test yourself. 6 hands-on exercises, each with a worked answer at the back where the workbook gives one.
No login, no card, no account. Before the download we ask you to follow Mickai (two quick links). Free to download and use for personal learning, study groups and inside your own team. Please do not resell the workbooks or republish them as your own. Link people to trust-agent.ai instead.
Test yourself
Questions and answers
Does this lab run autonomous agents?
No. It runs scripted role adapters through a fixed serial graph. There are no model calls, dynamic plans or parallel workers.
When should I add another role?
When an explicit responsibility and measured benefit justify the extra calls, context transfer and failure modes compared with a simpler baseline.
Can a cyclic plan run its independent root first?
Not in this contract. The full topological order is materialised and a cycle is rejected before any worker call.
Why check unknown dependency IDs before graph sorting?
The graph library can otherwise introduce an implied predecessor. Every node here must have an explicit task, role and prompt.
What binds a reply to its caller?
The host-supplied task ID and exact integer attempt number must match the reply. This synchronous lab has no separate workflow-run ID.
Do known reference IDs prove an answer is true?
No. They establish allowed lineage only. A claim can cite an allowed source and still contradict it or lack support.
Do invalid replies consume the budget?
Yes. Every worker invocation counts before it starts, including calls that return malformed JSON or raise ordinary exceptions.
What happens with budget 2 in the demo?
Research succeeds and the first draft reply fails. The draft cannot retry, so it is budget_exhausted; the reviewer is blocked.
Does a call budget provide a timeout?
No. It limits invocation count, not duration. A worker that never returns can still hang the serial run.
Is replaying a failed run safe for external effects?
Not established. This lab has no external effects, durable ledger or effect-specific idempotency. A real adapter needs an ambiguous-outcome and recovery policy.
When you have finished
Get your certificate of completion
Type your name and download a certificate for this workbook as a PDF, ready to print or to add to LinkedIn. It is made on your own device, so your name is never sent to us. It is a self-declared certificate, not an accredited qualification.
Learn the language
Key terms
- Orchestration
- Assigning work, checking handoffs and controlling progress, resources and stopping conditions.
- Role contract
- The defined inputs, outputs, responsibilities and permitted capabilities of a worker role.
- Dependency
- A requirement for an accepted predecessor result before a task can run.
- Directed acyclic graph
- A directed dependency graph with no cycle returning to a task through its predecessors.
- Topological order
- An order in which each task appears after all of its predecessors.
- Handoff
- A bounded result passed from completed work to a later task under an explicit contract.
6 of the workbook's 10 terms. The complete glossary is in the workbook.
Follow the evidence
Sources and checks
Facts last checked: .
Examples in this workbook were run on: Python 3.12.10 on Windows; standard library only. Fourteen learner tests, five exact demo output lines and three detected logic faults. Fixed serial DAG and scripted workers, with no model or network calls. No autonomy, concurrency, timeout, factual-quality or production-recovery acceptance is claimed. (2026-09-27).
These workbooks use AI assistance. See how the workbooks are made.
- Building effective agentsAnthropic
- AI agent orchestration patternsMicrosoft Learn
- Topological sorting and cycle detectionPython Software Foundation
- Dataclasses and frozen instancesPython Software Foundation
- JSON decoding and parser hooksPython Software Foundation
- Unit testing and assertionsPython Software Foundation
Created by Mickarle Wagstaff-Irons - Micky Irons with the Mickai team. Published by Mickai LTD. Last updated 27 September 2026.
NextKeep going
Where to go next
Recommended for you
Testing agent loops: budgets, termination and failure recovery
Test an agent loop locally: enforce budgets, replay model decisions, reject late results and verify bounded retries without paying for model calls.
Recommended for you
Threat-model an agent and its tool boundary
Map an agent tool boundary and test a local gateway for identity, ownership, exact-action approval, expiry, replay and stale writes using fictional notes.