agents · Level 3
Testing agent loops: budgets, termination and failure recovery
Build a deterministic local harness and make its failure boundaries observable.

Start with the essentials
The short answer
Test the harness separately from the model. Replace model responses, tool calls and time with controlled fixtures, then check terminal states and forbidden side effects. This workbook builds a small Python loop with bounded steps, tool attempts and cooperative deadlines. Its tests prove those local rules, while separate integration and model evaluations are still needed for a real service.
What you will learn
- Separate harness correctness from model answer quality.
- Define terminal outcomes and count attempted work consistently.
- Inject a fake model, tool and clock into a small loop.
- Test that denied or invalid actions never reach a tool.
- Bound retries and recognise the limits of cooperative deadlines.
- Write an evidence record that states what the tests do not establish.
Who it is for
Learners who can read small Python functions and want to test an agent harness before connecting real services.
Before you start
- Understand the loop-and-harness course.
- Be able to create Python files, use functions and run a command in a terminal.
Read a sample · Chapter 02 of 06
Build the model-step and validation boundary
Create a new local folder for this lab and use Python 3.12 or later. The worked example uses only the standard library. Save each named file as UTF-8; copy the selectable web code if a PDF reader changes indentation.
Create harness.py. Put the following imports and support types at its top. A frozen Result keeps the top-level report fields fixed after return. Its trace is a tuple of event labels; it deliberately excludes raw prompts, tool inputs and exception messages. This is useful minimal evidence, not a durable audit system.
from dataclasses import dataclass
from math import isfinite
from time import monotonic
class TemporaryToolError(Exception):
pass
@dataclass(frozen=True)
class Result:
status: str
steps: int
tool_calls: int
trace: tuple
answer: str = ""
def positive_int(value):
return type(value) is int and value > 0Append the next block to harness.py. It starts run(), validates configuration and implements the model half of the loop. Do not execute the file until you have appended chapter 3, which completes the indented while body. Python treats True as an integer in many contexts; the exact type check deliberately rejects boolean values used accidentally as budgets.
def run(model, tool, *, max_steps=3, max_tool_calls=2,
seconds=1.0, clock=monotonic):
if not positive_int(max_steps):
raise ValueError("max_steps must be a positive integer")
if not positive_int(max_tool_calls):
raise ValueError("max_tool_calls must be positive")
if (type(seconds) not in (int, float)
or not isfinite(seconds) or seconds <= 0):
raise ValueError("seconds must be positive and finite")
deadline = clock() + seconds
history, trace = [], []
steps = calls = 0
def expired():
return clock() >= deadline
def finish(status, answer=""):
return Result(status, steps, calls, tuple(trace), answer)
while steps < max_steps:
if expired():
return finish("deadline")
steps += 1 # An attempted call consumes a step.
trace.append("model")
try:
action = model(tuple(history))
except Exception:
return finish("model_error")
if expired():
return finish("deadline")
if not isinstance(action, dict):
return finish("invalid_action")
kind = action.get("kind")
if kind == "finish":
answer = action.get("answer")
if type(answer) is not str or not answer.strip():
return finish("invalid_action")
if len(answer) > 200:
return finish("invalid_action")
return finish("finished", answer)
if kind != "lookup":
return finish("invalid_action")
key = action.get("key")
if type(key) is not str or not 1 <= len(key) <= 80:
return finish("invalid_action")The model receives an immutable tuple of previous tool-result strings. Passing dependencies as arguments is dependency injection: the caller can supply controlled substitutes without rewriting the loop. The clock defaults to time.monotonic for elapsed-time comparisons. Its absolute value has no calendar meaning, and only differences are relevant here. In tests we supply a clock whose value changes only when the test says so.
The action contract is intentionally narrow. Unknown kinds, blank final answers, long answers and invalid lookup keys return invalid_action. Extra dictionary fields are ignored and never forwarded. This is not a general schema validator. A production adapter should validate the full protocol, permissions and resource limits before exposing powerful tools. A valid final answer can still be nonsense: structural validity and answer quality are different questions.
Try it yourself · Activity 02
15 minChoose a terminal status
Read the model half of the loop and predict the result for three first responses.
- Use a plain string instead of a dictionary.
- Use a finish action whose answer contains only spaces.
- Use a lookup action whose key is 81 characters long.
Would returning an error string as a successful answer hide the failure?
Worked answer
All three responses produce invalid_action after one model attempt and before any tool attempt. They fail different validation checks but share the same external status. Returning a normal-looking answer containing an error message would make it harder for the caller to distinguish failure from completion.
Read a sample · Chapter 03 of 06
Add bounded retries and cooperative deadlines
Append this block directly after chapter 2 in harness.py. Preserve its leading spaces: the for loop belongs inside the while loop, and the final return belongs outside it but inside run().
# Two attempts at most, including the first attempt.
for attempt in range(2):
if expired():
return finish("deadline")
if calls >= max_tool_calls:
return finish("tool_limit")
calls += 1
trace.append("lookup")
try:
value = tool(key)
except TemporaryToolError:
trace.append("temporary_error")
if expired():
return finish("deadline")
if attempt == 1:
return finish("tool_error")
continue
except Exception:
return finish("tool_error")
if expired():
return finish("deadline")
if type(value) is not str or len(value) > 200:
return finish("invalid_result")
history.append(value)
break
return finish("step_limit")Only TemporaryToolError is retried, and only once. The original attempt and its retry both consume the tool budget. Other exceptions stop immediately. This is an explicit policy for a fictional read-only lookup. It is not permission to retry payments, deletions or other actions with consequences. If a service times out after performing an action, repeating it may duplicate the action. Such operations need their own idempotency and reconciliation design.
The word cooperative matters. If a synchronous tool never returns, this function cannot regain control to check its clock. A fake deadline test proves rejection of a late result after the callback returns; it does not prove pre-emption. Real adapters need transport timeouts and, where necessary, cancellable work or an isolated worker that can be stopped. Even cancellation does not imply that a remote action was rolled back.
Failure status precedence is also part of this example. A normal model or tool return is followed by a deadline check. A generic dependency exception reports model_error or tool_error; a temporary tool exception checks the deadline before attempting recovery. Document that policy rather than promising that every run which used too much wall-clock time always reports deadline. The trace helps explain the path that actually occurred.
Count budgets are not a money meter. A model request may consume different token counts and prices from another request. This lab enforces attempts only. A service that promises a spending bound must reserve and reconcile provider usage, account for concurrent requests and use enforceable service limits. Logging an estimated cost after a response cannot prevent that response from already incurring a charge.
Try it yourself · Activity 03
15 minTrace the retry budget
Assume the first lookup raises TemporaryToolError and a second lookup would succeed.
- Set max_tool_calls to 1 and predict the result.
- Set it to 2 and predict the calls before a later finish action.
- Describe why this policy is limited to a read-only fictional tool.
What if the tool performed a change before its connection failed?
Worked answer
With one tool allowance, the retry is refused and the result is tool_limit with one tool attempt. With two allowances, both attempts are counted and the model may then finish if a model step and time remain. For a side-effecting tool, the caller might not know whether the first attempt committed, so retrying blindly could duplicate work.
Read a sample · Chapter 04 of 06
Replay a model and take control of time
A scripted model returns a known sequence of actions. That lets you test the harness without depending on sampling, network availability or a provider bill. Save the next block in fixtures.py beside harness.py.
from harness import run, TemporaryToolError
class ScriptedModel:
def __init__(self, actions):
self.actions = iter(actions)
self.seen = []
def __call__(self, history):
self.seen.append(history)
return next(self.actions)
class FakeClock:
def __init__(self):
self.now = 0.0
def __call__(self):
return self.now
def lookup(key):
return {"opening": "09:00"}[key]
if __name__ == "__main__":
model = ScriptedModel([
{"kind": "lookup", "key": "opening"},
{"kind": "finish", "answer": "Opens at 09:00."},
])
result = run(model, lookup, clock=FakeClock())
print(result.status, result.steps, result.tool_calls)
print(result.answer)
print(result.trace)Run python fixtures.py in the lab folder. You should see the output below. The trace describes two model calls with one lookup between them. The answer is fictional fixture text, not a statement about a real organisation. A scripted sequence that is exhausted raises StopIteration; the harness reports model_error rather than silently inventing another response.
finished 2 1
Opens at 09:00.
('model', 'lookup', 'model')FakeClock begins at zero and stays there until a test changes now. It does not sleep. Advancing it inside a tool simulates elapsed time between a call and its return. Keep model decisions, data fixtures and time independent: a failure should tell you whether a policy boundary broke, not whether the computer happened to run slowly that morning.
A replay is only as informative as its cases. Replaying one success trace cannot establish robustness against malformed requests, repeated actions or a different protocol. For a real integration, create a sanitised fixture from the documented adapter boundary. Retain event order, schema version and relevant outcomes. Exclude personal data and credentials; use invented records wherever possible. Replaying internal model reasoning is unnecessary for testing this contract.
Try it yourself · Activity 04
15 minMake the success path late
Create a copy of the example that advances FakeClock.now inside lookup before returning.
- Keep seconds at 1.0.
- Advance the clock by exactly 1.0 in the tool.
- Check that the final model action was never requested.
Which part of this experiment is simulated and which part is executed?
Worked answer
The Python control flow and status handling execute normally. Time is simulated. The tool returns, the post-call check observes equality with the deadline, and the run returns deadline with one model step, one tool call and an empty answer. The scripted finish action remains unused. This demonstrates late-result rejection, not cancellation of a hung tool.
Read a sample · Chapter 05 of 06
Test the absence of forbidden work
Save the next three code blocks, in order, in test_harness.py. The later blocks continue the LoopTests class before the final runner. Run python -m unittest -v test_harness from the same folder.
import unittest
from unittest.mock import Mock
from harness import run, TemporaryToolError
from fixtures import ScriptedModel, FakeClock, lookup
class LoopTests(unittest.TestCase):
def test_normal_completion(self):
model = ScriptedModel([
{"kind": "lookup", "key": "opening"},
{"kind": "finish", "answer": "09:00"},
])
tool = Mock(side_effect=lookup)
r = run(model, tool, clock=FakeClock())
self.assertEqual((r.status, r.steps, r.tool_calls),
("finished", 2, 1))
self.assertEqual(model.seen, [(), ("09:00",)])
tool.assert_called_once_with("opening")
def test_repetition_stops_at_step_limit(self):
model = Mock(return_value={"kind": "lookup", "key": "x"})
tool = Mock(return_value="same answer")
r = run(model, tool, max_steps=3, max_tool_calls=9,
clock=FakeClock())
self.assertEqual(r.status, "step_limit")
self.assertEqual(model.call_count, 3)
self.assertEqual(tool.call_count, 3)Continue the class with these two methods. Keep the four-space indentation before each def.
def test_invalid_action_never_reaches_tool(self):
tool = Mock()
r = run(lambda _: {"kind": "erase"}, tool,
clock=FakeClock())
self.assertEqual(r.status, "invalid_action")
tool.assert_not_called()
def test_temporary_failure_retries_once(self):
model = ScriptedModel([
{"kind": "lookup", "key": "opening"},
{"kind": "finish", "answer": "09:00"},
])
tool = Mock(side_effect=[TemporaryToolError(), "09:00"])
r = run(model, tool, clock=FakeClock())
self.assertEqual(r.status, "finished")
self.assertEqual(r.tool_calls, 2)
self.assertEqual(tool.call_count, 2)Append the remaining methods and the final test runner below.
def test_retry_cannot_exceed_tool_budget(self):
model = lambda _: {"kind": "lookup", "key": "opening"}
tool = Mock(side_effect=TemporaryToolError())
r = run(model, tool, max_tool_calls=1, clock=FakeClock())
self.assertEqual(r.status, "tool_limit")
self.assertEqual(tool.call_count, 1)
def test_result_after_deadline_is_rejected(self):
clock = FakeClock()
def slow_tool(key):
clock.now += 1.0
return "09:00"
model = Mock(return_value={"kind": "lookup", "key": "x"})
r = run(model, slow_tool, seconds=1.0, clock=clock)
self.assertEqual(r.status, "deadline")
self.assertEqual(model.call_count, 1)
self.assertEqual(r.tool_calls, 1)
self.assertEqual(r.answer, "")
if __name__ == "__main__":
unittest.main()The six tests cover success, repeated model decisions, an invalid action, one successful recovery, an exhausted retry allowance and a deadline boundary. Mock records calls and can return a sequence of values or raise supplied exceptions through side_effect. Use assertions about the absence of calls as well as assertions about the final status. A loop that returns tool_limit after making an extra tool call is still wrong.
A passing suite should be challenged. In a disposable copy of harness.py, change the tool budget comparison from >= to >. The one-attempt retry test should fail because an extra call is now allowed. Restore the correct source and rerun. This small mutation check asks whether the test can detect the particular defect it claims to cover. It does not mean the suite detects every possible defect.
Avoid asserting incidental details unless they form part of your contract. A test can care that an invalid action never reaches a tool without caring whether validation lives in one function or five. Exact trace order is helpful for this tiny example, but a real concurrent harness might promise only a partial order. Write the requirement first, then decide what evidence establishes it.
Try it yourself · Activity 05
15 minAdd a non-retryable failure test
Extend LoopTests so the tool raises ValueError on its first call.
- Choose a model that requests a lookup.
- Assert tool_error.
- Assert that the tool was called exactly once and no answer was returned.
Would a test that checked only the status detect an unwanted retry?
Worked answer
Use a Mock whose side_effect is ValueError("fixture failure"), run one lookup, then assert status equals tool_error, tool.call_count equals 1 and answer equals an empty string. Status alone would miss a harness that retried several times before eventually returning the same status.
Keep learning
The complete workbook
Build a small, inspectable agent loop and a repeatable test suite using only the Python standard library. Learn to test repeated decisions, invalid actions, exhausted budgets, late results and temporary failures. The lab makes no model requests and performs no external actions.
- 01Specify the boundary before testing the loopIn the workbook · 1 exercise
A model can propose an action. The harness decides whether that action may run, records the outcome and eventually stops. Test that machinery with predictable inputs before asking whether the model gives useful answers.
- 02Build the model-step and validation boundaryRead here · 1 exercise
Create a new local folder for this lab and use Python 3.12 or later. The worked example uses only the standard library. Save each named file as UTF-8; copy the selectable web code if a PDF reader changes indentation.
- 03Add bounded retries and cooperative deadlinesRead here · 1 exercise
Append this block directly after chapter 2 in harness.py. Preserve its leading spaces: the for loop belongs inside the while loop, and the final return belongs outside it but inside run().
- 04Replay a model and take control of timeRead here · 1 exercise
A scripted model returns a known sequence of actions. That lets you test the harness without depending on sampling, network availability or a provider bill. Save the next block in fixtures.py beside harness.py.
- 05Test the absence of forbidden workRead here · 1 exercise
Save the next three code blocks, in order, in test_harness.py. The later blocks continue the LoopTests class before the final runner. Run python -m unittest -v test_harness from the same folder.
- 06Turn a local result into an honest release decisionIn the workbook · 1 exercise
You now have a small contract, an executable harness and tests that exercise important boundaries. The next step is to describe the scope precisely and decide what still needs evidence before connecting real systems.
Also inside: a 8-point checklist, a glossary of 10 terms and 10 questions and answers to test yourself. 6 hands-on exercises, each with a worked answer at the back where the workbook gives one.
No login, no card, no account. Before the download we ask you to follow Mickai (two quick links). Free to download and use for personal learning, study groups and inside your own team. Please do not resell the workbooks or republish them as your own. Link people to trust-agent.ai instead.
Test yourself
Questions and answers
What does a finished status establish?
A structurally valid finish action arrived within the checked deadline. It does not establish that the answer is factually correct or useful.
Does a failed model call consume a step?
Yes. This lab counts attempted model calls, so an exception still consumes one step.
Why supply the model and clock as arguments?
Tests can substitute deterministic behaviour without changing the loop implementation. That isolates harness rules from network, model and timing variability.
Why reject a boolean used as a budget?
The contract requires a positive integer count, not a truth value. Exact type checking prevents True from silently becoming a one-step budget.
Can the deadline interrupt a tool that never returns?
No. This synchronous example only checks time when it regains control. Real adapters need appropriate timeouts and cancellation or isolation mechanisms.
Why do failed tool attempts count towards the allowance?
They still use resources. Counting only successful calls would let repeated failures bypass the intended bound.
What does a fake clock test prove?
It tests the control-flow response to chosen time values. It does not measure real latency, CPU performance or network cancellation.
Why assert that a tool was not called?
A correct final status can hide an earlier forbidden action. Call assertions check the side-effect boundary directly.
What is the point of deliberately changing >= to >?
It introduces a known extra-attempt defect. A failing boundary test provides evidence that this test detects that specific defect.
What remains after the local tests pass?
Adapter integration, real model evaluation, permissions, cancellation, concurrency and durable-state behaviour still need their own evidence before a real service is released.
When you have finished
Get your certificate of completion
Type your name and download a certificate for this workbook as a PDF, ready to print or to add to LinkedIn. It is made on your own device, so your name is never sent to us. It is a self-declared certificate, not an accredited qualification.
Learn the language
Key terms
- Harness
- The code that validates actions, manages state and budgets, invokes tools and stops a run.
- Fixture
- Controlled data and setup used to repeat a test.
- Dependency injection
- Supplying a dependency as an argument so a caller can choose a real implementation or a test substitute.
- Mock
- A test object that can record calls and return controlled values or raise exceptions.
- Terminal status
- An explicit outcome that says why a run stopped.
- Cooperative deadline
- A time boundary enforced only when control returns to code that checks it.
6 of the workbook's 10 terms. The complete glossary is in the workbook.
Follow the evidence
Sources and checks
Facts last checked: .
Examples in this workbook were run on: Python 3.12.10 on Windows 11, standard library only. Six learner tests, 68 additional course checks and a deliberate budget mutation; no model or network calls. (2026-09-27).
These workbooks use AI assistance. See how the workbooks are made.
- unittest: test cases, assertions and running testsPython Software Foundation
- unittest.mock: side effects and call assertionsPython Software Foundation
- time.monotonic: elapsed-time measurementPython Software Foundation
Created by Mickarle Wagstaff-Irons - Micky Irons with the Mickai team. Published by Mickai LTD. Last updated 27 September 2026.
NextKeep going
Where to go next
Recommended for you
Tool calling and structured outputs: validate every response
Treat every model reply as untrusted input: write JSON Schema contracts, parse and validate each tool call, retry with feedback, and run tools with limits, dry runs and logs.
Recommended for you
Design a repeatable evaluation suite for an AI application
Build an offline AI evaluation harness with test cases, a rubric, regression comparisons, uncertainty checks and a reproducible release record.