agents · Level 3

Testing agent loops: budgets, termination and failure recovery

Build a deterministic local harness and make its failure boundaries observable.

By Mickarle Wagstaff-Irons - Micky Irons

  • Level 3Building
  • 150 min
  • 6 chapters
  • Free PDF, no account
The Loopagents / 03

Start with the essentials

The short answer

Test the harness separately from the model. Replace model responses, tool calls and time with controlled fixtures, then check terminal states and forbidden side effects. This workbook builds a small Python loop with bounded steps, tool attempts and cooperative deadlines. Its tests prove those local rules, while separate integration and model evaluations are still needed for a real service.

What you will learn

  • Separate harness correctness from model answer quality.
  • Define terminal outcomes and count attempted work consistently.
  • Inject a fake model, tool and clock into a small loop.
  • Test that denied or invalid actions never reach a tool.
  • Bound retries and recognise the limits of cooperative deadlines.
  • Write an evidence record that states what the tests do not establish.

Who it is for

Learners who can read small Python functions and want to test an agent harness before connecting real services.

Before you start

  • Understand the loop-and-harness course.
  • Be able to create Python files, use functions and run a command in a terminal.

Read a sample · Chapter 02 of 06

02

Build the model-step and validation boundary

Create a new local folder for this lab and use Python 3.12 or later. The worked example uses only the standard library. Save each named file as UTF-8; copy the selectable web code if a PDF reader changes indentation.

Create harness.py. Put the following imports and support types at its top. A frozen Result keeps the top-level report fields fixed after return. Its trace is a tuple of event labels; it deliberately excludes raw prompts, tool inputs and exception messages. This is useful minimal evidence, not a durable audit system.

python · 20 lines
from dataclasses import dataclass
from math import isfinite
from time import monotonic


class TemporaryToolError(Exception):
    pass


@dataclass(frozen=True)
class Result:
    status: str
    steps: int
    tool_calls: int
    trace: tuple
    answer: str = ""


def positive_int(value):
    return type(value) is int and value > 0

Append the next block to harness.py. It starts run(), validates configuration and implements the model half of the loop. Do not execute the file until you have appended chapter 3, which completes the indented while body. Python treats True as an integer in many contexts; the exact type check deliberately rejects boolean values used accidentally as budgets.

python · 46 lines
def run(model, tool, *, max_steps=3, max_tool_calls=2,
        seconds=1.0, clock=monotonic):
    if not positive_int(max_steps):
        raise ValueError("max_steps must be a positive integer")
    if not positive_int(max_tool_calls):
        raise ValueError("max_tool_calls must be positive")
    if (type(seconds) not in (int, float)
            or not isfinite(seconds) or seconds <= 0):
        raise ValueError("seconds must be positive and finite")

    deadline = clock() + seconds
    history, trace = [], []
    steps = calls = 0

    def expired():
        return clock() >= deadline

    def finish(status, answer=""):
        return Result(status, steps, calls, tuple(trace), answer)

    while steps < max_steps:
        if expired():
            return finish("deadline")
        steps += 1  # An attempted call consumes a step.
        trace.append("model")
        try:
            action = model(tuple(history))
        except Exception:
            return finish("model_error")
        if expired():
            return finish("deadline")
        if not isinstance(action, dict):
            return finish("invalid_action")
        kind = action.get("kind")
        if kind == "finish":
            answer = action.get("answer")
            if type(answer) is not str or not answer.strip():
                return finish("invalid_action")
            if len(answer) > 200:
                return finish("invalid_action")
            return finish("finished", answer)
        if kind != "lookup":
            return finish("invalid_action")
        key = action.get("key")
        if type(key) is not str or not 1 <= len(key) <= 80:
            return finish("invalid_action")

The model receives an immutable tuple of previous tool-result strings. Passing dependencies as arguments is dependency injection: the caller can supply controlled substitutes without rewriting the loop. The clock defaults to time.monotonic for elapsed-time comparisons. Its absolute value has no calendar meaning, and only differences are relevant here. In tests we supply a clock whose value changes only when the test says so.

The action contract is intentionally narrow. Unknown kinds, blank final answers, long answers and invalid lookup keys return invalid_action. Extra dictionary fields are ignored and never forwarded. This is not a general schema validator. A production adapter should validate the full protocol, permissions and resource limits before exposing powerful tools. A valid final answer can still be nonsense: structural validity and answer quality are different questions.

Try it yourself · Activity 02

15 min

Choose a terminal status

Read the model half of the loop and predict the result for three first responses.

  1. Use a plain string instead of a dictionary.
  2. Use a finish action whose answer contains only spaces.
  3. Use a lookup action whose key is 81 characters long.

Would returning an error string as a successful answer hide the failure?

Worked answer

All three responses produce invalid_action after one model attempt and before any tool attempt. They fail different validation checks but share the same external status. Returning a normal-looking answer containing an error message would make it harder for the caller to distinguish failure from completion.

Read a sample · Chapter 03 of 06

03

Add bounded retries and cooperative deadlines

Append this block directly after chapter 2 in harness.py. Preserve its leading spaces: the for loop belongs inside the while loop, and the final return belongs outside it but inside run().

python · 26 lines
        # Two attempts at most, including the first attempt.
        for attempt in range(2):
            if expired():
                return finish("deadline")
            if calls >= max_tool_calls:
                return finish("tool_limit")
            calls += 1
            trace.append("lookup")
            try:
                value = tool(key)
            except TemporaryToolError:
                trace.append("temporary_error")
                if expired():
                    return finish("deadline")
                if attempt == 1:
                    return finish("tool_error")
                continue
            except Exception:
                return finish("tool_error")
            if expired():
                return finish("deadline")
            if type(value) is not str or len(value) > 200:
                return finish("invalid_result")
            history.append(value)
            break
    return finish("step_limit")

Only TemporaryToolError is retried, and only once. The original attempt and its retry both consume the tool budget. Other exceptions stop immediately. This is an explicit policy for a fictional read-only lookup. It is not permission to retry payments, deletions or other actions with consequences. If a service times out after performing an action, repeating it may duplicate the action. Such operations need their own idempotency and reconciliation design.

The word cooperative matters. If a synchronous tool never returns, this function cannot regain control to check its clock. A fake deadline test proves rejection of a late result after the callback returns; it does not prove pre-emption. Real adapters need transport timeouts and, where necessary, cancellable work or an isolated worker that can be stopped. Even cancellation does not imply that a remote action was rolled back.

Failure status precedence is also part of this example. A normal model or tool return is followed by a deadline check. A generic dependency exception reports model_error or tool_error; a temporary tool exception checks the deadline before attempting recovery. Document that policy rather than promising that every run which used too much wall-clock time always reports deadline. The trace helps explain the path that actually occurred.

Count budgets are not a money meter. A model request may consume different token counts and prices from another request. This lab enforces attempts only. A service that promises a spending bound must reserve and reconcile provider usage, account for concurrent requests and use enforceable service limits. Logging an estimated cost after a response cannot prevent that response from already incurring a charge.

Try it yourself · Activity 03

15 min

Trace the retry budget

Assume the first lookup raises TemporaryToolError and a second lookup would succeed.

  1. Set max_tool_calls to 1 and predict the result.
  2. Set it to 2 and predict the calls before a later finish action.
  3. Describe why this policy is limited to a read-only fictional tool.

What if the tool performed a change before its connection failed?

Worked answer

With one tool allowance, the retry is refused and the result is tool_limit with one tool attempt. With two allowances, both attempts are counted and the model may then finish if a model step and time remain. For a side-effecting tool, the caller might not know whether the first attempt committed, so retrying blindly could duplicate work.

Read a sample · Chapter 04 of 06

04

Replay a model and take control of time

A scripted model returns a known sequence of actions. That lets you test the harness without depending on sampling, network availability or a provider bill. Save the next block in fixtures.py beside harness.py.

python · 34 lines
from harness import run, TemporaryToolError


class ScriptedModel:
    def __init__(self, actions):
        self.actions = iter(actions)
        self.seen = []

    def __call__(self, history):
        self.seen.append(history)
        return next(self.actions)


class FakeClock:
    def __init__(self):
        self.now = 0.0

    def __call__(self):
        return self.now


def lookup(key):
    return {"opening": "09:00"}[key]


if __name__ == "__main__":
    model = ScriptedModel([
        {"kind": "lookup", "key": "opening"},
        {"kind": "finish", "answer": "Opens at 09:00."},
    ])
    result = run(model, lookup, clock=FakeClock())
    print(result.status, result.steps, result.tool_calls)
    print(result.answer)
    print(result.trace)

Run python fixtures.py in the lab folder. You should see the output below. The trace describes two model calls with one lookup between them. The answer is fictional fixture text, not a statement about a real organisation. A scripted sequence that is exhausted raises StopIteration; the harness reports model_error rather than silently inventing another response.

text · 3 lines
finished 2 1
Opens at 09:00.
('model', 'lookup', 'model')

FakeClock begins at zero and stays there until a test changes now. It does not sleep. Advancing it inside a tool simulates elapsed time between a call and its return. Keep model decisions, data fixtures and time independent: a failure should tell you whether a policy boundary broke, not whether the computer happened to run slowly that morning.

A replay is only as informative as its cases. Replaying one success trace cannot establish robustness against malformed requests, repeated actions or a different protocol. For a real integration, create a sanitised fixture from the documented adapter boundary. Retain event order, schema version and relevant outcomes. Exclude personal data and credentials; use invented records wherever possible. Replaying internal model reasoning is unnecessary for testing this contract.

Try it yourself · Activity 04

15 min

Make the success path late

Create a copy of the example that advances FakeClock.now inside lookup before returning.

  1. Keep seconds at 1.0.
  2. Advance the clock by exactly 1.0 in the tool.
  3. Check that the final model action was never requested.

Which part of this experiment is simulated and which part is executed?

Worked answer

The Python control flow and status handling execute normally. Time is simulated. The tool returns, the post-call check observes equality with the deadline, and the run returns deadline with one model step, one tool call and an empty answer. The scripted finish action remains unused. This demonstrates late-result rejection, not cancellation of a hung tool.

Read a sample · Chapter 05 of 06

05

Test the absence of forbidden work

Save the next three code blocks, in order, in test_harness.py. The later blocks continue the LoopTests class before the final runner. Run python -m unittest -v test_harness from the same folder.

python · 27 lines
import unittest
from unittest.mock import Mock
from harness import run, TemporaryToolError
from fixtures import ScriptedModel, FakeClock, lookup


class LoopTests(unittest.TestCase):
    def test_normal_completion(self):
        model = ScriptedModel([
            {"kind": "lookup", "key": "opening"},
            {"kind": "finish", "answer": "09:00"},
        ])
        tool = Mock(side_effect=lookup)
        r = run(model, tool, clock=FakeClock())
        self.assertEqual((r.status, r.steps, r.tool_calls),
                         ("finished", 2, 1))
        self.assertEqual(model.seen, [(), ("09:00",)])
        tool.assert_called_once_with("opening")

    def test_repetition_stops_at_step_limit(self):
        model = Mock(return_value={"kind": "lookup", "key": "x"})
        tool = Mock(return_value="same answer")
        r = run(model, tool, max_steps=3, max_tool_calls=9,
                clock=FakeClock())
        self.assertEqual(r.status, "step_limit")
        self.assertEqual(model.call_count, 3)
        self.assertEqual(tool.call_count, 3)

Continue the class with these two methods. Keep the four-space indentation before each def.

python · 17 lines
    def test_invalid_action_never_reaches_tool(self):
        tool = Mock()
        r = run(lambda _: {"kind": "erase"}, tool,
                clock=FakeClock())
        self.assertEqual(r.status, "invalid_action")
        tool.assert_not_called()

    def test_temporary_failure_retries_once(self):
        model = ScriptedModel([
            {"kind": "lookup", "key": "opening"},
            {"kind": "finish", "answer": "09:00"},
        ])
        tool = Mock(side_effect=[TemporaryToolError(), "09:00"])
        r = run(model, tool, clock=FakeClock())
        self.assertEqual(r.status, "finished")
        self.assertEqual(r.tool_calls, 2)
        self.assertEqual(tool.call_count, 2)

Append the remaining methods and the final test runner below.

python · 24 lines
    def test_retry_cannot_exceed_tool_budget(self):
        model = lambda _: {"kind": "lookup", "key": "opening"}
        tool = Mock(side_effect=TemporaryToolError())
        r = run(model, tool, max_tool_calls=1, clock=FakeClock())
        self.assertEqual(r.status, "tool_limit")
        self.assertEqual(tool.call_count, 1)

    def test_result_after_deadline_is_rejected(self):
        clock = FakeClock()

        def slow_tool(key):
            clock.now += 1.0
            return "09:00"

        model = Mock(return_value={"kind": "lookup", "key": "x"})
        r = run(model, slow_tool, seconds=1.0, clock=clock)
        self.assertEqual(r.status, "deadline")
        self.assertEqual(model.call_count, 1)
        self.assertEqual(r.tool_calls, 1)
        self.assertEqual(r.answer, "")


if __name__ == "__main__":
    unittest.main()

The six tests cover success, repeated model decisions, an invalid action, one successful recovery, an exhausted retry allowance and a deadline boundary. Mock records calls and can return a sequence of values or raise supplied exceptions through side_effect. Use assertions about the absence of calls as well as assertions about the final status. A loop that returns tool_limit after making an extra tool call is still wrong.

A passing suite should be challenged. In a disposable copy of harness.py, change the tool budget comparison from >= to >. The one-attempt retry test should fail because an extra call is now allowed. Restore the correct source and rerun. This small mutation check asks whether the test can detect the particular defect it claims to cover. It does not mean the suite detects every possible defect.

Avoid asserting incidental details unless they form part of your contract. A test can care that an invalid action never reaches a tool without caring whether validation lives in one function or five. Exact trace order is helpful for this tiny example, but a real concurrent harness might promise only a partial order. Write the requirement first, then decide what evidence establishes it.

Try it yourself · Activity 05

15 min

Add a non-retryable failure test

Extend LoopTests so the tool raises ValueError on its first call.

  1. Choose a model that requests a lookup.
  2. Assert tool_error.
  3. Assert that the tool was called exactly once and no answer was returned.

Would a test that checked only the status detect an unwanted retry?

Worked answer

Use a Mock whose side_effect is ValueError("fixture failure"), run one lookup, then assert status equals tool_error, tool.call_count equals 1 and answer equals an empty string. Status alone would miss a harness that retried several times before eventually returning the same status.

Keep learning

The complete workbook

Build a small, inspectable agent loop and a repeatable test suite using only the Python standard library. Learn to test repeated decisions, invalid actions, exhausted budgets, late results and temporary failures. The lab makes no model requests and performs no external actions.

  1. 01
    Specify the boundary before testing the loop

    A model can propose an action. The harness decides whether that action may run, records the outcome and eventually stops. Test that machinery with predictable inputs before asking whether the model gives useful answers.

    In the workbook · 1 exercise
  2. 02
    Build the model-step and validation boundary

    Create a new local folder for this lab and use Python 3.12 or later. The worked example uses only the standard library. Save each named file as UTF-8; copy the selectable web code if a PDF reader changes indentation.

    Read here · 1 exercise
  3. 03
    Add bounded retries and cooperative deadlines

    Append this block directly after chapter 2 in harness.py. Preserve its leading spaces: the for loop belongs inside the while loop, and the final return belongs outside it but inside run().

    Read here · 1 exercise
  4. 04
    Replay a model and take control of time

    A scripted model returns a known sequence of actions. That lets you test the harness without depending on sampling, network availability or a provider bill. Save the next block in fixtures.py beside harness.py.

    Read here · 1 exercise
  5. 05
    Test the absence of forbidden work

    Save the next three code blocks, in order, in test_harness.py. The later blocks continue the LoopTests class before the final runner. Run python -m unittest -v test_harness from the same folder.

    Read here · 1 exercise
  6. 06
    Turn a local result into an honest release decision

    You now have a small contract, an executable harness and tests that exercise important boundaries. The next step is to describe the scope precisely and decide what still needs evidence before connecting real systems.

    In the workbook · 1 exercise

Also inside: a 8-point checklist, a glossary of 10 terms and 10 questions and answers to test yourself. 6 hands-on exercises, each with a worked answer at the back where the workbook gives one.

No login, no card, no account. Before the download we ask you to follow Mickai (two quick links). Free to download and use for personal learning, study groups and inside your own team. Please do not resell the workbooks or republish them as your own. Link people to trust-agent.ai instead.

Test yourself

Questions and answers

What does a finished status establish?

A structurally valid finish action arrived within the checked deadline. It does not establish that the answer is factually correct or useful.

Does a failed model call consume a step?

Yes. This lab counts attempted model calls, so an exception still consumes one step.

Why supply the model and clock as arguments?

Tests can substitute deterministic behaviour without changing the loop implementation. That isolates harness rules from network, model and timing variability.

Why reject a boolean used as a budget?

The contract requires a positive integer count, not a truth value. Exact type checking prevents True from silently becoming a one-step budget.

Can the deadline interrupt a tool that never returns?

No. This synchronous example only checks time when it regains control. Real adapters need appropriate timeouts and cancellation or isolation mechanisms.

Why do failed tool attempts count towards the allowance?

They still use resources. Counting only successful calls would let repeated failures bypass the intended bound.

What does a fake clock test prove?

It tests the control-flow response to chosen time values. It does not measure real latency, CPU performance or network cancellation.

Why assert that a tool was not called?

A correct final status can hide an earlier forbidden action. Call assertions check the side-effect boundary directly.

What is the point of deliberately changing >= to >?

It introduces a known extra-attempt defect. A failing boundary test provides evidence that this test detects that specific defect.

What remains after the local tests pass?

Adapter integration, real model evaluation, permissions, cancellation, concurrency and durable-state behaviour still need their own evidence before a real service is released.

When you have finished

Get your certificate of completion

Type your name and download a certificate for this workbook as a PDF, ready to print or to add to LinkedIn. It is made on your own device, so your name is never sent to us. It is a self-declared certificate, not an accredited qualification.

Learn the language

Key terms

Harness
The code that validates actions, manages state and budgets, invokes tools and stops a run.
Fixture
Controlled data and setup used to repeat a test.
Dependency injection
Supplying a dependency as an argument so a caller can choose a real implementation or a test substitute.
Mock
A test object that can record calls and return controlled values or raise exceptions.
Terminal status
An explicit outcome that says why a run stopped.
Cooperative deadline
A time boundary enforced only when control returns to code that checks it.

6 of the workbook's 10 terms. The complete glossary is in the workbook.

Follow the evidence

Sources and checks

Facts last checked: .

Examples in this workbook were run on: Python 3.12.10 on Windows 11, standard library only. Six learner tests, 68 additional course checks and a deliberate budget mutation; no model or network calls. (2026-09-27).

These workbooks use AI assistance. See how the workbooks are made.

  1. unittest: test cases, assertions and running testsPython Software Foundation
  2. unittest.mock: side effects and call assertionsPython Software Foundation
  3. time.monotonic: elapsed-time measurementPython Software Foundation

Created by Mickarle Wagstaff-Irons - Micky Irons with the Mickai team. Published by Mickai LTD. Last updated 27 September 2026.

NextKeep going

Where to go next