agents · Level 3

The loop harness: how AI agents actually run

The scaffolding that turns one model call into a working agent: context, tools, permissions, budgets, stop conditions, logs and tests.

By Mickarle Wagstaff-Irons - Micky Irons

  • Level 3Building
  • 90 min
  • 6 chapters
  • Free PDF, no account
The Loopagents / 03

Start with the essentials

The short answer

A loop harness, also called an agent harness, is the program that runs a language model in a loop. On each turn it assembles the context, calls the model, parses any tool request, runs the tool if the rules allow, feeds the result back and repeats until a stop condition is met. The harness, not the model, enforces permissions, budgets, logging and safety.

What you will learn

  • You will be able to explain what a loop harness is and how it differs from the model inside it.
  • You will be able to read a minimal agent loop and say what each part does.
  • You will know how to set permissions, sandboxes, budgets and stop conditions for an agent.
  • You will understand why logging, tamper-evident audit trails and testing matter, and how to start both.
  • You will be able to recognise prompt injection and name the guardrails that limit its damage.
  • You will have a complete harness design on paper for a real task of your own.

Who it is for

Builders and technically curious readers who want to understand or design an AI agent. You should be able to read a short Python example, but you do not need to write any code to do the exercises.

Before you start

  • Comfort with prompts, tokens and context windows (the What is AI workbook covers them). Having read AI agents explained helps but is not essential.

Keep learning

The complete workbook

This workbook opens up the part of an AI agent that most tutorials skip: the harness around the model. You will follow one full turn of the loop, read a minimal loop in Python, and meet every component that helps make an agent safe and dependable, from permissions and budgets to audit logs and tests. It ends with a paper design exercise for a real task.

  1. 01
    What a harness is, and why it matters

    On its own, a language model can only produce text in response to its input. It cannot open a file, search the web or send an email. Something has to sit around it and do those things on its behalf. That something is the harness.

    In the workbook · Reading
  2. 02
    One turn, step by step

    Read the loop first as a story with six beats, then as code.

    In the workbook · 1 exercise
  3. 03
    What the model sees: prompt, tools and memory

    The model is stateless. Each call starts from nothing, so on every turn the harness must rebuild everything the model needs to know. Getting that assembly right is most of the craft.

    In the workbook · Reading
  4. 04
    Control: permissions, sandboxes, budgets and stops

    A harness earns its keep by saying no. These controls decide what an agent may do, how much it may spend and when it must stop.

    In the workbook · 1 exercise
  5. 05
    Trust: logs, tests and prompt injection

    You will not watch every step. Logs, tests and defences against hostile text are how you keep trust in a system that acts on its own.

    In the workbook · 1 exercise
  6. 06
    Design a harness on paper

    The best way to make the parts stick is to design a whole harness for a real task before writing any code.

    In the workbook · 1 exercise

Also inside: a 9-point checklist, a glossary of 12 terms and 10 questions and answers to test yourself. 4 hands-on exercises, each with a worked answer at the back where the workbook gives one.

No login, no card, no account. Before the download we ask you to follow Mickai (two quick links). Free to download and use for personal learning, study groups and inside your own team. Please do not resell the workbooks or republish them as your own. Link people to trust-agent.ai instead.

Test yourself

Questions and answers

What is a loop harness?

It is the program that runs a language model in a loop. It builds the context, calls the model, runs any tool the model asks for, feeds the result back and repeats until a stop condition is met. The harness also enforces permissions, budgets and logging.

Is the harness the same thing as the model?

No. The model turns text into text. The harness is ordinary software around it that supplies context, runs tools, applies limits and keeps records. Two agents on the same model can behave very differently because their harnesses differ.

What ends the loop?

A named stop condition. Usually the model replies with an answer and no tool request. Other causes are a passed success check, a spent budget, a repeating pattern, a human pressing stop, an unrecoverable error or a policy breach. Record the reason every time.

Why does an agent need budgets?

Because a loop can run without end. Limits on steps, tokens, time and money stop a confused agent from looping, growing its context or spending without control. Check them before each call and stop with a report when one runs out.

Which actions should need human approval?

Anything hard to undo or that leaves your control, such as deleting data, sending messages, publishing or paying. Show the approver the exact action. Do not ask for approval on trivial steps, or people will stop reading the requests.

What is prompt injection?

It is hostile or unintended instructions hidden in text the model reads, such as a web page or email. Because instructions and data share one channel, the model may obey them. Reduce the risk with least privilege, approvals, allow-lists and treating outside text as untrusted.

What is a tamper-evident log?

A log where each entry includes a fingerprint made from its content and the previous fingerprint. Editing or deleting an old entry breaks every later fingerprint, so the change can be detected. It does not prevent editing, and the latest fingerprint should be stored separately.

How do I test an agent when its answers vary?

Write scenarios with checks on the end result rather than exact wording, and run each several times. Include failure cases and hostile text. Test the harness with a fake model that returns scripted replies, and re-run everything after any change.

Why compact the context instead of sending the whole history?

The history grows every turn, and each call pays for all of it again in cost and delay. It also eventually exceeds the context window. Compaction keeps recent turns and a short summary, while pinning the instructions and the task.

Do I need a framework to build a harness?

No. The core loop is short, as the example in this workbook shows. A framework can help with plumbing, but you should still understand and own the permissions, budgets, stop conditions and logs, because the agent's safety depends on them.

When you have finished

Get your certificate of completion

Type your name and download a certificate for this workbook as a PDF, ready to print or to add to LinkedIn. It is made on your own device, so your name is never sent to us. It is a self-declared certificate, not an accredited qualification.

Learn the language

Key terms

Agent harness
The program that runs a model in a loop, supplying context and tools and enforcing limits. Also called a loop harness.
Tool call
A structured request from the model to run a named tool with given arguments. It is only a request until the harness runs it.
Context assembly
Building the full prompt for each model call from the system prompt, tool definitions, task, history and any saved notes.
Compaction
Replacing older parts of the conversation with a shorter summary so the context stays within its limit.
Sandbox
A restricted environment in which tools run, limiting what they can read, change or reach.
Least privilege
Giving a system only the access it needs for its task and nothing more.

6 of the workbook's 12 terms. The complete glossary is in the workbook.

Follow the evidence

Sources and checks

Facts last checked: .

These workbooks use AI assistance. See how the workbooks are made.

  1. OWASP Top 10 for LLM Applications (includes prompt injection)OWASP Gen AI Security Project
  2. Prompt injection is not SQL injection (it may be worse)National Cyber Security Centre (NCSC)
  3. ReAct: Synergizing Reasoning and Acting in Language Models (Yao et al., 2022)arXiv

Created by Mickarle Wagstaff-Irons - Micky Irons with the Mickai team. Published by Mickai LTD. Last updated 25 September 2026.

NextKeep going

Where to go next