agents · Level 3

Tool calling and structured outputs: validate every response

Let an assistant call your code safely: JSON Schema contracts, a three-stage check on every reply, retries, timeouts, dry runs, logs and untrusted tool results.

By Mickarle Wagstaff-Irons - Micky Irons

  • Level 3Building
  • 95 min
  • 5 chapters
  • Free PDF, no account
The Loopagents / 03

Start with the essentials

The short answer

Tool calling lets a language model propose a call to a function you describe, with its arguments as JSON. The model never runs anything: your code decides whether and how. Treat every reply as untrusted input. Parse it, validate it against a JSON Schema and check your business rules, then run only allowlisted tools, with bounded arguments, timeouts, dry runs for changes and a log.

What you will learn

  • You will be able to explain what tool calling is and why your code, not the model, decides what runs.
  • You will be able to write JSON Schema contracts for a tool's arguments and results, and check them with the jsonschema package.
  • You will be able to build a gate that parses, schema-validates and rule-checks every reply, with limited retries and feedback.
  • You will be able to run tools behind an allowlist, with bounds, timeouts, dry runs, idempotency keys and a log.
  • You will be able to explain why schema-valid is not the same as correct, and why text returned by a tool is data, never instructions.

Who it is for

Anyone who has built a simple chatbot and now wants an assistant to call code safely. You should be able to read and run short Python scripts from a terminal. No model account or API key is needed, and apart from installing one package, nothing uses the network.

Before you start

  • You should have built or read through a simple chatbot (the Build your first chatbot workbook covers it) and be able to run Python scripts from a terminal (Python: from first script to a useful automation covers that). You need Python 3.11 or later.

Read a sample · Chapter 02 of 05

02

JSON Schema: the contract

A contract written as data can be checked by code, every time.

JSON Schema is a vocabulary, written in JSON, for describing valid data. Its current version is draft 2020-12. The model interfaces and the open tool protocol in this workbook's sources use it to describe tool arguments, and the protocol can describe results with it too. The jsonschema package checks it in Python.

Scroll sideways to see every column.

JSON Schema: the contract · Table 1
KeywordMeans
typeObject, array, string, number, integer, boolean or null
requiredNames that must be present
additionalProperties: falseNo names beyond those listed
enumOne of a fixed list
minimum, maximumInclusive numeric bounds
pattern, maxLengthShape and length of text

Save this as contracts.py. ARGUMENTS doubles as the allowlist: a tool that is not a key here can never run. RESULTS is the contract for what a tool sends back.

python · 47 lines
"""contracts.py: JSON Schemas (draft 2020-12) for replies, arguments and results."""
from jsonschema import Draft202012Validator

UNIT_KIND = {"m": "length", "km": "length", "g": "mass", "kg": "mass"}
ENVELOPE = {
    "type": "object",
    "properties": {"tool": {"type": "string"}, "arguments": {"type": "object"}},
    "required": ["tool", "arguments"],
    "additionalProperties": False,
}
ARGUMENTS = {  # the allowlist: a tool not named here can never run
    "convert_units": {
        "type": "object",
        "properties": {
            "value": {"type": "number", "minimum": -1000000, "maximum": 1000000},
            "from_unit": {"enum": list(UNIT_KIND)},
            "to_unit": {"enum": list(UNIT_KIND)},
        },
        "required": ["value", "from_unit", "to_unit"],
        "additionalProperties": False,
    },
    "lookup_stock": {
        "type": "object",
        "properties": {"ticker": {"enum": ["ZEPHYR", "SNAIL"]}},
        "required": ["ticker"],
        "additionalProperties": False,
    },
}
RESULTS = {
    "lookup_stock": {
        "type": "object",
        "properties": {"ticker": {"type": "string"}, "price": {"type": "number"},
                       "summary": {"type": "string", "maxLength": 200}},
        "required": ["ticker", "price", "summary"],
        "additionalProperties": False,
    },
}


def problems(instance, schema):
    """Check the schema itself, then return every error message, sorted."""
    Draft202012Validator.check_schema(schema)
    return sorted(e.message for e in Draft202012Validator(schema).iter_errors(instance))


if __name__ == "__main__":
    print(problems({"value": "5", "from_unit": "miles"}, ARGUMENTS["convert_units"]))

text · 2 lines
python contracts.py
["'5' is not of type 'number'", "'miles' is not one of ['m', 'km', 'g', 'kg']", "'to_unit' is a required property"]

Every problem is reported at once, ready to feed back to a model. The jsonschema documentation adds two surprises: format (such as "format": "date") is not checked unless you switch that on, and default fills in nothing.

Who enforces the contract? A prompt can only ask. Some services offer more.

Scroll sideways to see every column.

JSON Schema: the contract · Table 2
ApproachWhat you getStill check
Ask in the promptA request onlyEverything
JSON mode, where offeredOutput that parses as JSON, not necessarily your schemaSchema, rules, meaning
Structured outputs, where offeredGeneration constrained to your schemaUnsupported keywords, refusals, cut-off replies, rules, meaning

One provider lists minimum and maximum as unsupported, and providers note that refusals and cut-off replies may not match. Keep your own gate whatever a service promises.

Try it yourself · Activity 02

15 min

Write the note contract

Add a contract for a third tool, save_note.

  1. In ARGUMENTS, add save_note with two required strings, name and text, and nothing else.
  2. Allow only lower-case letters, digits and hyphens in name, 1 to 40 characters, and at most 500 characters of text.
  3. Test it with problems on a good note and on the name ../secrets.

Why is a strict pattern worth more than a warning in the tool's description?

Worked answer

Add this inside ARGUMENTS, after lookup_stock (later chapters need it): "save_note": {"type": "object", "properties": {"name": {"type": "string", "pattern": "^[a-z0-9-]{1,40}$"}, "text": {"type": "string", "maxLength": 500}}, "required": ["name", "text"], "additionalProperties": False},. A good note gives []; ../secrets gives ["'../secrets' does not match '^[a-z0-9-]{1,40}$'"]. Keep ^ and $: patterns are not implicitly anchored. The jsonschema package uses Python's regular expressions, whose $ also matches before a final line break, so "shopping\n" passes too; the runner's folder check still keeps it inside practice-notes.

Read a sample · Chapter 03 of 05

03

Validate every response, then retry

Three gates, in order. A rejected reply gets feedback and a limited second chance.

  1. Parse. json.loads turns text into data. Anything that is not JSON stops here.
  2. Validate. Check the envelope, the allowlist, then the arguments.
  3. Check the rules a schema cannot express, such as which units convert into which.

Save this as gate.py and run it.

python · 37 lines
"""gate.py: parse, then validate, then check business rules."""
import json

from contracts import ARGUMENTS, ENVELOPE, UNIT_KIND, problems


def refuse(constant):
    raise ValueError(f"{constant} is not allowed in JSON")


def check_call(reply):
    """Return (call, None) if the reply may run, or (None, reason) if not."""
    if len(reply) > 2000:
        return None, "reply too long"
    try:
        call = json.loads(reply, parse_constant=refuse)
    except ValueError as error:  # json.JSONDecodeError is a ValueError
        return None, f"not JSON: {error}"
    found = problems(call, ENVELOPE)
    if found:
        return None, "envelope: " + "; ".join(found)
    tool, args = call["tool"], call["arguments"]
    if tool not in ARGUMENTS:
        return None, f"tool not allowed: {tool}"
    found = problems(args, ARGUMENTS[tool])
    if found:
        return None, "arguments: " + "; ".join(found)
    if tool == "convert_units" and UNIT_KIND[args["from_unit"]] != UNIT_KIND[args["to_unit"]]:
        return None, f"rule: cannot convert {args['from_unit']} to {args['to_unit']}"
    return call, None


if __name__ == "__main__":
    from fake_model import REPLIES
    for name, reply in REPLIES.items():
        call, reason = check_call(reply)
        print(f"{name:<13}", "RUN" if call else "REJECT " + reason)

text · 8 lines
python gate.py
good          RUN
malformed     REJECT not JSON: Expecting ',' delimiter: line 1 column 51 (char 50)
wrong_type    REJECT arguments: 'five' is not of type 'number'
extra_arg     REJECT arguments: Additional properties are not allowed ('round' was unexpected)
unknown_tool  REJECT tool not allowed: delete_all_notes
out_of_range  REJECT arguments: 1000000000.0 is greater than the maximum of 1000000
smuggled      REJECT envelope: Additional properties are not allowed ('system' was unexpected)

Python's json accepts NaN and Infinity, which the JSON standard forbids, so parse_constant refuses them. A number too large for a float, such as 1e400, still becomes infinity; here the maximum stops it. Its documentation also recommends limiting the size of untrusted JSON, hence the 2,000-character cap.

Now save retry.py. Each rejection goes back as short feedback, and after three attempts it gives up rather than guessing.

python · 20 lines
"""retry.py: send feedback and ask again, a limited number of times."""
from fake_model import fake_model
from gate import check_call


def get_valid_call(script, max_attempts=3):
    messages = [{"role": "user", "content": "Convert 5 km to metres."}]
    for attempt in range(max_attempts):
        reply = fake_model(script, attempt)
        call, reason = check_call(reply)
        print(f"attempt {attempt + 1}:", "accepted" if call else reason)
        if call:
            return call
        messages += [{"role": "assistant", "content": reply},
                     {"role": "user", "content": f"Rejected: {reason}. Reply with one "
                      "JSON object that matches the tool schema, and nothing else."}]
    return None  # give up and tell the person; never guess


print(get_valid_call(["unknown_tool", "wrong_type", "good"]))

text · 5 lines
python retry.py
attempt 1: tool not allowed: delete_all_notes
attempt 2: arguments: 'five' is not of type 'number'
attempt 3: accepted
{'tool': 'convert_units', 'arguments': {'value': 5, 'from_unit': 'km', 'to_unit': 'm'}}

The stand-in ignores feedback. A real model usually corrects itself, but not always, so the limit matters. Keep feedback factual and free of secrets.

Try it yourself · Activity 03

15 min

Break the gate on purpose

Add two replies to REPLIES, predict each verdict, then run.

  1. Add "mixed": convert 3 kg into km. Add "nan": the good call with NaN in place of 5.
  2. Run gate.py, then again without parse_constant=refuse. Put it back.
  3. In retry.py, add max_attempts=2 to the get_valid_call call and run it.

Which gate should catch nan if parsing lets it through?

Worked answer

"mixed": convert('{"value": 3, "from_unit": "kg", "to_unit": "km"}'), prints REJECT rule: cannot convert kg to km: it is schema-valid, so only the rule stops it. "nan": convert('{"value": NaN, "from_unit": "km", "to_unit": "m"}'), prints REJECT not JSON: NaN is not allowed in JSON, but RUN without parse_constant: NaN is neither below the minimum nor above the maximum. A rule rejecting values that are not finite would catch it. With max_attempts=2, retry.py prints two rejections, then None.

Keep learning

The complete workbook

This workbook shows how to let an assistant call your code without trusting what it says. With a scripted stand-in for a model, so every result is reproducible and no account is needed, you will write JSON Schema contracts, build a gate that parses, validates and rule-checks each reply, retry with feedback, and run tools with timeouts, dry runs, idempotency and a log. It ends with instructions hidden in tool results, and why validation alone cannot stop them.

  1. 01
    The model proposes, your code decides

    A model never runs your code. It writes a request, and your program decides what happens.

    In the workbook · 1 exercise
  2. 02
    JSON Schema: the contract

    A contract written as data can be checked by code, every time.

    Read here · 1 exercise
  3. 03
    Validate every response, then retry

    Three gates, in order. A rejected reply gets feedback and a limited second chance.

    Read here · 1 exercise
  4. 04
    Run tools safely

    Passing the gate earns a call the right to run, on your conditions.

    In the workbook · 1 exercise
  5. 05
    Untrusted text coming back from tools

    Tool results flow into the model's next prompt. If they contain instructions, the model may follow them.

    In the workbook · 1 exercise

Also inside: a 8-point checklist, a glossary of 12 terms and 10 questions and answers to test yourself. 5 hands-on exercises, each with a worked answer at the back where the workbook gives one.

No login, no card, no account. Before the download we ask you to follow Mickai (two quick links). Free to download and use for personal learning, study groups and inside your own team. Please do not resell the workbooks or republish them as your own. Link people to trust-agent.ai instead.

Test yourself

Questions and answers

What is tool calling, and does the model run my code?

Tool calling lets a language model ask for one of your functions to be run. You describe each tool with a name, a description and a schema for its arguments. The model replies with a structured request, usually JSON. It runs nothing itself: your code decides whether to run the request, and sends back the result. Some services also run built-in tools on their side, which this workbook does not cover.

Is asking the model to reply in JSON enough?

No. A prompt is a request, not a constraint, so replies can include prose, miss fields or break the format. Always parse, schema-validate and rule-check the reply in code, and reject or retry anything that fails.

If a service offers structured outputs, do I still need to validate?

Yes. Providers document that only some schema keywords are supported, and that refusals or replies cut off at a length limit may not match. A schema also cannot check meaning. Your own validation is the part you control.

What does additionalProperties false do?

It rejects any property the schema does not list. For tool arguments this stops a model adding options you never offered, such as a flag that skips a dry run, and it catches extra fields carrying smuggled instructions.

How many times should I retry a rejected reply?

A small, fixed number, such as the three used here. Send short, factual feedback naming what failed. If the limit is reached, stop and tell the person plainly. Never retry an action that changes something without an idempotency key or a check that it has not already happened.

If a reply passes the schema, is it correct?

Not necessarily. A schema checks shape: types, required fields, bounds and allowed values. A well-formed call can still pick the wrong unit, the wrong record, or an action someone tricked the model into. Add business rules, check results against the request, and limit what any accepted call can do.

Why make tools dry-run by default and use idempotency keys?

A dry run shows what would change before anything does, so a person or a later check can stop it. An idempotency key lets you recognise a repeated request, so a retry or a duplicate reply does not act twice.

Can I filter prompt injection out of tool results?

Not reliably. The NCSC warns against blocking known phrases, because attacks can be reworded endlessly, and notes that current models do not separate instructions from data. Treat results as untrusted data, keep tools narrow, validate results, dry-run changes and ask a person before anything hard to undo.

Do I need a model account or API key for these exercises?

No. A scripted stand-in plays the model, so every result is reproducible and no model is called. With a real model, the same gate, runner and log apply once its tool request is in the same shape; only the replies become less predictable. Never paste real keys, passwords or personal data into practice code.

Which versions of Python and jsonschema do I need?

Python 3.11 or later, because the runner catches the built-in TimeoutError, which concurrent.futures has raised since 3.11. Install jsonschema 4.26.0 exactly, as shown, so your output matches; it validates JSON Schema draft 2020-12. The workbook was tested with Python 3.12.10 on Windows 11.

When you have finished

Get your certificate of completion

Type your name and download a certificate for this workbook as a PDF, ready to print or to add to LinkedIn. It is made on your own device, so your name is never sent to us. It is a self-declared certificate, not an accredited qualification.

Learn the language

Key terms

Tool calling
A model proposing that your code run a described function with given arguments. Also called function calling.
Tool call
The structured request itself: a tool name and its arguments. It is only a request until your code runs it.
JSON
A plain text format for data built from objects (names and values), arrays, strings, numbers, true, false and null.
JSON Schema
A vocabulary, written in JSON, for describing valid data. This workbook uses draft 2020-12.
Validation
Checking data against a contract before using it. Here: parse, then schema, then business rules.
Allowlist
A fixed list of what is permitted. Anything not on it is refused.

6 of the workbook's 12 terms. The complete glossary is in the workbook.

Follow the evidence

Sources and checks

Facts last checked: .

Examples in this workbook were run on: Python 3.12.10 on Windows 11, in a fresh virtual environment with jsonschema 4.26.0 (validating JSON Schema draft 2020-12 through Draft202012Validator), installed by pip 25.0.1 together with its dependencies attrs 26.1.0, jsonschema-specifications 2025.9.1, referencing 0.37.0, rpds-py 2026.6.3 and typing_extensions 4.16.0. No model was called and no API key was used; only pip used the network, to download the packages. (2026-09-26).

These workbooks use AI assistance. See how the workbooks are made.

  1. JSON Schema specification (current version 2020-12)JSON Schema
  2. JSON Schema: A Media Type for Describing JSON Documents (Core, draft 2020-12)JSON Schema
  3. JSON Schema Validation: A Vocabulary for Structural Validation of JSON (draft 2020-12)JSON Schema
  4. jsonschema 4.26.0 documentationpython-jsonschema (Read the Docs)
  5. jsonschema: Schema Validationpython-jsonschema (Read the Docs)
  6. jsonschema: Frequently Asked Questionspython-jsonschema (Read the Docs)
  7. RFC 8259: The JavaScript Object Notation (JSON) Data Interchange FormatIETF (RFC Editor)
  8. json: JSON encoder and decoderPython Software Foundation
  9. re: regular expression operationsPython Software Foundation
  10. concurrent.futures: launching parallel tasksPython Software Foundation
  11. pathlib: object-oriented filesystem pathsPython Software Foundation
  12. venv: creation of virtual environmentsPython Software Foundation
  13. OWASP Top 10 for LLM Applications 2025OWASP Gen AI Security Project
  14. LLM01:2025 Prompt InjectionOWASP Gen AI Security Project
  15. LLM05:2025 Improper Output HandlingOWASP Gen AI Security Project
  16. LLM06:2025 Excessive AgencyOWASP Gen AI Security Project
  17. Prompt injection is not SQL injection (it may be worse)National Cyber Security Centre (NCSC)
  18. Model Context Protocol specification (2025-06-18): ToolsModel Context Protocol
  19. Function calling (vendor documentation, cited for the general pattern only)OpenAI
  20. Structured Outputs (vendor documentation, cited for the general pattern only)OpenAI
  21. Tool use overview (vendor documentation, cited for the general pattern only)Anthropic
  22. Structured outputs (vendor documentation, cited for the general pattern only)Anthropic

Created by Mickarle Wagstaff-Irons - Micky Irons with the Mickai team. Published by Mickai LTD. Last updated 26 September 2026.

NextKeep going

Where to go next