agents · Level 3
Tool calling and structured outputs: validate every response
Let an assistant call your code safely: JSON Schema contracts, a three-stage check on every reply, retries, timeouts, dry runs, logs and untrusted tool results.

Start with the essentials
The short answer
Tool calling lets a language model propose a call to a function you describe, with its arguments as JSON. The model never runs anything: your code decides whether and how. Treat every reply as untrusted input. Parse it, validate it against a JSON Schema and check your business rules, then run only allowlisted tools, with bounded arguments, timeouts, dry runs for changes and a log.
What you will learn
- You will be able to explain what tool calling is and why your code, not the model, decides what runs.
- You will be able to write JSON Schema contracts for a tool's arguments and results, and check them with the jsonschema package.
- You will be able to build a gate that parses, schema-validates and rule-checks every reply, with limited retries and feedback.
- You will be able to run tools behind an allowlist, with bounds, timeouts, dry runs, idempotency keys and a log.
- You will be able to explain why schema-valid is not the same as correct, and why text returned by a tool is data, never instructions.
Who it is for
Anyone who has built a simple chatbot and now wants an assistant to call code safely. You should be able to read and run short Python scripts from a terminal. No model account or API key is needed, and apart from installing one package, nothing uses the network.
Before you start
- You should have built or read through a simple chatbot (the Build your first chatbot workbook covers it) and be able to run Python scripts from a terminal (Python: from first script to a useful automation covers that). You need Python 3.11 or later.
Read a sample · Chapter 02 of 05
JSON Schema: the contract
A contract written as data can be checked by code, every time.
JSON Schema is a vocabulary, written in JSON, for describing valid data. Its current version is draft 2020-12. The model interfaces and the open tool protocol in this workbook's sources use it to describe tool arguments, and the protocol can describe results with it too. The jsonschema package checks it in Python.
Scroll sideways to see every column.
| Keyword | Means |
|---|---|
type | Object, array, string, number, integer, boolean or null |
required | Names that must be present |
additionalProperties: false | No names beyond those listed |
enum | One of a fixed list |
minimum, maximum | Inclusive numeric bounds |
pattern, maxLength | Shape and length of text |
Save this as contracts.py. ARGUMENTS doubles as the allowlist: a tool that is not a key here can never run. RESULTS is the contract for what a tool sends back.
"""contracts.py: JSON Schemas (draft 2020-12) for replies, arguments and results."""
from jsonschema import Draft202012Validator
UNIT_KIND = {"m": "length", "km": "length", "g": "mass", "kg": "mass"}
ENVELOPE = {
"type": "object",
"properties": {"tool": {"type": "string"}, "arguments": {"type": "object"}},
"required": ["tool", "arguments"],
"additionalProperties": False,
}
ARGUMENTS = { # the allowlist: a tool not named here can never run
"convert_units": {
"type": "object",
"properties": {
"value": {"type": "number", "minimum": -1000000, "maximum": 1000000},
"from_unit": {"enum": list(UNIT_KIND)},
"to_unit": {"enum": list(UNIT_KIND)},
},
"required": ["value", "from_unit", "to_unit"],
"additionalProperties": False,
},
"lookup_stock": {
"type": "object",
"properties": {"ticker": {"enum": ["ZEPHYR", "SNAIL"]}},
"required": ["ticker"],
"additionalProperties": False,
},
}
RESULTS = {
"lookup_stock": {
"type": "object",
"properties": {"ticker": {"type": "string"}, "price": {"type": "number"},
"summary": {"type": "string", "maxLength": 200}},
"required": ["ticker", "price", "summary"],
"additionalProperties": False,
},
}
def problems(instance, schema):
"""Check the schema itself, then return every error message, sorted."""
Draft202012Validator.check_schema(schema)
return sorted(e.message for e in Draft202012Validator(schema).iter_errors(instance))
if __name__ == "__main__":
print(problems({"value": "5", "from_unit": "miles"}, ARGUMENTS["convert_units"]))python contracts.py
["'5' is not of type 'number'", "'miles' is not one of ['m', 'km', 'g', 'kg']", "'to_unit' is a required property"]Every problem is reported at once, ready to feed back to a model. The jsonschema documentation adds two surprises: format (such as "format": "date") is not checked unless you switch that on, and default fills in nothing.
Who enforces the contract? A prompt can only ask. Some services offer more.
Scroll sideways to see every column.
| Approach | What you get | Still check |
|---|---|---|
| Ask in the prompt | A request only | Everything |
| JSON mode, where offered | Output that parses as JSON, not necessarily your schema | Schema, rules, meaning |
| Structured outputs, where offered | Generation constrained to your schema | Unsupported keywords, refusals, cut-off replies, rules, meaning |
One provider lists minimum and maximum as unsupported, and providers note that refusals and cut-off replies may not match. Keep your own gate whatever a service promises.
Try it yourself · Activity 02
15 minWrite the note contract
Add a contract for a third tool, save_note.
- In
ARGUMENTS, addsave_notewith two required strings,nameandtext, and nothing else. - Allow only lower-case letters, digits and hyphens in
name, 1 to 40 characters, and at most 500 characters oftext. - Test it with
problemson a good note and on the name../secrets.
Why is a strict pattern worth more than a warning in the tool's description?
Worked answer
Add this inside ARGUMENTS, after lookup_stock (later chapters need it): "save_note": {"type": "object", "properties": {"name": {"type": "string", "pattern": "^[a-z0-9-]{1,40}$"}, "text": {"type": "string", "maxLength": 500}}, "required": ["name", "text"], "additionalProperties": False},. A good note gives []; ../secrets gives ["'../secrets' does not match '^[a-z0-9-]{1,40}$'"]. Keep ^ and $: patterns are not implicitly anchored. The jsonschema package uses Python's regular expressions, whose $ also matches before a final line break, so "shopping\n" passes too; the runner's folder check still keeps it inside practice-notes.
Read a sample · Chapter 03 of 05
Validate every response, then retry
Three gates, in order. A rejected reply gets feedback and a limited second chance.
- Parse.
json.loadsturns text into data. Anything that is not JSON stops here. - Validate. Check the envelope, the allowlist, then the arguments.
- Check the rules a schema cannot express, such as which units convert into which.
Save this as gate.py and run it.
"""gate.py: parse, then validate, then check business rules."""
import json
from contracts import ARGUMENTS, ENVELOPE, UNIT_KIND, problems
def refuse(constant):
raise ValueError(f"{constant} is not allowed in JSON")
def check_call(reply):
"""Return (call, None) if the reply may run, or (None, reason) if not."""
if len(reply) > 2000:
return None, "reply too long"
try:
call = json.loads(reply, parse_constant=refuse)
except ValueError as error: # json.JSONDecodeError is a ValueError
return None, f"not JSON: {error}"
found = problems(call, ENVELOPE)
if found:
return None, "envelope: " + "; ".join(found)
tool, args = call["tool"], call["arguments"]
if tool not in ARGUMENTS:
return None, f"tool not allowed: {tool}"
found = problems(args, ARGUMENTS[tool])
if found:
return None, "arguments: " + "; ".join(found)
if tool == "convert_units" and UNIT_KIND[args["from_unit"]] != UNIT_KIND[args["to_unit"]]:
return None, f"rule: cannot convert {args['from_unit']} to {args['to_unit']}"
return call, None
if __name__ == "__main__":
from fake_model import REPLIES
for name, reply in REPLIES.items():
call, reason = check_call(reply)
print(f"{name:<13}", "RUN" if call else "REJECT " + reason)python gate.py
good RUN
malformed REJECT not JSON: Expecting ',' delimiter: line 1 column 51 (char 50)
wrong_type REJECT arguments: 'five' is not of type 'number'
extra_arg REJECT arguments: Additional properties are not allowed ('round' was unexpected)
unknown_tool REJECT tool not allowed: delete_all_notes
out_of_range REJECT arguments: 1000000000.0 is greater than the maximum of 1000000
smuggled REJECT envelope: Additional properties are not allowed ('system' was unexpected)Python's json accepts NaN and Infinity, which the JSON standard forbids, so parse_constant refuses them. A number too large for a float, such as 1e400, still becomes infinity; here the maximum stops it. Its documentation also recommends limiting the size of untrusted JSON, hence the 2,000-character cap.
Now save retry.py. Each rejection goes back as short feedback, and after three attempts it gives up rather than guessing.
"""retry.py: send feedback and ask again, a limited number of times."""
from fake_model import fake_model
from gate import check_call
def get_valid_call(script, max_attempts=3):
messages = [{"role": "user", "content": "Convert 5 km to metres."}]
for attempt in range(max_attempts):
reply = fake_model(script, attempt)
call, reason = check_call(reply)
print(f"attempt {attempt + 1}:", "accepted" if call else reason)
if call:
return call
messages += [{"role": "assistant", "content": reply},
{"role": "user", "content": f"Rejected: {reason}. Reply with one "
"JSON object that matches the tool schema, and nothing else."}]
return None # give up and tell the person; never guess
print(get_valid_call(["unknown_tool", "wrong_type", "good"]))python retry.py
attempt 1: tool not allowed: delete_all_notes
attempt 2: arguments: 'five' is not of type 'number'
attempt 3: accepted
{'tool': 'convert_units', 'arguments': {'value': 5, 'from_unit': 'km', 'to_unit': 'm'}}The stand-in ignores feedback. A real model usually corrects itself, but not always, so the limit matters. Keep feedback factual and free of secrets.
Try it yourself · Activity 03
15 minBreak the gate on purpose
Add two replies to REPLIES, predict each verdict, then run.
- Add
"mixed": convert 3kgintokm. Add"nan": the good call withNaNin place of 5. - Run
gate.py, then again withoutparse_constant=refuse. Put it back. - In
retry.py, addmax_attempts=2to theget_valid_callcall and run it.
Which gate should catch nan if parsing lets it through?
Worked answer
"mixed": convert('{"value": 3, "from_unit": "kg", "to_unit": "km"}'), prints REJECT rule: cannot convert kg to km: it is schema-valid, so only the rule stops it. "nan": convert('{"value": NaN, "from_unit": "km", "to_unit": "m"}'), prints REJECT not JSON: NaN is not allowed in JSON, but RUN without parse_constant: NaN is neither below the minimum nor above the maximum. A rule rejecting values that are not finite would catch it. With max_attempts=2, retry.py prints two rejections, then None.
Keep learning
The complete workbook
This workbook shows how to let an assistant call your code without trusting what it says. With a scripted stand-in for a model, so every result is reproducible and no account is needed, you will write JSON Schema contracts, build a gate that parses, validates and rule-checks each reply, retry with feedback, and run tools with timeouts, dry runs, idempotency and a log. It ends with instructions hidden in tool results, and why validation alone cannot stop them.
- 01The model proposes, your code decidesIn the workbook · 1 exercise
A model never runs your code. It writes a request, and your program decides what happens.
- 02JSON Schema: the contractRead here · 1 exercise
A contract written as data can be checked by code, every time.
- 03Validate every response, then retryRead here · 1 exercise
Three gates, in order. A rejected reply gets feedback and a limited second chance.
- 04Run tools safelyIn the workbook · 1 exercise
Passing the gate earns a call the right to run, on your conditions.
- 05Untrusted text coming back from toolsIn the workbook · 1 exercise
Tool results flow into the model's next prompt. If they contain instructions, the model may follow them.
Also inside: a 8-point checklist, a glossary of 12 terms and 10 questions and answers to test yourself. 5 hands-on exercises, each with a worked answer at the back where the workbook gives one.
No login, no card, no account. Before the download we ask you to follow Mickai (two quick links). Free to download and use for personal learning, study groups and inside your own team. Please do not resell the workbooks or republish them as your own. Link people to trust-agent.ai instead.
Test yourself
Questions and answers
What is tool calling, and does the model run my code?
Tool calling lets a language model ask for one of your functions to be run. You describe each tool with a name, a description and a schema for its arguments. The model replies with a structured request, usually JSON. It runs nothing itself: your code decides whether to run the request, and sends back the result. Some services also run built-in tools on their side, which this workbook does not cover.
Is asking the model to reply in JSON enough?
No. A prompt is a request, not a constraint, so replies can include prose, miss fields or break the format. Always parse, schema-validate and rule-check the reply in code, and reject or retry anything that fails.
If a service offers structured outputs, do I still need to validate?
Yes. Providers document that only some schema keywords are supported, and that refusals or replies cut off at a length limit may not match. A schema also cannot check meaning. Your own validation is the part you control.
What does additionalProperties false do?
It rejects any property the schema does not list. For tool arguments this stops a model adding options you never offered, such as a flag that skips a dry run, and it catches extra fields carrying smuggled instructions.
How many times should I retry a rejected reply?
A small, fixed number, such as the three used here. Send short, factual feedback naming what failed. If the limit is reached, stop and tell the person plainly. Never retry an action that changes something without an idempotency key or a check that it has not already happened.
If a reply passes the schema, is it correct?
Not necessarily. A schema checks shape: types, required fields, bounds and allowed values. A well-formed call can still pick the wrong unit, the wrong record, or an action someone tricked the model into. Add business rules, check results against the request, and limit what any accepted call can do.
Why make tools dry-run by default and use idempotency keys?
A dry run shows what would change before anything does, so a person or a later check can stop it. An idempotency key lets you recognise a repeated request, so a retry or a duplicate reply does not act twice.
Can I filter prompt injection out of tool results?
Not reliably. The NCSC warns against blocking known phrases, because attacks can be reworded endlessly, and notes that current models do not separate instructions from data. Treat results as untrusted data, keep tools narrow, validate results, dry-run changes and ask a person before anything hard to undo.
Do I need a model account or API key for these exercises?
No. A scripted stand-in plays the model, so every result is reproducible and no model is called. With a real model, the same gate, runner and log apply once its tool request is in the same shape; only the replies become less predictable. Never paste real keys, passwords or personal data into practice code.
Which versions of Python and jsonschema do I need?
Python 3.11 or later, because the runner catches the built-in TimeoutError, which concurrent.futures has raised since 3.11. Install jsonschema 4.26.0 exactly, as shown, so your output matches; it validates JSON Schema draft 2020-12. The workbook was tested with Python 3.12.10 on Windows 11.
When you have finished
Get your certificate of completion
Type your name and download a certificate for this workbook as a PDF, ready to print or to add to LinkedIn. It is made on your own device, so your name is never sent to us. It is a self-declared certificate, not an accredited qualification.
Learn the language
Key terms
- Tool calling
- A model proposing that your code run a described function with given arguments. Also called function calling.
- Tool call
- The structured request itself: a tool name and its arguments. It is only a request until your code runs it.
- JSON
- A plain text format for data built from objects (names and values), arrays, strings, numbers, true, false and null.
- JSON Schema
- A vocabulary, written in JSON, for describing valid data. This workbook uses draft 2020-12.
- Validation
- Checking data against a contract before using it. Here: parse, then schema, then business rules.
- Allowlist
- A fixed list of what is permitted. Anything not on it is refused.
6 of the workbook's 12 terms. The complete glossary is in the workbook.
Follow the evidence
Sources and checks
Facts last checked: .
Examples in this workbook were run on: Python 3.12.10 on Windows 11, in a fresh virtual environment with jsonschema 4.26.0 (validating JSON Schema draft 2020-12 through Draft202012Validator), installed by pip 25.0.1 together with its dependencies attrs 26.1.0, jsonschema-specifications 2025.9.1, referencing 0.37.0, rpds-py 2026.6.3 and typing_extensions 4.16.0. No model was called and no API key was used; only pip used the network, to download the packages. (2026-09-26).
These workbooks use AI assistance. See how the workbooks are made.
- JSON Schema specification (current version 2020-12)JSON Schema
- JSON Schema: A Media Type for Describing JSON Documents (Core, draft 2020-12)JSON Schema
- JSON Schema Validation: A Vocabulary for Structural Validation of JSON (draft 2020-12)JSON Schema
- jsonschema 4.26.0 documentationpython-jsonschema (Read the Docs)
- jsonschema: Schema Validationpython-jsonschema (Read the Docs)
- jsonschema: Frequently Asked Questionspython-jsonschema (Read the Docs)
- RFC 8259: The JavaScript Object Notation (JSON) Data Interchange FormatIETF (RFC Editor)
- json: JSON encoder and decoderPython Software Foundation
- re: regular expression operationsPython Software Foundation
- concurrent.futures: launching parallel tasksPython Software Foundation
- pathlib: object-oriented filesystem pathsPython Software Foundation
- venv: creation of virtual environmentsPython Software Foundation
- OWASP Top 10 for LLM Applications 2025OWASP Gen AI Security Project
- LLM01:2025 Prompt InjectionOWASP Gen AI Security Project
- LLM05:2025 Improper Output HandlingOWASP Gen AI Security Project
- LLM06:2025 Excessive AgencyOWASP Gen AI Security Project
- Prompt injection is not SQL injection (it may be worse)National Cyber Security Centre (NCSC)
- Model Context Protocol specification (2025-06-18): ToolsModel Context Protocol
- Function calling (vendor documentation, cited for the general pattern only)OpenAI
- Structured Outputs (vendor documentation, cited for the general pattern only)OpenAI
- Tool use overview (vendor documentation, cited for the general pattern only)Anthropic
- Structured outputs (vendor documentation, cited for the general pattern only)Anthropic
Created by Mickarle Wagstaff-Irons - Micky Irons with the Mickai team. Published by Mickai LTD. Last updated 26 September 2026.
NextKeep going
Where to go next
Recommended for you
The loop harness: how AI agents actually run
How an agent harness runs a model in a loop: context, tools, permissions, budgets, stop conditions, logging and testing, with a minimal Python loop and a paper design exercise.
Recommended for you
RAG and knowledge bases, step by step
Build a knowledge base a model can answer from: collect, clean, chunk, embed, retrieve, cite and evaluate, with security basics and a hands-on exercise using five short documents.
Recommended for you
Host a chatbot on your website, safely
Put a chatbot on your site without leaking keys or running up bills: server function, limits, CORS, privacy, accessibility, fallbacks, monitoring and a launch checklist.