applied · Level 3
Test-driven development with an AI coding assistant
Let a clear contract and failing examples guide the patch.

Start with the essentials
The short answer
Test-driven development starts with an observable example that fails, adds an implementation to make it pass, then improves the structure while preserving behaviour. With an AI coding assistant, keep the contract and expected outcomes independent of the proposed code. This workbook uses a fictional delivery-fee function to practise tests, patch review, boundary cases and honest handover.
What you will learn
- Translate a small request into explicit examples and rejection rules.
- Recognise an intended assertion failure instead of an environment error.
- Give an assistant a bounded implementation brief without weakening tests.
- Choose boundary cases from the contract rather than the proposed code.
- Use isolated mutations and refactoring to probe test usefulness.
- Record executed checks and limitations in a reviewable handover.
Who it is for
Builders who can read simple Python functions and want a disciplined way to review AI-assisted changes.
Before you start
- Read Python functions, conditionals and exceptions.
- Run a local Python file from a terminal.
- Understand a Git working tree and a small diff.
Read a sample · Chapter 01 of 06
Turn a request into observable examples
Before asking an assistant to write code, decide how you will recognise correct behaviour. A short, explicit contract gives you something independent of the generated implementation to test.
This lab calculates a fictional delivery fee. Its values are teaching inputs, not a real merchant policy or Mickai pricing. The function accepts a subtotal in integer pence and a Boolean express flag, then returns only the delivery fee in integer pence. It does not calculate the basket total, collect payment or decide whether a real order can be shipped. That small boundary lets us inspect every decision.
Scroll sideways to see every column.
| Input or situation | Required behaviour |
|---|---|
| Subtotal type | An integer, excluding True and False |
| Subtotal range | 0 through 1,000,000 pence, inclusive |
| Express argument | Exactly True or False; omitted means False |
| Standard, subtotal below 5,000 | Return 499 |
| Standard, subtotal at least 5,000 | Return 0 |
| Express, any valid subtotal | Return 899 |
| Wrong type / outside range | Raise TypeError / ValueError respectively |
The threshold includes exactly 5,000. Express remains paid even above it. Zero is a valid calculation input here; that does not mean a checkout must accept an empty basket. Keep application policy separate from this deliberately narrow function. Strings, floats and Boolean subtotals are rejected rather than converted. Invalid input must not look like a legitimate free-delivery result.
Examples are an executable interpretation of this contract. A test oracle is the reason an expected answer is correct. Here that reason is the written table and arithmetic, not whatever value the function currently returns. If you copy expected results from a generated implementation, both can share the same mistake. Resolve an ambiguous requirement before turning it into a permanent assertion.
Use a new practice directory and Python 3.12 for the recorded environment. The checked run used Python 3.12.10 on Windows 11, with no third-party packages, network calls or model service. Save plain UTF-8 files. Commands use python; use your known interpreter command if it is named differently. Confirm the version rather than assuming which installation a terminal has selected.
mkdir tdd-lab
cd tdd-lab
python --versionTry it yourself · Activity 01
10 minWrite the oracle before the code
Predict fees directly from the contract.
- Write the standard fees for 4,999, 5,000 and 5,001 pence.
- Write the express fee at 5,000 pence.
- Decide what should happen for the string "5000" and Boolean True.
Which example distinguishes inclusive from exclusive comparison?
Worked answer
The standard fees are 499, 0 and 0. Express at 5,000 is 899. The string and Boolean subtotal both raise TypeError. The exact 5,000 standard example distinguishes >= from >; an example far above the threshold cannot do that.
Read a sample · Chapter 02 of 06
Make the first tests fail for the right reason
The red stage is useful evidence only when the intended test actually runs and rejects missing behaviour. A syntax error or an undiscovered test is a different problem.
Create fee.py with this intentionally incomplete implementation. It is a temporary starting point, not a usable answer. Returning zero makes a concrete false prediction that our first examples can reject. Keep it inside the practice directory so nobody mistakes it for a finished delivery rule.
def delivery_fee(subtotal_pence, express=False):
return 0Now create test_first.py exactly as shown. Each method names one observable condition and compares a literal expected fee with the public function result. The second case deliberately combines express delivery with the free standard threshold, so the distinction appears early. The file imports fee, which must be fee.py beside it, rather than an unrelated installed module.
import unittest
from fee import delivery_fee
class FirstExamples(unittest.TestCase):
def test_standard_below_threshold(self):
self.assertEqual(delivery_fee(4999), 499)
def test_express_still_costs_above_threshold(self):
self.assertEqual(delivery_fee(5000, express=True), 899)
if __name__ == '__main__':
unittest.main()python -m unittest -v test_first
Expected red result for the temporary implementation:
Ran 2 tests
FAILED (failures=2)The recorded run produced two assertion failures: zero was returned where 499 and 899 were required. Your timings and file paths may differ. Read the assertion lines as well as the summary. An import failure means the example never reached delivery_fee; fix the filename, directory or interpreter problem first. A message saying zero tests ran is not evidence that the fee logic works.
Keep this red result as a short note or local log. Do not write a unit test that expects a whole runner transcript, since paths, line numbers and timings change. The useful evidence is the command, two discovered tests, the mismatched outputs and the reason those outputs contradict the contract.
The red, green, refactor cycle works in small increments. You establish one missing behaviour, implement enough to satisfy it, and improve structure while retaining that behaviour. This workbook starts with two closely related examples to make the competing delivery rules visible; it is not a rule that every change must begin with exactly two tests.
Try it yourself · Activity 02
15 minDiagnose the red run
Run the temporary implementation and read the failure.
- Confirm that both named tests execute.
- Identify the expected and actual numbers in each failure.
- Write a different diagnosis for an import error or a zero-test result.
Why is a red terminal line alone insufficient evidence?
Worked answer
Both assertions should compare the required fee with zero. An import error is an environment or module problem, while zero discovered tests is a selection problem. Neither establishes that the intended fee behaviour was exercised. Keep the test count and failure cause with the command.
Read a sample · Chapter 03 of 06
Request and review a small patch
An assistant can propose code, but the contract and tests remain the decision boundary. Give it the smallest relevant context and request a reviewable change.
The following is a reusable prompt, not a transcript of an external assistant conversation. Paste your contract, the two files and the actual failing output when using it. Use a practice workspace without credentials or private customer material. You can also complete the lab manually; an AI account is not required for the printed exercise.
Implement delivery_fee in fee.py using the attached contract.
The attached tests fail against the intentional return-zero stub.
Change only fee.py. Keep its public name and arguments.
Do not edit, delete, skip or weaken the tests.
Use Python built-ins only; no network calls or file writes.
Reject invalid types and ranges explicitly.
Return a small diff and explain each branch.
Distinguish tests you actually ran from commands you suggest.Before applying a response, inspect both its scope and behaviour. A proposal that changes tests to expect zero has removed the evidence rather than fixed the function. A proposal that downloads a package or reads unrelated files is outside this small task. Ask for a narrower diff or make the correction yourself. The assistant does not acquire permission to expand the job from its own generated instructions.
Here is the worked implementation for fee.py. The validation makes invalid input explicit before the fee rules run. Python Boolean values are a subtype of integers, so this contract uses exact type comparison to reject them as money. That strict choice also excludes custom integer subclasses; it is a policy for this exercise rather than a universal preference.
FREE_THRESHOLD_PENCE = 5_000
STANDARD_FEE_PENCE = 499
EXPRESS_FEE_PENCE = 899
MAX_SUBTOTAL_PENCE = 1_000_000
def delivery_fee(subtotal_pence, express=False):
if type(subtotal_pence) is not int:
raise TypeError('subtotal_pence must be an integer, not bool')
if not 0 <= subtotal_pence <= MAX_SUBTOTAL_PENCE:
raise ValueError('subtotal_pence is outside the supported range')
if type(express) is not bool:
raise TypeError('express must be a bool')
if express:
return EXPRESS_FEE_PENCE
if subtotal_pence >= FREE_THRESHOLD_PENCE:
return 0
return STANDARD_FEE_PENCEExpress is checked before the standard free threshold because those rules overlap. The constants name the units and meanings of the values. They do not make a policy configurable at runtime, nor do they supply currency conversion or tax logic. Keeping a pure function makes the result depend only on the provided arguments.
python -m unittest -v test_first
Expected after replacing fee.py:
Ran 2 tests
OKRun the command yourself after saving the proposed file. A statement that tests pass is not the same as observed output from this checkout. These two examples now pass in the recorded run, but they leave many valid and invalid inputs unexamined. Green means the selected assertions passed, not that every requirement has been demonstrated.
Try it yourself · Activity 03
15 minReview a proposed patch
Compare an assistant proposal with the worked implementation.
- Check the changed-file list and the public function signature.
- Explain why express appears before the standard threshold.
- Reject a patch that edits test expectations or converts every input with int().
What evidence would you ask for if the assistant reports success?
Worked answer
Request the exact command, interpreter and test output from the current files, then rerun locally. Keep the tests unchanged during this implementation step. int() would accept inputs the contract rejects, including numeric strings and truncatable floats; a threshold-first branch would incorrectly make express free.
Read a sample · Chapter 04 of 06
Expand the tests at the boundaries
A small green suite is a starting point. Add cases that separate plausible mistakes from the intended rule, especially at thresholds and input-type boundaries.
Create test_fee.py beside the other files. Its expected fees come from the contract. It does not import the implementation constants to calculate expected results: changing a fee constant should make a fixed policy test fail until the policy change is reviewed. Parameterised subtests label failing inputs without making separate test methods for every row.
import unittest
from fee import delivery_fee
class DeliveryFeeTests(unittest.TestCase):
def test_standard_boundary_examples(self):
for subtotal, expected in [(0, 499), (4999, 499),
(5000, 0), (1_000_000, 0)]:
with self.subTest(subtotal=subtotal):
self.assertEqual(delivery_fee(subtotal), expected)
def test_express_is_always_paid(self):
for subtotal in [0, 4999, 5000, 1_000_000]:
with self.subTest(subtotal=subtotal):
self.assertEqual(delivery_fee(subtotal, True), 899)
def test_subtotal_rejects_other_types(self):
for value in [True, False, None, '5000', 5000.0]:
with self.subTest(value=value):
with self.assertRaises(TypeError):
delivery_fee(value)
def test_subtotal_rejects_outside_range(self):
for value in [-1, 1_000_001]:
with self.subTest(value=value):
with self.assertRaises(ValueError):
delivery_fee(value)
def test_express_requires_boolean(self):
for value in [None, 0, 1, 'yes', []]:
with self.subTest(value=value):
with self.assertRaises(TypeError):
delivery_fee(5000, value)
def test_omitted_express_means_standard(self):
self.assertEqual(delivery_fee(2500),
delivery_fee(2500, False))
def test_result_is_integer_pence(self):
for value, express in [(0, False), (5000, False),
(5000, True)]:
self.assertIs(type(delivery_fee(value, express)), int)
def test_calls_do_not_change_later_results(self):
self.assertEqual(delivery_fee(5000, True), 899)
self.assertEqual(delivery_fee(0), 499)
self.assertEqual(delivery_fee(5000), 0)
self.assertEqual(delivery_fee(5000, True), 899)
if __name__ == '__main__':
unittest.main()python -m unittest -v test_first test_fee
Expected combined result:
Ran 10 tests
OKThere are two first-example methods and eight expanded methods. Subtests cover more input cases than that runner count suggests. Read counts accurately when reporting results. The recorded release additionally checked 106 selected subtotals around the threshold and limits, making 318 direct calls across standard, explicit False and express modes. Those checks supplement the printed ten-test suite; they are not a benchmark or proof over every possible Python object.
The type and range tests express separate contract decisions. TypeError is required for a string subtotal, and ValueError for an integer just outside the permitted range. Testing only that some exception occurs could accept a NameError caused by a typo. The suite does not lock down exception message wording because wording is not part of this contract.
The repeat-call case checks that an express calculation does not contaminate a later standard result. This is a useful guard against accidental mutable state, not a concurrency test. Likewise, checking an integer return type catches a float result even if 499.0 compares numerically equal to 499. Each assertion should have a failure you can explain.
Try it yourself · Activity 04
20 minAdd one discriminating case
Choose a plausible defect and a test that distinguishes it.
- Predict the result for standard delivery at exactly 5,000.
- Explain why True is a poor subtotal test if it is silently treated as 1.
- Add a new valid or invalid case and state what mistaken implementation it rejects.
Would copying the function formula into the test create independent evidence?
Worked answer
The exact threshold must return zero. True must be rejected even though Python relates bool and int. A new case is useful when it distinguishes a plausible defect, such as a fractional subtotal or a non-Boolean express value. Repeating the implementation formula risks repeating its error; tie the expected outcome to the written contract.
Read a sample · Chapter 05 of 06
Challenge the tests and refactor with evidence
Passing tests become more informative when you know they reject realistic defects. Deliberately alter one rule in an isolated copy, observe the failure, then restore the correct version.
A mutation is a deliberate small code defect used to probe the tests. Make it in a separate practice copy, and never distribute that copy as the working solution. Run the suite, identify the expected failure and restore the correct code before proceeding. If a mutation does not fail, investigate whether the tests miss the behaviour or whether the edit changed nothing observable under this contract.
Scroll sideways to see every column.
| Deliberate defect | Distinguishing check |
|---|---|
| Change >= to > at the threshold | Standard 5,000 must remain free |
| Accept isinstance(subtotal, int) | Boolean subtotals must raise TypeError |
| Check free threshold before express | Express at 5,000 must still cost 899 |
All three defects above were made separately against the printed implementation and rejected by the expanded learner suite. This establishes sensitivity to these three changes. It does not give an overall mutation score or establish that the tests detect every meaningful defect. A surviving mutant can sometimes be equivalent for the supported inputs; classify it before adding an assertion merely to increase a number.
Refactoring changes organisation while preserving supported behaviour. For a small exercise, rename a private constant consistently or extract a validation helper without changing the public function name, arguments, results or exceptions. Run the unchanged suite immediately. Avoid combining a refactor with a new discount rule: if behaviour changes intentionally, write and review that new contract separately.
When the lab is in an existing practice Git repository, inspect its status and diff before accepting a patch. Keep tests visible alongside source. If the files are new and untracked, an ordinary diff does not show their contents; open them directly as well. The staged diff shows what is prepared for the next commit, which may differ from the current working files.
git status --short
git diff -- fee.py test_first.py test_fee.py
git diff --cached -- fee.py test_first.py test_fee.pyAsk an assistant to explain a failing mutation or suggest a smaller refactor, then compare that suggestion with the actual diff. Do not approve a change solely because its explanation sounds convincing. Inspect removed assertions, skipped tests, broad exception handlers and added side effects. Retain useful test failures in your notes instead of repeatedly regenerating both code and tests until the output becomes green.
Try it yourself · Activity 05
15 minDemonstrate a caught defect
Use an isolated copy for one mutation from the table.
- Run the correct eight-test file and record success.
- Introduce exactly one listed defect and identify the failed assertion.
- Restore the correct implementation and rerun before making a small refactor.
What would you investigate if the mutation stayed green?
Worked answer
Check that you edited the file being imported and ran the intended suite. Then inspect whether a distinguishing input is absent or the mutation is behaviourally equivalent within the contract. The >= to > change should fail at exactly 5,000; restoring >= should recover the suite. Do not leave the mutated file in the final deliverable.
Read a sample · Chapter 06 of 06
Hand over a tested change with honest limits
A useful handover connects the request, changed files and observed checks. It also states which behaviours remain outside the evidence, so the next person knows what to verify.
The finished lab contains fee.py, test_first.py and test_fee.py. Its policy is fictional, with no network or persistent state. Record the interpreter version, exact test command, discovered method count and result. Include a short description of each deliberately caught defect. Separate a check you ran from a suggested future check, and record any assistant contribution as a proposal reviewed against the contract.
Change: implement fictional delivery_fee contract.
Files: fee.py, test_first.py, test_fee.py.
Run: python -m unittest -v test_first test_fee
Recorded environment: Python 3.12.10, Windows 11.
Result: 10 test methods passed.
Probe: three named defects were each rejected in isolation.
Scope: pure local fee calculation, integer pence only.
Not checked: checkout integration, concurrency, public service.A unit test of delivery_fee does not establish that a website passes the correct subtotal or handles its exceptions well. A real caller might use pounds instead of pence, pass a string from a form or forget the express choice. An integration test should exercise that boundary with representative input and check both displayed results and error handling. Do not add a database or browser to this lab merely to make the test suite look larger.
If a real feature requires currencies, discounts, postcode eligibility, refunds or policy changes, revisit the specification first. Decide the source of each rule and test its interactions. Avoid turning this fictional tariff into business advice. The lesson is how to establish and preserve a small contract, not which delivery prices to charge.
AI assistance can reduce typing and suggest overlooked cases, but speed is not evidence of correctness. Ask for one bounded patch at a time, inspect the changed files, rerun the acceptance checks and keep an explanation of decisions that matter. Where expected outcomes are unclear, investigate the requirement rather than asking a model to invent certainty. You remain responsible for choosing the oracle and accepting the result.
The next step is to use the same approach on a small feature in an app you already understand. Select a pure function or narrow boundary first. Write a failing example before implementation, add the edge cases that matter, and finish with a reviewable diff and clear evidence. Broader application evaluation complements these checks when outcomes depend on models, external tools or changing data.
Try it yourself · Activity 06
15 minWrite a reviewable handover
Summarise the completed lab for another developer.
- State the contract and the three final files.
- Record the actual local command and result, with no invented execution claim.
- Name one missing integration check and the input boundary it would cover.
Can another person distinguish observed results from planned work?
Worked answer
A sound note identifies the fictional integer-pence rule, exact files and real test result. It describes the mutation probes separately from normal tests. A useful next integration check sends a form subtotal through its parsing layer into the function and verifies both valid output and rejected values. It does not claim the local unit suite proves a production checkout works.
Keep learning
The complete workbook
Build a small local Python function from a written contract. Observe the initial failures, review a bounded implementation, expand the tests and prove that three deliberate defects are caught. AI use is optional; every runnable file is included.
- 01Turn a request into observable examplesRead here · 1 exercise
Before asking an assistant to write code, decide how you will recognise correct behaviour. A short, explicit contract gives you something independent of the generated implementation to test.
- 02Make the first tests fail for the right reasonRead here · 1 exercise
The red stage is useful evidence only when the intended test actually runs and rejects missing behaviour. A syntax error or an undiscovered test is a different problem.
- 03Request and review a small patchRead here · 1 exercise
An assistant can propose code, but the contract and tests remain the decision boundary. Give it the smallest relevant context and request a reviewable change.
- 04Expand the tests at the boundariesRead here · 1 exercise
A small green suite is a starting point. Add cases that separate plausible mistakes from the intended rule, especially at thresholds and input-type boundaries.
- 05Challenge the tests and refactor with evidenceRead here · 1 exercise
Passing tests become more informative when you know they reject realistic defects. Deliberately alter one rule in an isolated copy, observe the failure, then restore the correct version.
- 06Hand over a tested change with honest limitsRead here · 1 exercise
A useful handover connects the request, changed files and observed checks. It also states which behaviours remain outside the evidence, so the next person knows what to verify.
Also inside: a 9-point checklist, a glossary of 8 terms and 10 questions and answers to test yourself. 6 hands-on exercises, each with a worked answer at the back where the workbook gives one.
No login, no card, no account. Before the download we ask you to follow Mickai (two quick links). Free to download and use for personal learning, study groups and inside your own team. Please do not resell the workbooks or republish them as your own. Link people to trust-agent.ai instead.
Test yourself
Questions and answers
Are the delivery prices a real merchant policy?
No. Every amount and rule is an invented teaching fixture, not a Mickai service or commercial recommendation.
What makes the exact 5,000 case useful?
It distinguishes an inclusive free threshold from an exclusive comparison. A value far above the threshold does not distinguish them.
Does an import error establish the red stage?
It shows the suite could not reach the intended behaviour. Fix the environment or module path, then observe the intended assertion fail.
Should a coding assistant change expected values to make tests pass?
Not during implementation of an unchanged contract. Review any proposed policy change separately, with a reason and new acceptance examples.
Why reject a Boolean subtotal explicitly?
Python relates bool to int, but the lab contract accepts money as an integer excluding True and False. Exact type validation enforces that decision.
Why test specific exception classes?
A broad exception check could accept an unrelated defect such as NameError. TypeError and ValueError express different contract failures here.
Why not calculate expected fees from imported constants?
A mistaken constant change could alter both implementation and expected result. These policy examples keep fixed expected amounts derived from the contract.
Do three caught mutations prove all defects are detectable?
No. They demonstrate sensitivity to three specified changes. Other faults, equivalent mutations and untested inputs remain possible.
Does git diff show a new untracked file?
An ordinary working-tree diff does not show untracked file contents. Use status to identify them and open them directly; inspect the staged diff separately.
Does this suite verify a checkout website?
No. It checks a local pure function. Form parsing, units, display, persistence and other application boundaries require their own evidence.
When you have finished
Get your certificate of completion
Type your name and download a certificate for this workbook as a PDF, ready to print or to add to LinkedIn. It is made on your own device, so your name is never sent to us. It is a self-declared certificate, not an accredited qualification.
Learn the language
Key terms
- Contract
- The inputs, outcomes and errors a caller may rely on.
- Oracle
- The independent reason an expected test result is correct.
- Assertion
- A check comparing observed behaviour with an expectation.
- Boundary case
- An input at or near a change in the required rule.
- Subtest
- A labelled case within one unittest test method.
- Mutation
- A deliberate defect used to challenge a test suite.
6 of the workbook's 8 terms. The complete glossary is in the workbook.
Follow the evidence
Sources and checks
Facts last checked: .
Examples in this workbook were run on: Python 3.12.10 on Windows 11, standard library only. Two intended initial assertion failures; ten final test methods pass. Three isolated mutations are caught, and 106 subtotal cases exercise 318 direct calls. No network or model service. (2026-09-27).
These workbooks use AI assistance. See how the workbooks are made.
- Unit testing framework, Python 3.12Python documentation
- Built-in types and Boolean valuesPython documentation
- Inspecting working-tree and staged differencesGit documentation
- Test Driven DevelopmentMartin Fowler
Created by Mickarle Wagstaff-Irons - Micky Irons with the Mickai team. Published by Mickai LTD. Last updated 27 September 2026.
NextKeep going
Where to go next
Recommended for you
Git for AI builders: commits, branches and pull requests
Learn Git from your first commit to a reviewed pull request, with safe habits for AI-written code: read every diff, commit small, undo safely and keep secrets out.
Recommended for you
Design a repeatable evaluation suite for an AI application
Build an offline AI evaluation harness with test cases, a rubric, regression comparisons, uncertainty checks and a reproducible release record.