inference · Level 4
Model routing, caching and graceful fallbacks
Reuse only equivalent answers, spend within a declared budget and make failure visible.

Start with the essentials
The short answer
A model router chooses an eligible model for a request. A response cache can reuse an earlier answer only when the request context and access scope are equivalent. A fallback makes another eligible attempt after a classified transient failure, within the remaining budget. This course implements those contracts with scripted providers and invented costs; its results demonstrate policy mechanics, not real model quality or provider savings.
What you will learn
- Distinguish complete-response caching from provider prompt-prefix caching.
- Define an exact cache key that includes scope, context and policy/model versions.
- Implement completion-based TTL, least-recently-used eviction and cache bypass.
- Bound fallback attempts and cost while keeping refusals terminal.
- Compare routing policies with explicit quality and request denominators.
- Reproduce tests and specify the missing controls for a real provider adapter.
Who it is for
Builders who have hosted a chatbot and want a small, inspectable routing policy before adding live providers.
Before you start
- Read Python functions, dataclasses, dictionaries and unit tests.
- Understand chatbot requests, model outputs and server-side identity.
- Use Python 3.12 in an empty folder; no model, API account, GPU or network is required.
Read a sample · Chapter 01 of 06
Define the answer contract before the route
A cheaper call is useful only when the resulting answer still meets the task.
Imagine a fictional learning service that handles short lookups and longer explanations. We want to avoid paying for an expensive route when a simpler one is adequate, while preserving an explicit standard for an acceptable answer. The policy needs evidence about task types and outcomes. A short prompt is not necessarily easy, and a long prompt is not necessarily difficult. This lab therefore uses a trusted task label instead of pretending that prompt length measures reasoning difficulty.
The experiment is deliberately synthetic. FAST costs one invented unit and takes ten simulated milliseconds. STRONG costs four units and takes thirty. These names are fixture labels, not commercial products or claims about model sizes. A scripted function supplies answers and failures. The simulator does not sleep, call a network, measure execution latency or estimate tokens. Its clock advances by the assigned duration of each attempted model call.
Scroll sideways to see every column.
| Decision | Contract in this lab |
|---|---|
| Eligible route | Simple tasks may start with FAST; complex tasks start with STRONG. |
| Reused answer | Exact request, scope, context version, selected model, policy and revision must match. |
| Another attempt | Only a transient FAST failure in route mode may try STRONG. |
| Stop | Refusal, exhausted eligible routes or insufficient remaining cost budget returns a visible status. |
Create router_lab.py and append its sections in order. The first section below defines immutable request and model records and an integer clock. Request.validate rejects empty identifiers and unsupported task labels before a lookup or call. The dataclass is an internal contract, not an HTTP request parser: the caller must construct a Request after authentication, authorisation and input-size checks. Neither a type annotation nor a scope string performs those checks.
"""Serial, synthetic routing lab. No provider calls or real billing."""
from collections import OrderedDict
from dataclasses import dataclass
def natural(value, name, minimum=0):
if type(value) is not int or value < minimum:
raise ValueError(name)
@dataclass(frozen=True)
class Request:
scope: str
prompt: str
context_version: str = 'docs-v1'
task: str = 'simple'
cacheable: bool = True
def validate(self):
for value in (self.scope, self.prompt, self.context_version):
if type(value) is not str or not value.strip():
raise ValueError('empty or non-string request field')
if self.task not in ('simple', 'complex'):
raise ValueError('task')
if type(self.cacheable) is not bool:
raise ValueError('cacheable')@dataclass(frozen=True)
class Model:
name: str
units: int
duration: int
FAST = Model('fast-v1', 1, 10)
STRONG = Model('strong-v1', 4, 30)
class Clock:
def __init__(self):
self.now = 0
def advance(self, milliseconds):
natural(milliseconds, 'milliseconds')
self.now += millisecondsTreat scope as an opaque identifier for a permitted answer-sharing group, including the relevant permission generation. It could represent one user within a tenant and a current authorisation version. A client must not choose another person's scope. Changing permissions requires a new scope or explicit invalidation before looking up an answer. The lab checks equality of supplied scopes; it does not prove that they were issued correctly.
context_version identifies the exact information and instruction context affecting the answer. In this fixture it is a manually supplied string. A real adapter must derive and maintain a version for the actual system instructions, retrieved documents, tool schema, output format and generation settings. Omitting one changing input can make two different tasks look equivalent. Keep raw user text exact here: even trailing whitespace forms a different request.
The three policies are strong, fast and route. The first two always choose their named fixture and never fall back. The route policy uses the task label, then permits one upward fallback for a simple task after a transient failure. It never downgrades a complex task to FAST. Eligibility is an application decision that should be established before availability or price chooses among candidates.
Try it yourself · Activity 01
15 minWrite an eligibility rule
A lookup can use either fixture, but an explanation must return the scripted reasoned answer.
- State which fixture is eligible for each task.
- Explain why a user-supplied simple label is not adequate classification.
- Name the state that must change when answer-sharing permissions change.
Which parts of your rule would need measured model-quality evidence instead of the fixture labels?
Worked answer
The lookup may begin with FAST and fall back to STRONG after a transient failure. The explanation must use STRONG because the fixture deliberately makes FAST incomplete for that task. The host supplies the task class after its own checks. A permission change must alter the trusted scope generation or invalidate the relevant entries before a cache read; an unchanged string cannot express revoked access.
Read a sample · Chapter 02 of 06
Cache a complete answer only within its scope
Freshness, equivalence and eviction solve different problems.
A complete-response cache returns stored answer text without generating a new answer. Provider prompt caching reuses processing of an eligible repeated prompt prefix while generation continues; it does not mean that your application has stored and replayed a final answer. OpenAI's prompt-caching guide describes prefix reuse. Its eligibility and retention are provider-specific. This lab implements only the application response cache and makes no provider token-discount claim.
RFC 9111 describes HTTP cache freshness, reuse and restrictions around shared and authenticated responses. It is useful background for asking whether a response may be reused, but the following class is not an HTTP cache and does not implement that standard. It has no headers, validators or revalidation protocol. Do not place it in front of private HTTP responses and assume its small key replaces the required access controls.
class Cache:
def __init__(self, clock, capacity=4, ttl=100):
natural(capacity, 'capacity', 1)
natural(ttl, 'ttl', 1)
self.clock, self.capacity, self.ttl = clock, capacity, ttl
self.items = OrderedDict()
def purge(self):
for key, (expires, _) in list(self.items.items()):
if self.clock.now >= expires:
del self.items[key]
def get(self, key):
self.purge()
if key not in self.items:
return None
self.items.move_to_end(key)
return self.items[key][1]
def put(self, key, text):
self.purge()
self.items[key] = (self.clock.now + self.ttl, text)
self.items.move_to_end(key)
while len(self.items) > self.capacity:
self.items.popitem(last=False)The cache stores an expiry timestamp and text for each key. put runs only after the first model attempt has completed successfully, so expiry is completion time plus TTL. get removes entries when current time is greater than or equal to expiry. An entry is usable just before that boundary and unusable at it. A hit changes recency but does not extend the original expiry. This is a fixed lifetime, not a sliding one.
For example, a FAST answer started at time zero finishes at ten. With TTL 100 it expires at 110, not at 100. A hit at 109 succeeds, while a read at 110 misses. If repeated reads restarted the lifetime, a frequently requested answer could remain indefinitely despite the intended freshness interval. These endpoint choices are observable requirements, so the tests cover both completion time and exact expiry.
OrderedDict supplies the recency order. A successful read moves its key to the newest end. Insertion does the same, and overflowing capacity removes the oldest end. Purging expired entries happens before both operations so expired entries do not displace live ones. This simple purge scans the whole small cache. Capacity bounds the entry count only; there is no bound on the size of a prompt or answer in bytes.
The simulated clock is advanced only forwards through Clock.advance. Production elapsed durations should use a monotonic clock, whose differences are unaffected by wall-clock corrections, as described in Python's time documentation. Do not persist a process-local monotonic timestamp and assume it remains meaningful after a restart. This lab stores nothing on disk and does not implement a production clock adapter.
TTL limits reuse but does not establish factual freshness. A changed document can make an answer obsolete one millisecond after insertion. Update context_version when the source changes, or disable caching for the request. The cacheable field is a trusted host decision. False bypasses both reads and writes; it should be used for requests whose answers must not be retained or repeated under the chosen policy.
Try it yourself · Activity 02
15 minTrace expiry and eviction
Use capacity two and TTL 100. Insert A at time zero and B at time ten; read A at time twenty; insert C at time thirty.
- Which key is evicted at thirty?
- At what time does A expire despite the read?
- At time one hundred, purge and state the remaining key.
Would a longer TTL fix an omitted document version, or only make the incorrect reuse last longer?
Worked answer
Reading A makes it most recently used, so inserting C evicts B. A still expires at 100 because a read does not extend TTL. At exactly 100 the inclusive expiry check removes A, leaving C until its expiry at 130. Capacity pressure and time expiry are separate reasons to remove an entry.
Read a sample · Chapter 03 of 06
Make fallback a bounded state transition
A transient failure, a refusal and a programming error should not share a retry path.
Append the remaining router_lab.py source below. The key contains the full immutable Request plus the primary Model, policy and policy revision. The model record includes its versioned name. A policy revision can intentionally separate results after the routing or answer contract changes. The per-request spending budget is not part of the key: a valid cached answer can be served with zero new units even when that budget is zero.
@dataclass(frozen=True)
class Result:
status: str
text: str
model: str
hit: bool
attempts: int
units: int
elapsed: int
class Router:
def __init__(self, provider, clock, cache=None,
policy='route', revision='policy-v1'):
if policy not in ('route', 'fast', 'strong'):
raise ValueError('policy')
if type(revision) is not str or not revision.strip():
raise ValueError('revision')
if cache is not None and cache.clock is not clock:
raise ValueError('cache must share the clock')
self.provider, self.clock, self.cache = provider, clock, cache
self.policy, self.revision = policy, revision def run(self, request, budget=5):
request.validate()
natural(budget, 'budget')
primary = FAST
if self.policy == 'strong' or (
self.policy == 'route' and request.task == 'complex'):
primary = STRONG
key = (request, primary, self.policy, self.revision)
usable = self.cache is not None and request.cacheable
if usable:
text = self.cache.get(key)
if text is not None:
return Result('ok', text, primary.name, True, 0, 0, 0) started, spent, attempts = self.clock.now, 0, 0
models = [primary]
if self.policy == 'route' and primary == FAST:
models.append(STRONG)
for model in models:
if spent + model.units > budget:
return Result('budget', '', '', False, attempts,
spent, self.clock.now - started)
spent += model.units
attempts += 1
self.clock.advance(model.duration)
status, text = self.provider(model, request)
if status not in ('ok', 'transient', 'refused'):
raise ValueError('unknown provider status')
if type(text) is not str or (status == 'ok' and not text):
raise ValueError('invalid provider text') if status == 'ok':
if usable and attempts == 1:
self.cache.put(key, text)
return Result('ok', text, model.name, False, attempts,
spent, self.clock.now - started)
if status == 'refused':
return Result('refused', '', model.name, False,
attempts, spent, self.clock.now - started)
return Result('unavailable', '', '', False, attempts,
spent, self.clock.now - started)
def fixture(model, request):
"""Invented outputs and failures, not a model quality assessment."""
if request.prompt == 'restricted':
return 'refused', ''
if request.prompt == 'outage' and model == FAST:
return 'transient', ''
if request.task == 'complex':
return 'ok', 'reasoned' if model == STRONG else 'incomplete'
return 'ok', 'answer'run validates the request before consulting the cache. On a miss it selects a finite list of eligible models: one normally, two only for a simple request in route mode. Before each attempt it compares spent plus that model's units with the budget. The default budget of five allows a one-unit FAST failure followed by a four-unit STRONG attempt. Budget four allows the first attempt but stops before the second; the returned result still reports the one unit already spent.
Every attempted call is charged its fixed synthetic units and duration, including a transient failure or refusal. This is an accounting assumption for the fixture, not a statement about provider billing. Real prices can depend on input, output and cached tokens and on the service contract. Reserving an upper bound before starting a live call and reconciling its actual usage is a separate adapter responsibility. This program never charges money.
The provider returns one of three statuses. ok requires a non-empty string. transient permits the next eligible attempt if there is one. refused stops immediately, never trying to find a provider that will answer the same refused request. Unrecognised statuses, invalid output and unexpected provider exceptions raise errors rather than being disguised as temporary outages. A real application needs a visible outer error handler and accounting for such adapter faults; that handler is outside this lab.
Only a successful first attempt enters the cache. A successful fallback returns its actual model name but is not cached under the failed primary route. That keeps the next request free to try the recovered primary. It also means repeated outages can repeatedly pay for a failed first call. There is no circuit breaker, failure-memory window, shared retry budget or backoff here. Those are additional policies that must be designed and tested, not inferred from this two-attempt loop.
Google's discussion of cascading failures explains why retries need limits and why retries at multiple layers can multiply load. This fixture has one synchronous routing layer with no hidden retries. In a live stack, inspect the client SDK, gateway and server together. Passing two attempts to an adapter that silently retries three times does not yield a two-call system. Fallback also needs enough remaining time and spare downstream capacity.
Scroll sideways to see every column.
| Observed outcome | What the caller receives |
|---|---|
| Primary success | Answer, actual model, one attempt, charged units and simulated elapsed time. |
| Cache hit | Stored answer, primary model, zero new attempts, zero units and assumed zero lookup time. |
| Primary transient then success | Fallback answer, actual fallback model, two attempts and both charges. |
| Refusal | Explicit refused status, no answer and no fallback. |
| No remaining eligible route | Unavailable status, with attempted work still counted. |
| Insufficient budget | Budget status; no new attempt is started. |
Try it yourself · Activity 03
20 minAccount for a failed first call
A simple outage request has budget four. FAST returns transient and STRONG would cost four more units.
- Predict status, attempted model count, cost units and elapsed time.
- Repeat the prediction with budget five.
- Explain why the successful fallback is not cached.
What could repeated fallback do to a second provider already near its capacity limit?
Worked answer
With budget four, the result is budget, one attempt, one unit and ten simulated milliseconds; starting STRONG would bring total spending to five. With budget five, STRONG succeeds and the result is ok, two attempts, five units and forty milliseconds. The fallback answer is returned but not cached, so a future request can try the primary again. Neither outcome represents actual provider charges or timeout enforcement.
Read a sample · Chapter 04 of 06
Compare policies on the same complete trace
Keep cost, answer acceptance and refusals on declared denominators.
Save demo.py below beside router_lab.py and run python demo.py. The six requests are two identical simple questions for reader A, one complex question for A, the same simple question for reader B, a temporary FAST outage and one deliberately refused request. The scripted answer for a complex task is reasoned only on STRONG. These labels define the fixture; they are not a learned quality judge or a benchmark dataset.
from router_lab import Cache, Clock, Request, Router, fixture
TRACE = [
(Request('reader-a', 'sum'), 'answer'),
(Request('reader-a', 'sum'), 'answer'),
(Request('reader-a', 'why', task='complex'), 'reasoned'),
(Request('reader-b', 'sum'), 'answer'),
(Request('reader-a', 'outage'), 'answer'),
(Request('reader-a', 'restricted'), None),
]def evaluate(policy, cached):
clock = Clock()
cache = Cache(clock) if cached else None
router = Router(fixture, clock, cache, policy=policy)
results = [router.run(request) for request, _ in TRACE]
accepted = sum(r.status == 'ok' and r.text == expected
for r, (_, expected) in zip(results, TRACE)
if expected is not None)
refusals = sum(r.status == 'refused' for r in results)
print(f'{policy} cache={int(cached)}: '
f'units={sum(r.units for r in results)} '
f'ms={sum(r.elapsed for r in results)} '
f'accepted={accepted}/5 refused={refusals}/1 '
f'hits={sum(r.hit for r in results)}/6 '
f'attempts={sum(r.attempts for r in results)}')
return results
if __name__ == '__main__':
for policy, cached in [('strong', False), ('fast', False),
('route', False), ('route', True)]:
evaluate(policy, cached)strong cache=0: units=24 ms=180 accepted=5/5 refused=1/1 hits=0/6 attempts=6
fast cache=0: units=6 ms=60 accepted=3/5 refused=1/1 hits=0/6 attempts=6
route cache=0: units=13 ms=110 accepted=5/5 refused=1/1 hits=0/6 attempts=7
route cache=1: units=12 ms=100 accepted=5/5 refused=1/1 hits=1/6 attempts=6The all-STRONG baseline attempts six calls, spending 24 units and advancing time 180 milliseconds. Five answerable requests meet their exact expected strings and the restricted request is refused. All-FAST costs six units and sixty milliseconds, but only three of five answerable requests pass: its complex answer is incomplete and its outage has no fallback. A lower bill alone hides those two failures.
Routing without caching uses FAST for the ordinary simple requests, STRONG for the complex one and an extra STRONG attempt for the outage. This costs 13 units across seven attempts and takes 110 simulated milliseconds. It passes all five answerable cases and preserves the refusal. The extra attempt is counted even though there are only six user requests. Confusing those two denominators would hide work amplification.
Enabling the cache saves the immediate second request for reader A. The reader-B request must miss despite identical prompt text because its scope differs. The complex answer uses a different route and key. The fallback is intentionally not stored, and refusals are not stored. The final row therefore has one hit out of six submitted requests, six attempted calls, twelve units and one hundred simulated milliseconds.
Against all-STRONG, the last row saves twelve of twenty-four invented units, or fifty percent, on this selected trace. Against the same routing policy without a cache, caching alone saves one of thirteen units, about 7.7 percent. Those are different comparisons. Neither number predicts production savings. The fixture intentionally includes one repeat and favourable task labels; changing the traffic, expiry or quality outcomes changes the result.
A hit is assigned zero simulated time because this lab does not model lookup overhead. The total elapsed value is the sum of sequential invented call durations; it is neither measured wall-clock latency nor throughput under concurrent load. Do not divide the six requests by it to claim real serving capacity. Report the model of time alongside the number.
accepted uses the five answerable requests as its denominator. refused uses the one expected refusal. hits uses all six submitted requests. An exact-string acceptance check is suitable for these literal scripted outputs, not for judging open-ended explanations. A real evaluation needs a representative held-out set, an explicit rubric, error categories and a reviewer or validated checking method. Keep appropriately refused requests visible without falsely labelling them successful answers.
Try it yourself · Activity 04
20 minSeparate the savings mechanisms
Use the four exact output rows to explain the final result.
- Compute routing-plus-cache savings against all-STRONG.
- Compute caching-only savings against uncached routing.
- Explain the different acceptance, refusal, hit and attempt counts.
Which traffic changes would eliminate the one cache hit without changing the routing rule?
Worked answer
The combined reduction is (24 - 12) / 24 = 50 percent of invented units. Caching alone saves (13 - 12) / 13, about 7.7 percent. Five answerable requests are checked for expected text; one restricted request is checked for refusal; one of six submitted requests is a hit. The uncached routing row has seven attempts because one user request needed two calls. These are fixture comparisons, not a forecast of cost or model quality.
Read a sample · Chapter 05 of 06
Test the boundaries that attractive averages hide
A fast demo is incomplete if it reuses the wrong reader's answer or retries a refusal.
Save test_router.py below and run python -m unittest -v test_router.py. Fourteen test methods cover route eligibility, zero-cost hits, request dimensions, version separation, exact expiry, non-sliding lifetime, recency eviction, bypass, fallback, refusal, spending and invalid input. They run without downloading a model or requiring an account. The source here is the same source executed for the published evidence.
import unittest
from dataclasses import replace
from router_lab import Cache, Clock, Request, Router, FAST, STRONG, fixture
class RoutingTests(unittest.TestCase):
def setUp(self):
self.clock = Clock()
self.cache = Cache(self.clock, capacity=2, ttl=100)
self.calls = []
def provider(model, request):
self.calls.append((model.name, request))
return fixture(model, request)
self.router = Router(provider, self.clock, self.cache)
self.request = Request('a', 'sum')
def test_routing_and_quality(self):
self.assertEqual(self.router.run(self.request).model, FAST.name)
complex_request = replace(self.request, task='complex')
result = self.router.run(complex_request)
self.assertEqual((result.model, result.text), (STRONG.name, 'reasoned')) def test_hit_no_new_call_or_units(self):
self.router.run(self.request)
hit = self.router.run(self.request, budget=0)
self.assertEqual((hit.hit, hit.attempts, hit.units, hit.elapsed),
(True, 0, 0, 0))
self.assertEqual(len(self.calls), 1)
def test_every_request_dimension_separates(self):
for change in ({'scope':'b'}, {'prompt':'sum '},
{'context_version':'docs-v2'}, {'task':'complex'}):
with self.subTest(change=change):
self.setUp()
self.router.run(self.request)
changed = replace(self.request, **change)
self.assertFalse(self.router.run(changed).hit)
self.assertEqual(len(self.calls), 2)
def test_policy_revision_and_model_separate(self):
self.router.run(self.request)
for policy, revision in [('route','v2'), ('strong','policy-v1')]:
router = Router(fixture, self.clock, self.cache, policy, revision)
self.assertFalse(router.run(self.request).hit) def test_expiry_starts_at_completion_and_is_exclusive(self):
self.router.run(self.request)
self.clock.advance(99)
self.assertTrue(self.router.run(self.request).hit)
self.clock.advance(1)
self.assertFalse(self.router.run(self.request).hit)
self.assertEqual(len(self.calls), 2)
def test_hit_does_not_extend_ttl(self):
self.cache.put('a', 'A')
self.clock.advance(90)
self.assertEqual(self.cache.get('a'), 'A')
self.clock.advance(10)
self.assertIsNone(self.cache.get('a'))
def test_lru_touch_and_expired_entries(self):
self.cache.put('a','A')
self.cache.put('b','B')
self.cache.get('a')
self.cache.put('c','C')
self.assertEqual(list(self.cache.items), ['a','c'])
self.clock.advance(100)
self.cache.put('d','D')
self.assertEqual(list(self.cache.items), ['d']) def test_non_cacheable_bypasses_reads_and_writes(self):
request = replace(self.request, cacheable=False)
self.router.run(request)
self.assertFalse(self.router.run(request).hit)
self.assertEqual(len(self.calls), 2)
self.assertFalse(self.cache.items)
def test_transient_fallback_is_not_cached(self):
request = replace(self.request, prompt='outage')
result = self.router.run(request)
self.assertEqual((result.status, result.model, result.attempts,
result.units, result.elapsed),
('ok', STRONG.name, 2, 5, 40))
self.assertFalse(self.router.run(request).hit)
self.assertEqual(len(self.calls), 4)
def test_refusal_is_terminal_and_not_cached(self):
request = replace(self.request, prompt='restricted')
self.assertEqual(self.router.run(request).status, 'refused')
self.assertEqual(len(self.calls), 1)
self.assertEqual(self.router.run(request).status, 'refused')
self.assertEqual(len(self.calls), 2) def test_budget_checked_before_each_attempt(self):
self.assertEqual(self.router.run(self.request, 0).status, 'budget')
self.assertFalse(self.calls)
request = replace(self.request, prompt='outage')
result = self.router.run(request, 4)
self.assertEqual((result.status, result.units, result.attempts),
('budget', 1, 1))
self.assertEqual(len(self.calls), 1)
def test_exhaustion_and_no_quality_downgrade(self):
def failed(model, request):
self.calls.append(model)
return 'transient', ''
router = Router(failed, self.clock)
result = router.run(self.request)
self.assertEqual((result.status, len(self.calls)), ('unavailable', 2))
self.calls.clear()
result = router.run(replace(self.request, task='complex'))
self.assertEqual((result.status, self.calls), ('unavailable', [STRONG])) def test_invalid_input_and_programming_errors(self):
for change in ({'scope':''}, {'prompt':7}, {'task':'unknown'},
{'context_version':' '}, {'cacheable':1}):
with self.assertRaises(ValueError):
self.router.run(replace(self.request, **change))
for budget in (-1, True, 1.5):
with self.assertRaises(ValueError):
self.router.run(self.request, budget)
self.assertFalse(self.calls)
for status, text in [('wrong',''), ('ok',''), ('ok',None)]:
router = Router(lambda m,r: (status,text), self.clock)
with self.assertRaises(ValueError):
router.run(self.request)
def test_cache_configuration_rejected(self):
for capacity, ttl in [(0,1),(1,0),(True,1),(1,1.5)]:
with self.assertRaises(ValueError):
Cache(self.clock, capacity, ttl)
with self.assertRaises(ValueError):
Router(fixture, Clock(), self.cache)
if __name__ == '__main__':
unittest.main()The scope test makes a request for one reader, then changes one request dimension at a time. Every such change must cause a new call. It is an equality-contract check rather than an authentication test: a malicious caller who can forge the trusted Request could still provide somebody else's scope. Put that distinction in a review note instead of calling a passing unit test a security audit.
The expiry test deliberately asks at completion plus TTL minus one, then exactly at expiry. The first must hit and the second must miss. The separate non-sliding test reads an item shortly before expiry and confirms that the read did not extend it. The recency test reads A after writing A and B, inserts C, then checks that B was evicted. Inspecting only final cache size would miss the wrong eviction choice.
The budget test checks both stopping before any call and stopping before the second call after spending on a failure. The refusal test requires only one provider call for each restricted request. The exhaustion test confirms that a complex request never falls down to FAST, even when STRONG fails. These observations establish what the code does for selected cases; they do not establish a reliable model classifier or sufficient capacity during a live outage.
Author review also ran four syntactically valid faulty copies in isolated temporary directories. Omitting scope from the key, accepting an entry exactly at expiry, allowing refusal to reach fallback and ignoring already-spent units each caused its targeted test to fail. Original file hashes stayed unchanged. This makes those checks demonstrably sensitive to four plausible defects, while leaving other defects and production behaviours outside the tested set.
Keep errors distinguishable in reports. An unavailable response is not a refusal; a budget stop is not a model-quality failure; a malformed adapter response is not a transient backend outage. The tests raise invalid-output errors directly. A production adapter must define a narrow mapping from real provider statuses and exceptions into the contract, and must preserve observability when it rejects unexpected responses.
Try it yourself · Activity 05
20 minIntroduce and detect one logic fault
In a disposable copy, change spent + model.units > budget to model.units > budget.
- Predict the result for an outage request with budget four.
- Run the budget test and explain the failure.
- Restore the original expression and rerun the full suite.
What additional observations would you need if the SDK could retry inside one apparent attempt?
Worked answer
The faulty comparison sees STRONG's four-unit price as affordable while forgetting the one unit already spent. It starts the fallback and spends five despite a budget of four. test_budget_checked_before_each_attempt expects budget, one unit and one attempt, so it fails. Restoring the cumulative comparison returns the expected stop. This is a valid-Python logic error, not a syntax failure.
Read a sample · Chapter 06 of 06
Hand over a policy that can be evaluated honestly
A live adapter needs measurements and controls that the scripted lab intentionally leaves open.
The reusable output is a decision contract and an executable example. Before connecting a provider, record each candidate model's exact revision, allowed data boundary, supported output contract and observed quality on the relevant tasks. Establish which routes are eligible first. A provider that cannot meet the required schema or data-handling rules is not an acceptable fallback merely because it is available.
Scroll sideways to see every column.
| Before live integration | Evidence or control to provide |
|---|---|
| Identity and context | Host-issued scope and permission version; exact instruction, source and generation configuration versions. |
| Quality | Held-out examples, explicit rubric, per-task results, refusals and unacceptable-output categories. |
| Accounting | Attempt IDs, token usage, applicable prices, reservation and actual-usage reconciliation. |
| Time and load | End-to-end deadline, per-attempt timeout, cancellation, bounded queue and downstream capacity checks. |
| Cache operation | Byte limits, eviction metrics, invalidation, retention rules and concurrent lookup behaviour. |
| Change control | A recorded policy revision, limited rollout, acceptance owner and rollback conditions. |
This cache is single-process and single-threaded. It has no locks, single-flight request coalescing, distributed coordination, encrypted storage or durable invalidation. Two concurrent misses could both call a provider in a future adaptation unless coordinated. Large request text could exhaust memory despite the entry-count limit. Keep this teaching cache out of an untrusted service boundary until those concerns have explicit implementations and tests.
A live call must respect the request's remaining end-to-end time. A model duration field does not enforce a timeout; our field simply advances an invented clock. Cancellation may also need to reach a provider to stop downstream work. Handle in-progress streams and partial output deliberately: the lab only caches complete successful strings and never demonstrates resuming or replacing a partly delivered answer.
Google's overload chapter discusses accepting work that available capacity can serve and rejecting excess work deliberately. Carry that principle into routing: moving every failed request to a second backend can overload it. A bounded per-request fallback is only one limit. The application may need admission control, a service-wide retry budget and an explicitly designed degradation path with observable recovery.
A useful degradation might be a clearly labelled unavailable result with a later retry option, or an approved static explanation for a narrow task. Do not silently substitute a low-quality answer and count it as successful. If stale information is allowed for a particular product feature, define its age, scope and label explicitly. This lab serves no expired entry and has no stale-on-error mode.
For the next experiment, use a replayable, representative request set that the service is permitted to process. Run the same set across the candidate policies, controlling model versions and cache warm-up. Record per-request status, actual route, attempts, cache state, usage and timing, plus the quality judgement. Separate cold and warm cache results; show failures and refusals alongside accepted answers. Obtain owner-agreed thresholds before deciding whether to promote a policy.
The published evidence covers the exact local lab, fourteen tests, four deterministic demo lines and four detected defects. It does not cover actual model quality, provider billing, authentication, concurrent races, network timeout enforcement, sustained load or production deployment. A handover should name those remaining checks and the person or team responsible for acceptance instead of turning the fixture's fifty-percent result into a marketing claim.
Try it yourself · Activity 06
15 minWrite a live-adapter acceptance note
Specify the first controlled evaluation before connecting this design to a real chatbot.
- Name the model/configuration snapshots and permitted evaluation set.
- List the accounting, quality, cache and time fields to capture.
- State the acceptance owner, rollback trigger and unimplemented controls.
Can another builder tell which results were executed locally and which still require a real service?
Worked answer
The note should identify exact model and policy revisions, host-issued scope/context rules and a permitted held-out request set. Capture actual model, attempts, usage, cache hits, end-to-end timing, output quality, refusals and errors for every request. The service owner sets acceptance thresholds and rollback triggers. Mark authentication, real timeouts, cancellation, concurrency, byte limits, outage load and provider-specific billing as unverified until their own checks pass.
Keep learning
The complete workbook
Six chapters teach a complete three-file Python lab, scoped TTL and LRU caching, explicit failure categories, budget checks, exact experiment output and reproducible tests. A worked workbook separates local implementation evidence from the measurements needed before a live service rollout.
- 01Define the answer contract before the routeRead here · 1 exercise
A cheaper call is useful only when the resulting answer still meets the task.
- 02Cache a complete answer only within its scopeRead here · 1 exercise
Freshness, equivalence and eviction solve different problems.
- 03Make fallback a bounded state transitionRead here · 1 exercise
A transient failure, a refusal and a programming error should not share a retry path.
- 04Compare policies on the same complete traceRead here · 1 exercise
Keep cost, answer acceptance and refusals on declared denominators.
- 05Test the boundaries that attractive averages hideRead here · 1 exercise
A fast demo is incomplete if it reuses the wrong reader's answer or retries a refusal.
- 06Hand over a policy that can be evaluated honestlyRead here · 1 exercise
A live adapter needs measurements and controls that the scripted lab intentionally leaves open.
Also inside: a 8-point checklist, a glossary of 10 terms and 10 questions and answers to test yourself. 6 hands-on exercises, each with a worked answer at the back where the workbook gives one.
No login, no card, no account. Before the download we ask you to follow Mickai (two quick links). Free to download and use for personal learning, study groups and inside your own team. Please do not resell the workbooks or republish them as your own. Link people to trust-agent.ai instead.
Test yourself
Questions and answers
Are FAST and STRONG real model benchmarks?
No. They are scripted fixtures with invented costs, durations and output quality labels. No live model was called.
Can a client supply its own sharing scope?
The lab assumes a trusted host constructs Request. A live service must authenticate and authorise the caller before issuing the scope and checking the cache.
Does a cache hit extend the entry lifetime?
No. A hit changes recency for eviction but expiry stays completion time plus TTL. The exact expiry timestamp is already too late for reuse.
Does prompt caching replay a complete answer?
The provider prefix reuse discussed here continues generating an answer. The application response cache in this lab instead returns complete stored text.
Why is fallback success not cached?
It is returned with its actual model name but is not stored under a failed primary route. A later request can try the recovered primary again.
Does a refusal trigger another model call?
No. A refusal is terminal. Only a transient FAST failure in route mode may try STRONG, within the remaining cumulative budget.
How much did caching alone save in the fixture?
One of thirteen invented units, about 7.7 percent, compared with the same routing policy without caching. This is not a provider savings forecast.
Why can attempts exceed request count?
One user request can make a primary call and a fallback call. Uncached routing submits six requests but attempts seven scripted calls.
What do the four faulty copies prove?
The targeted tests detect omitted scope, an inclusive expiry bug, refusal fallback and forgetting already-spent units. They do not prove exhaustive correctness or production reliability.
Does the lab enforce real network timeouts?
No. Its clock advances by invented durations. A real adapter needs an end-to-end deadline, per-attempt timeout, cancellation and separate verification.
When you have finished
Get your certificate of completion
Type your name and download a certificate for this workbook as a PDF, ready to print or to add to LinkedIn. It is made on your own device, so your name is never sent to us. It is a self-declared certificate, not an accredited qualification.
Learn the language
Key terms
- Model routing
- Selecting an eligible model for a request using an explicit policy and task contract.
- Response cache
- A store of complete answers that may be reused for requests meeting a declared equivalence rule.
- Cache key
- The values used to decide whether a stored result belongs to the current request.
- Scope
- A host-issued identity for a group of requests permitted to share an answer under current permissions.
- Time to live
- The interval after insertion or completion during which an entry may be reused under the chosen freshness policy.
- Least recently used
- An eviction rule that removes the entry whose last successful use is oldest.
6 of the workbook's 10 terms. The complete glossary is in the workbook.
Follow the evidence
Sources and checks
Facts last checked: .
Examples in this workbook were run on: Python 3.12.10 on Windows, standard library only: fourteen tests, four exact experiment outputs and four detected logic faults. Scripted providers, invented durations and cost units, and a six-request fixture. No live model, billing, real timeout, concurrent cache, authentication system or production load test was executed. (2026-09-27).
These workbooks use AI assistance. See how the workbooks are made.
- HTTP caching and response reuseIETF RFC 9111
- Prompt cachingOpenAI documentation
- Ordered dictionariesPython Software Foundation
- Monotonic clocksPython Software Foundation
- Addressing cascading failuresGoogle SRE
- Handling overloadGoogle SRE
Created by Mickarle Wagstaff-Irons - Micky Irons with the Mickai team. Published by Mickai LTD. Last updated 27 September 2026.
NextKeep going
Where to go next
Recommended for you
Throughput, latency and batching: capacity planning for a model server
Build a tested batch queue simulator, compare burst and sparse arrivals, calculate timing metrics and design a measured model-server experiment.
Recommended for you
Design a repeatable evaluation suite for an AI application
Build an offline AI evaluation harness with test cases, a rubric, regression comparisons, uncertainty checks and a reproducible release record.