inference · Level 4

Throughput, latency and batching: capacity planning for a model server

Measure the waiting, the work and the finished requests before choosing a serving configuration.

By Mickarle Wagstaff-Irons - Micky Irons

  • Level 4Frontier
  • 180 min
  • 6 chapters
  • Free PDF, no account
The Streaminference / 04

Start with the essentials

The short answer

Throughput counts completed work per time window; latency measures how long each request waits for its result. Batching can share work across requests but can also make an early request wait for companions. This course uses a finite single-server simulator with an explicitly assumed service formula to expose that trade-off. Its timings are teaching examples, not hardware benchmarks or a recommendation for a production batch size.

What you will learn

  • Define response latency, queue delay, throughput and their measurement boundaries.
  • Trace a FIFO batch scheduler with an explicit capacity and waiting deadline.
  • Calculate exact cohort metrics and explain the chosen percentile method.
  • Reproduce burst and sparse arrival results using three complete Python files.
  • Challenge the scheduler with an independent reference and targeted logic faults.
  • Design a measured serving experiment with workload, errors, quality and saturation evidence.

Who it is for

Builders who understand inference engines and want to reason about serving trade-offs before tuning or buying capacity.

Before you start

  • Read Python functions, dataclasses, lists and unit tests.
  • Understand model inference, input and output tokens, and the difference between prefill and decode.
  • Use Python 3.12 in a fresh folder; this lab needs no model, GPU, account or network.

Read a sample · Chapter 01 of 06

01

Choose the question before measuring speed

A capacity claim needs a workload, a time boundary and an acceptable result.

Imagine a fictional learning service where several readers ask for a short explanation at once. The builder wants to know whether requests should start immediately or wait briefly to share a batch. A fast answer for one reader and many answers per second are different objectives. A configuration can improve the latter while making an individual reader wait longer. The first task is to define whose experience matters and what a successful response means.

Our lab uses requests with an arrival timestamp and a name. It does not contain a language model. Each batch takes an assumed 8 milliseconds of fixed work plus 2 milliseconds for every request in that batch. Those numbers are invented to make the arithmetic visible. Running the Python program measures its logic, not the service time of a GPU, CPU, inference engine or hosted API. Never put its output on a hardware comparison chart.

Scroll sideways to see every column.

Choose the question before measuring speed · Table 1
QuantityBoundary and unit in this workbook
Queue delayBatch start minus request arrival, in simulated milliseconds.
Service durationBatch finish minus batch start, in simulated milliseconds.
Response latencyRequest finish minus arrival; queue delay plus service duration.
ThroughputCompleted requests divided by the full observed cohort window, in requests per second.
ConcurrencyRequests present at a moment, including waiting and being served; not registered users.

For a streaming language model, the first useful output and the full response arrive at different times. NVIDIA NIM benchmarking documentation distinguishes time to first token from end-to-end latency and notes that definitions can differ between tools. Record the actual timestamp boundaries your tool uses. A first-token metric cannot be reconstructed from only the completion times in our simulator.

Hugging Face's streaming guide explains returning generated output incrementally. Streaming can let the reader see output before the full response is ready; it does not by itself prove that the server completes more requests per second. Our jobs have a single completion event and no token events. The lab therefore reports neither time to first token nor inter-token latency, and it cannot evaluate the smoothness of a streamed answer.

Write a workload contract before setting an objective. Record input lengths, output limits, arrival pattern, cancellation policy, model and engine revision, hardware, cache state and result-quality checks for a future real run. A short cached prompt and a long uncached document may consume different resources. Counting them both as one request is legitimate only when the mix is described alongside the number.

Use a service objective that includes failures. A proposed experiment might require a chosen proportion of valid responses to finish within an owner-agreed time, with a separate limit on rejected, timed-out or incorrect requests. Do not silently discard failed attempts to make the successful tail look faster. The educational fixtures here all complete; error handling is a missing dimension of their performance model.

Save queue_lab.py, demo.py and test_queue.py in one fresh folder as you reach them. The first file continues across chapters 2 and 3. Append its labelled code blocks in order, preserving indentation and adding one blank line between blocks. The code was executed with Python 3.12.10 on Windows using only the standard library; this is the tested runtime, not a claim about the newest Python version.

Try it yourself · Activity 01

15 min

Define a usable performance claim

Rewrite “the server handles 100 users” as an experiment question.

  1. Distinguish registered users, in-flight requests and requests per second.
  2. Name two workload properties that affect the comparison.
  3. Add a latency objective and a failure-accounting rule without inventing measured results.

Could two teams reproduce the same workload from the words in your claim?

Worked answer

A testable question is: under a specified arrival trace, prompt/output length mix and fixed engine revision, what completed request rate can the service sustain while meeting an agreed response-latency objective? Registered users are not simultaneous work. Count every submitted request as completed, failed, rejected, cancelled or unfinished, and report these outcomes separately. No capacity number is established until that experiment is run.

Read a sample · Chapter 02 of 06

02

Model the queue and the batch deadline

Make time advancement explicit so that waiting cannot disappear from the result.

Start queue_lab.py with the following implementation. A Job is one named arrival. Done records the original arrival, batch start, batch finish and batch number. Metrics defines the later summary. Frozen dataclasses discourage accidental edits to observations; they are not a security boundary. The simulator receives a tuple so its finite workload is explicit and does not pull from a live stream.

python · 26 lines
"""Integer-time teaching simulator, not measured model-server performance."""
from dataclasses import dataclass
from fractions import Fraction

@dataclass(frozen=True)
class Job:
    name: str
    arrival: int

@dataclass(frozen=True)
class Done:
    name: str
    arrival: int
    start: int
    finish: int
    batch: int

@dataclass(frozen=True)
class Metrics:
    horizon: int
    requests_per_second: Fraction
    mean_wait: Fraction
    mean_latency: Fraction
    p95_latency: int
    average_in_system: Fraction
    busy_fraction: Fraction

python · 20 lines
def bounded_int(value, low, high):
    if type(value) is not int or not low <= value <= high:
        raise ValueError("integer outside the lab range")

def simulate(jobs, max_batch, wait_ms, fixed_ms=8, per_item_ms=2):
    bounded_int(max_batch, 1, 64)
    bounded_int(wait_ms, 0, 10000)
    bounded_int(fixed_ms, 0, 10000)
    bounded_int(per_item_ms, 1, 10000)
    if type(jobs) is not tuple or not 1 <= len(jobs) <= 256:
        raise ValueError("supply a tuple of 1 to 256 jobs")
    names = set()
    for job in jobs:
        if type(job) is not Job:
            raise ValueError("Job required")
        if (type(job.name) is not str or not job.name.strip()
                or len(job.name) > 40 or job.name in names):
            raise ValueError("unique nonempty job name required")
        bounded_int(job.arrival, 0, 1000000)
        names.add(job.name)

python · 24 lines
    ordered = sorted(jobs, key=lambda job: job.arrival)
    now = 0
    cursor = 0
    batch = 0
    done = []
    while cursor < len(ordered):
        first = ordered[cursor]
        now = max(now, first.arrival)
        deadline = first.arrival + wait_ms
        fill_index = cursor + max_batch - 1
        fill_time = (ordered[fill_index].arrival
                     if fill_index < len(ordered) else deadline)
        start = max(now, min(deadline, fill_time))
        end = cursor
        while (end < len(ordered) and end - cursor < max_batch
               and ordered[end].arrival <= start):
            end += 1
        finish = start + fixed_ms + per_item_ms * (end - cursor)
        batch += 1
        done.extend(Done(job.name, job.arrival, start, finish, batch)
                    for job in ordered[cursor:end])
        cursor = end
        now = finish
    return tuple(done)

Input checks happen before simulation. Supply 1 to 256 jobs with unique, non-empty names of at most 40 Python characters and integer arrival times from 0 to 1,000,000 milliseconds. Capacity is 1 to 64 and batching delay is 0 to 10,000 milliseconds. Fixed cost may be zero; per-item cost must be positive, which ensures positive service duration and a nonzero observation window. Exact integer checks refuse booleans and fractional timestamps.

The requests are sorted by arrival without modifying the caller's tuple. Equal arrival times retain their input order. This gives a reproducible FIFO tie rule. The simulation has one server, so batches never overlap in service. A later arrival cannot pass an earlier queued request, and every accepted job appears exactly once in the result. There are no priorities, preemption, rejection or cancellation events.

The oldest unserved arrival starts the batching clock. Its deadline is arrival plus wait_ms. If enough requests to fill a batch have arrived sooner, dispatch can happen sooner. If the server is still busy, dispatch waits for that work to finish. The formula takes the later of server availability and the earlier of the deadline or full-batch arrival. A batching allowance therefore does not mean that total queue delay stays below that allowance.

At dispatch, include only requests that have arrived by the chosen start time, including those arriving exactly at that instant. Stop at max_batch. All members then finish together after fixed_ms plus per_item_ms times the actual group size. A partly filled batch pays the same fixed cost but has fewer members to share it. This is the entire service model; token lengths, memory use and compute contention are absent.

NVIDIA Triton documents dynamic batching with an optional queue delay and separate inflight batching for language-model execution. Our fixed whole-request batches illustrate waiting to form a group. They do not implement continuous batching, where active requests can change between generation iterations. A request arriving during one of our batches must wait until that whole batch finishes. Do not copy these mechanics into a token scheduler as if they modelled its internals.

Consider arrivals at 0, 1 and 2 with capacity 2 and delay 3. The first two requests fill the batch at time 1 and finish at 13. The third request's own deadline was 5, so it starts immediately at 13 and finishes at 23. Restarting its delay when the server becomes free would add unnecessary waiting. The test suite contains this exact case because a plausible-looking scheduler can make that mistake.

Try it yourself · Activity 02

15 min

Trace a busy server

Use arrivals 0, 1 and 2, capacity 2, delay 3 and service time 8 + 2n.

  1. Find the first batch start and finish.
  2. Calculate the third request's deadline and actual start.
  3. Separate its queue delay from its service duration and full latency.

Which timestamp would expose a bug that restarts the delay after every completed batch?

Worked answer

The first batch starts at 1 and finishes at 13. Request three arrived at 2, so its batching deadline was 5; the busy server makes it start at 13. Its queue delay is 11, service duration 10 and full latency 21 milliseconds. The 3-millisecond batching allowance does not cap total queueing behind active work.

Read a sample · Chapter 03 of 06

03

Calculate metrics over one declared window

Exact arithmetic makes accounting checkable; it cannot make a small sample representative.

Complete queue_lab.py with these functions. metrics accepts only the unmodified non-empty output of simulate. It is not a parser or validator for logs from a real service. An application importing arbitrary observations would first need to check unique request IDs, timestamp order, batch consistency, missing completions and units. Keeping that boundary explicit prevents a convenient summary helper from being mistaken for an ingestion contract.

python · 25 lines
def nearest_rank(values, percentile):
    bounded_int(percentile, 1, 100)
    if not values or any(type(v) is not int or v < 0 for v in values):
        raise ValueError("nonempty nonnegative integer observations required")
    ordered = sorted(values)
    rank = (len(ordered) * percentile + 99) // 100
    return ordered[rank - 1]

def metrics(done):
    """Accept only an unmodified nonempty result from simulate()."""
    if not done:
        raise ValueError("no observations")
    horizon = max(d.finish for d in done) - min(d.arrival for d in done)
    latency = [d.finish - d.arrival for d in done]
    wait = [d.start - d.arrival for d in done]
    batches = {d.batch: d.finish - d.start for d in done}
    return Metrics(
        horizon=horizon,
        requests_per_second=Fraction(1000 * len(done), horizon),
        mean_wait=Fraction(sum(wait), len(done)),
        mean_latency=Fraction(sum(latency), len(done)),
        p95_latency=nearest_rank(latency, 95),
        average_in_system=Fraction(sum(latency), horizon),
        busy_fraction=Fraction(sum(batches.values()), horizon),
    )

The observation window begins at the first arrival and ends at the final completion. It includes initial batching delay, idle gaps and the final drain after the last arrival. Subtracting first completion instead would omit early work. Subtracting last arrival would omit draining work. A finite burst rate measured with this full window is useful for comparing these fixtures, but it is not an estimate of sustained operating capacity.

Requests per second is 1,000 times the completed count divided by the window in milliseconds. The mean uses every request's latency, including time in the queue. Python Fraction keeps ratios exact until demo.py formats them for display. Constructing a rational from integer counts avoids a floating-point rounding decision during accounting. Decimal formatting is presentation; the exact underlying fraction remains available for assertions.

The percentile method is nearest rank: sort N observations, choose rank ceiling(N times p divided by 100), then convert the one-based rank to a zero-based index. For four samples, p95 selects rank 4, the largest value. It is not evidence about the 95th percentile of an unknown production distribution. Other tools may interpolate between observations; record the method instead of treating every field named p95 as interchangeable.

Scroll sideways to see every column.

Calculate metrics over one declared window · Table 2
Four serial burst requestsObserved latency in milliseconds
r0: arrival 0, finish 1010
r1: arrival 1, finish 2019
r2: arrival 2, finish 3028
r3: arrival 3, finish 4037
Mean; nearest-rank p9523.5; 37
Full window; completion rate40 milliseconds; 100 requests/second

Average occupancy counts both waiting and active requests. Each request contributes its full latency to the area under the number-in-system curve. Dividing that area by the same full window gives average_in_system. For this completely observed finite cohort, the identity average occupancy = requests per millisecond times mean latency follows directly from the sums. No steady-state assumption is needed for that accounting identity; it is not a forecast for a growing production queue.

For the serial burst, total request-time is 94 milliseconds over a 40-millisecond window, so average occupancy is 2.35. Multiplying 100 requests per second by 0.0235 seconds gives the same value. Multiplying by 23.5 without converting units would be wrong by a factor of 1,000. The test suite also counts occupancy at each integer tick rather than relying only on the sum-of-latencies formula.

Busy fraction sums each batch's service duration once and divides by the window. Summing it once per request would count simultaneous batch work multiple times. This is utilisation of our abstract single server, not GPU utilisation, memory saturation or kernel occupancy. A long idle gap can lower this fraction while the few submitted requests still have acceptable latency. Never interpret one such number without the arrival trace.

Try it yourself · Activity 03

20 min

Reconcile the numbers

Recalculate the four serial burst latencies from the table.

  1. Find the mean, nearest-rank p95 and full observation window.
  2. Compute completed requests per second.
  3. Calculate occupancy in two equivalent ways and state what the result does not establish.

What must change if the measurement window ends while some requests are still unfinished?

Worked answer

The latencies sum to 94, so their mean is 94/4 = 23.5 milliseconds and nearest-rank p95 is 37. Four completions over 40 milliseconds give 100 requests per second. Average occupancy is 94/40 = 2.35, also 100 times 0.0235. These equalities check this finite fixture; four arrivals do not establish a sustainable rate, confidence interval or production tail percentile.

Read a sample · Chapter 04 of 06

04

Compare burst arrivals with sparse demand

One batch policy can help one traffic pattern and hurt another.

Batching changes who waits.

Simulated burst timelines separate dashed queue delay from solid service. Grouping the burst reduces mean latency but slows its first request. Waiting adds latency without a benefit to sparse arrivals.
Executed observations from the published teaching simulator. All times are simulated milliseconds; batch service is the assumed formula 8 + 2n. Complete traces and metrics are provided as native text below. Open the full-size batching and latency diagram.

B is maximum batch size and W is the batching allowance in milliseconds. Requests are identified by r0 through r3. A circle marks arrival, a dashed line means waiting, and a solid bar means service; a short upright mark ends each bar at completion. Both burst timelines use the same 0-to-40 scale. Every batch member completes together. No timing was measured on a GPU or model server.

Burst arrivals at 0, 1, 2 and 3: all request times in simulated milliseconds
Configuration / requestArrivalStartFinishQueue delayFull latency
B1 W0 / r00010010
B1 W0 / r111020919
B1 W0 / r2220301828
B1 W0 / r3330402737
B4 W3 / r00319319
B4 W3 / r11319218
B4 W3 / r22319117
B4 W3 / r33319016

With B1 W0, each request takes 10 milliseconds of service and later arrivals wait behind earlier work. With B4 W3, all four start at 3 and finish at 19. The first request becomes slower: full latency rises from 10 to 19. The mean falls from 23.5 to 17.5 because later requests finish sooner. This is a trade-off for these arrivals and this assumed cost formula, not a universal batch-size recommendation.

Sparse arrivals at 0, 20, 40 and 60: all request times in simulated milliseconds
Configuration / requestArrivalStartFinishQueue delayFull latency
B4 W0 / r00010010
B4 W0 / r1202030010
B4 W0 / r2404050010
B4 W0 / r3606070010
B4 W3 / r00313313
B4 W3 / r1202333313
B4 W3 / r2404353313
B4 W3 / r3606373313

Every sparse request runs alone in both policies. W3 adds 3 milliseconds of queue delay without sharing the fixed service cost. In the graphic, the sparse bars use time since each request's own arrival, on a 0-to-15 scale. They describe all four sparse requests: latency is 10 with W0 and 13 with W3.

Four complete cohorts: full windows and latency in milliseconds; completion rate in requests/second
ConfigurationFull window Mean latencyNearest-rank p95 Requests/second
Burst B1 W04023.537100.000
Burst B4 W31917.519210.526
Sparse B4 W07010.01057.143
Sparse B4 W37313.01354.795

The window runs from first arrival to last completion, including idle gaps and final drain. Rates are formatted to three decimal places from exact Fraction values. Each cohort has only four observations, so nearest-rank p95 is its largest latency. The first burst has latencies 10, 19, 28 and 37; the grouped burst has 19, 18, 17 and 16. These are finite-cohort results, not sustained capacity or a production tail estimate.

Separate B2 W3 fixture with arrivals 0, 1 and 2: all times in simulated milliseconds
RequestArrival Own batching deadlineStart FinishQueue delay Full latency
r003113113
r114113012
r22513231121

The first two requests fill a batch at 1 and finish at 13. Request r2 arrived at 2 with deadline 5, but the server is busy. It starts immediately when the server becomes free at 13 and finishes at 23. Its delay is not restarted: queue delay is 11, service is 10 and full latency is 21. A batching allowance does not cap waiting behind active work; a full group can also start before an individual deadline.

This is a finite FIFO model of homogeneous requests with an invented service cost. It includes no token events, continuous batching, variable prompt/output lengths, memory pressure, failures, cancellations or admission control. The reference checks implementation under these assumptions, not their fit to real hardware. No GPU benchmark, production capacity or browser acceptance is implied.

Save demo.py below beside queue_lab.py. It runs the same service formula against two workload shapes. The burst arrives at 0, 1, 2 and 3 milliseconds. The sparse trace arrives at 0, 20, 40 and 60. Every request has the same abstract work requirement. Run python demo.py. The following four lines are the exact observed output from the checked implementation; long lines can wrap on the printed page.

python · 16 lines
from queue_lab import Job, metrics, simulate

def show(label, arrivals, batch, delay):
    jobs = tuple(Job(f"r{i}", at) for i, at in enumerate(arrivals))
    done = simulate(jobs, batch, delay)
    m = metrics(done)
    print(f"{label}: finish={[d.finish for d in done]}; "
          f"mean_ms={float(m.mean_latency):.1f}; "
          f"p95_ms={m.p95_latency}; "
          f"rps={float(m.requests_per_second):.3f}")

if __name__ == "__main__":
    show("burst B1 W0", (0, 1, 2, 3), 1, 0)
    show("burst B4 W3", (0, 1, 2, 3), 4, 3)
    show("sparse B4 W0", (0, 20, 40, 60), 4, 0)
    show("sparse B4 W3", (0, 20, 40, 60), 4, 3)

text · 4 lines
burst B1 W0: finish=[10, 20, 30, 40]; mean_ms=23.5; p95_ms=37; rps=100.000
burst B4 W3: finish=[19, 19, 19, 19]; mean_ms=17.5; p95_ms=19; rps=210.526
sparse B4 W0: finish=[10, 30, 50, 70]; mean_ms=10.0; p95_ms=10; rps=57.143
sparse B4 W3: finish=[13, 33, 53, 73]; mean_ms=13.0; p95_ms=13; rps=54.795

With burst arrivals and capacity 1, requests finish at 10, 20, 30 and 40. Later requests queue behind earlier work. Capacity 4 with delay 3 collects the burst into one group that starts at 3 and finishes at 19. The first request now takes 19 rather than 10 milliseconds, but the last takes 16 rather than 37. The mean falls from 23.5 to 17.5 while the earliest reader has become slower.

The grouped burst uses one fixed setup cost instead of four. That explains the higher finite-cohort completion rate under our assumed formula: 4,000/19 is about 210.526 requests per second. The speedup was built into the service-cost assumption and the chosen arrivals. It does not show that any real engine doubles throughput with a batch of four. Real batch service time must be observed across relevant request sizes and resource constraints.

Sparse arrivals do not find companions within the 3-millisecond delay. Each request starts alone after waiting, then takes 10 milliseconds of service. All response latencies rise from 10 to 13. The last completion moves from 70 to 73, lowering the full-window rate from about 57.143 to 54.795. Waiting has delivered no sharing benefit for this trace.

Maximum batch size is a ceiling, not a command to wait until that many requests exist. With capacity 4 and delay 0, the first burst request starts alone at time 0. Requests arriving while it is in service can still form a later batch when the server becomes free. Zero deliberate batching delay therefore does not imply every future batch has one member; existing queueing can provide a group.

Do not choose a policy from one average. Inspect the first reader, late readers, the latency distribution and the mix of arrival patterns. Repeat a real experiment across representative low, normal and burst demand. Keep the model, hardware, prompt/output mix and quality checks fixed while changing one scheduler setting. Otherwise an apparent scheduling gain may come from shorter outputs, caching or changed request contents.

A finite list always drains in this model because arrivals stop and each accepted request has positive finite service. That says nothing about a service that receives new work indefinitely. If offered demand repeatedly exceeds completion capacity, unfinished work can accumulate until a separate admission policy acts. This workbook does not implement such a policy; bounding the input tuple is a limit on a teaching fixture, not production backpressure.

Try it yourself · Activity 04

20 min

Explain the trade-off to a product owner

Compare the burst and sparse outputs without making a hardware claim.

  1. Identify who becomes slower in the grouped burst.
  2. Explain why sparse requests gain no benefit.
  3. Propose the next controlled comparison before selecting a production policy.

Would the decision change if the first response mattered more than the final completion rate?

Worked answer

The first burst request becomes slower, from 10 to 19 simulated milliseconds, while later requests benefit from shared fixed work. Sparse requests remain alone and each pays an extra 3 milliseconds. A next experiment should replay representative arrival and length distributions on the actual engine, preserve output-quality requirements and compare individual latency, tail latency, errors and completed throughput. The fixture does not select a universal batch size.

Read a sample · Chapter 05 of 06

05

Check the scheduler from a different direction

A second implementation should advance time differently and agree on the observable trace.

Save test_queue.py below. Run python -m unittest -v test_queue.py from the same folder. Thirteen tests cover exact traces, batching deadlines, capacity, stable ties, idle gaps, rank boundaries, invalid inputs, frozen observations and occupancy accounting. The final test contains 972 small comparisons; that is a count of reference fixtures inside one test, not 972 independent unit-test methods.

python · 8 lines
from dataclasses import FrozenInstanceError
from fractions import Fraction
from itertools import product
import unittest
from queue_lab import Done, Job, metrics, nearest_rank, simulate

def jobs(arrivals):
    return tuple(Job(f"r{i}", at) for i, at in enumerate(arrivals))

python · 25 lines
def tick_reference(requests, capacity, delay, fixed=8, per_item=2):
    """Independent one-millisecond clock for small teaching fixtures."""
    ordered = sorted(requests, key=lambda job: job.arrival)
    waiting = []
    pending = 0
    busy_until = 0
    result = []
    batch = 0
    for tick in range(1000):
        while pending < len(ordered) and ordered[pending].arrival <= tick:
            waiting.append(ordered[pending])
            pending += 1
        if tick < busy_until or not waiting:
            continue
        if len(waiting) < capacity and tick - waiting[0].arrival < delay:
            continue
        group = waiting[:capacity]
        waiting = waiting[capacity:]
        busy_until = tick + fixed + per_item * len(group)
        batch += 1
        result.extend(Done(j.name, j.arrival, tick, busy_until, batch)
                      for j in group)
        if len(result) == len(ordered):
            return tuple(result)
    raise AssertionError("reference fixture exceeded 1000 ticks")

python · 23 lines
class QueueTests(unittest.TestCase):
    def test_serial_burst(self):
        done = simulate(jobs((0, 1, 2, 3)), 1, 0)
        self.assertEqual([d.finish for d in done], [10, 20, 30, 40])
        self.assertEqual(metrics(done).mean_latency, Fraction(47, 2))
        self.assertEqual(metrics(done).p95_latency, 37)

    def test_grouped_burst(self):
        done = simulate(jobs((0, 1, 2, 3)), 4, 3)
        self.assertEqual([d.start for d in done], [3] * 4)
        self.assertEqual([d.finish for d in done], [19] * 4)
        self.assertEqual(metrics(done).requests_per_second, Fraction(4000, 19))

    def test_sparse_delay_penalty(self):
        for delay in (0, 3):
            done = simulate(jobs((0, 20, 40, 60)), 4, delay)
            self.assertEqual([d.finish - d.arrival for d in done],
                             [10 + delay] * 4)

    def test_busy_server_does_not_restart_delay(self):
        done = simulate(jobs((0, 1, 2)), 2, 3)
        self.assertEqual([(d.start, d.finish) for d in done],
                         [(1, 13), (1, 13), (13, 23)])

python · 20 lines
    def test_deadline_boundary_and_future_arrival(self):
        done = simulate(jobs((0, 3, 4)), 4, 3)
        self.assertEqual([d.batch for d in done], [1, 1, 2])
        self.assertEqual([d.start for d in done], [3, 3, 15])

    def test_capacity_is_enforced(self):
        done = simulate(jobs((0, 0, 0, 0, 0)), 2, 0)
        self.assertEqual([d.batch for d in done], [1, 1, 2, 2, 3])

    def test_sort_and_stable_ties(self):
        requests = (Job("later", 9), Job("a", 0), Job("b", 0))
        done = simulate(requests, 1, 0)
        self.assertEqual([d.name for d in done], ["a", "b", "later"])
        self.assertEqual(requests[0].name, "later")

    def test_window_includes_idle_gap_and_drain(self):
        m = metrics(simulate(jobs((100, 200)), 1, 0))
        self.assertEqual(m.horizon, 110)
        self.assertEqual(m.requests_per_second, Fraction(200, 11))
        self.assertEqual(m.busy_fraction, Fraction(2, 11))

python · 24 lines
    def test_nearest_rank_boundaries(self):
        values = list(range(1, 22))
        self.assertEqual(nearest_rank(values, 95), 20)
        self.assertEqual(nearest_rank(values, 100), 21)
        self.assertEqual(nearest_rank([37, 10, 28, 19], 95), 37)
        for bad in (0, 101, True, 50.5):
            with self.assertRaises(ValueError):
                nearest_rank(values, bad)

    def test_invalid_input(self):
        for bad in ((), [], (Job("x", 0), Job("x", 1)),
                    (Job(" ", 0),), (Job("x", True),),
                    (Job("x", -1),), (Job("x", 0.5),), ("x",),
                    jobs(range(257))):
            with self.assertRaises(ValueError):
                simulate(bad, 1, 0)
        for args in ((0, 0), (65, 0), (True, 0), (2, -1), (2, 1.5),
                     (2, 0, -1), (2, 0, 0, 0)):
            with self.assertRaises(ValueError):
                simulate(jobs((0,)), *args)
        with self.assertRaises(ValueError):
            metrics(())
        with self.assertRaises(ValueError):
            nearest_rank([], 95)

python · 26 lines
    def test_immutable_observations(self):
        done = simulate(jobs((0,)), 1, 0)
        with self.assertRaises(FrozenInstanceError):
            done[0].start = 5

    def test_exact_occupancy_accounting(self):
        done = simulate(jobs((0, 1, 2, 3, 30)), 4, 3)
        m = metrics(done)
        area = sum(sum(d.arrival <= t < d.finish for d in done)
                   for t in range(m.horizon))
        self.assertEqual(m.average_in_system, Fraction(area, m.horizon))
        self.assertEqual(m.average_in_system,
                         m.requests_per_second * m.mean_latency / 1000)
        self.assertEqual(m.mean_wait, Fraction(9, 5))

    def test_independent_reference_972_cases(self):
        for arrivals in product((0, 1, 5), repeat=3):
            for capacity, delay, fixed, per_item in product(
                    (1, 2, 4), (0, 1, 3), (0, 8), (1, 2)):
                with self.subTest(arrivals=arrivals, capacity=capacity,
                                  delay=delay, fixed=fixed, per_item=per_item):
                    requests = jobs(arrivals)
                    actual = simulate(requests, capacity, delay, fixed, per_item)
                    expected = tick_reference(requests, capacity, delay,
                                              fixed, per_item)
                    self.assertEqual(actual, expected)

python · 2 lines
if __name__ == "__main__":
    unittest.main()

The reference advances a clock one millisecond at a time. It enqueues arrivals, waits while the server is busy and dispatches when the queue is full or the oldest waiting request reaches its allowance. It does not use the event simulator's minimum of deadline and future full-batch arrival. Both versions share the declared service formula and result record type. Agreement therefore checks scheduling logic under that assumption, not whether the service formula models hardware.

The generated fixtures combine three arrival choices for each of three requests, three capacities, three delays, two fixed costs and two per-item costs: 27 times 3 times 3 times 2 times 2 = 972. This includes out-of-order input, equal timestamps, early fill and expired deadlines. The reference's 1,000-tick guard is ample for these small cases; it is not intended for the simulator's largest supported arrival range.

Golden examples prevent two implementations from agreeing on the same misunderstanding without challenge. The busy-server test requires the third request to start at 13, not to receive a fresh waiting period. The capacity test requires five simultaneous requests with capacity 2 to use batch IDs 1, 1, 2, 2, 3. The percentile test uses 21 values so that rounding down rather than up produces a visible rank error.

During author review, three syntactically valid faulty copies were run separately. One restarted the batching deadline from current server availability, one allowed a batch to exceed capacity by one, and one rounded the percentile rank down. The corresponding tests failed, and the original files retained their hashes. These observations establish that the selected checks detect those three mistakes; they do not establish exhaustive defect coverage.

The occupancy check independently counts requests present at each tick for an arrival trace beginning at zero. A request is present from its arrival inclusively to its finish exclusively. That endpoint rule avoids counting it after completion. For arbitrary traces starting later, iterate from the first arrival to the final finish when writing a similar reference. The production metrics helper already uses the full first-arrival-to-last-finish window.

Testing a simulator and validating a model are different jobs. Unit tests and reference comparisons can show that code follows a chosen mathematical contract. Model validation would compare those predictions against observations from the target system and identify where the assumptions fail. Because this course never runs a model server, it supplies implementation evidence only. A visually convincing chart would not change that limit.

Try it yourself · Activity 05

20 min

Make a known defect observable

In a disposable copy, change end - cursor < max_batch to end - cursor <= max_batch.

  1. Predict the five simultaneous requests' batch IDs when capacity is 2.
  2. Run the tests and locate the capacity failure.
  3. Restore the checked files and explain why this is a logic defect rather than a parse error.

Which test would detect an extra waiting period even when the final batch sizes stayed correct?

Worked answer

The faulty loop can take three requests in the first batch, then two, giving IDs 1, 1, 1, 2, 2. The expected IDs are 1, 1, 2, 2, 3, so test_capacity_is_enforced fails. Both versions are valid Python; only the boundary condition differs. Restore the strict comparison and require all thirteen tests to pass before treating the code as the reference again.

Read a sample · Chapter 06 of 06

06

Turn the lesson into a measured capacity plan

A deployable conclusion needs evidence from the actual workload and serving boundary.

The useful handover from this lab is a measurement plan, not a GPU shopping list. Start with the user-visible requirement and the actual service boundary. Record whether timestamps come from the client, gateway or engine, and use a consistent clock for durations. If layers use different clocks, do not subtract their raw timestamps without an explicit alignment method. Keep a source revision and configuration snapshot with every result.

Scroll sideways to see every column.

Turn the lesson into a measured capacity plan · Table 3
Experiment fieldWhat to record before comparing settings
WorkloadArrival trace, prompt and output length distribution, endpoint and quality requirements.
SystemModel/tokeniser revision, precision, hardware, engine version, scheduler limits and cache policy.
TimingSubmission, first useful output, completion, cancellation and timeout boundaries.
OutcomesSubmitted, completed, rejected, failed, cancelled and unfinished counts.
ResourcesMemory pressure, queued and active requests, utilisation and constrained resources.
DecisionOwner-agreed latency/error objectives, observed trade-offs, headroom and rollback condition.

A fixed-arrival workload submits according to an external schedule even when responses slow down. A fixed-concurrency workload sends the next request only when a slot becomes free. The latter can reduce offered traffic as latency rises, which may hide the queue growth a real burst would cause. Both can answer useful questions, but they answer different ones. Record the load-generation rule and whether the generator itself can sustain its schedule.

Separate warm-up from the declared measurement window and record how caches are handled. Run repeated trials and show variation instead of presenting one favourable run as a stable property. Keep failed and unfinished work visible at the cutoff. If you summarise a completed cohort, drain it and include that drain in the declared window as this lab does. If you use fixed wall-clock windows, account for requests crossing the boundaries explicitly.

Google's SRE monitoring chapter groups core signals into latency, traffic, errors and saturation. It also distinguishes successful-request latency from failed-request latency. Use those perspectives together: a fast rejection is not a fast successful explanation, and growing queues may matter before a resource meter reaches its maximum. A single average or device utilisation percentage does not establish a healthy serving configuration.

For streaming inference, include both the delay before the first useful token and the pacing of later output alongside final completion. Retain raw observations needed for the chosen aggregation and percentile method. Do not combine requests-per-second and output-tokens-per-second as though their denominators and work units were identical. The output-length distribution is part of interpreting both.

For an initial demand estimate, list expected arrival patterns and work sizes, then test candidate settings within the actual resource limits. Do not infer a safe concurrency limit by dividing GPU memory only by model weights; active-request state also matters. Do not infer a safe request rate from the invented 8 + 2n formula. The earlier inference and memory workbooks provide background, while final headroom and admission limits require measured evidence.

A release note should separate executed observations from planned experiments. This release executed a standard-library simulator, thirteen tests, 972 reference fixtures, four demo outputs and three deliberate logic defects. It did not execute a GPU workload, streaming client, continuous batcher, failure-rate study or production load test. Naming these boundaries lets the next builder extend the work without turning a teaching example into an unsupported operational promise.

Try it yourself · Activity 06

15 min

Write the next experiment handover

Prepare a short note that another builder can run and judge.

  1. Specify the real system and workload snapshot to use.
  2. Choose one batching setting to vary and keep other conditions fixed.
  3. Name the result fields, acceptance owner and remaining untested behaviours.

Can the recipient distinguish a completed local code check from a proposed production experiment?

Worked answer

The note should identify an exact engine/model/hardware configuration, a replayable arrival and request-length trace, fixed quality checks and a small chosen set of batching delays. Capture per-request timing and outcome, aggregate completed throughput, agreed latency percentiles, resource observations and unfinished work. The service owner supplies acceptance thresholds. State that actual serving, streaming, overload, cancellation and recovery remain untested by this simulator.

Keep learning

The complete workbook

Six chapters and six worked activities include a complete three-file Python lab, exact output, an independent clock-based reference and a reviewed workbook. Learn queue delay, response latency, nearest-rank percentiles, measurement windows and occupancy accounting, then write an experiment plan for real streaming inference.

  1. 01
    Choose the question before measuring speed

    A capacity claim needs a workload, a time boundary and an acceptable result.

    Read here · 1 exercise
  2. 02
    Model the queue and the batch deadline

    Make time advancement explicit so that waiting cannot disappear from the result.

    Read here · 1 exercise
  3. 03
    Calculate metrics over one declared window

    Exact arithmetic makes accounting checkable; it cannot make a small sample representative.

    Read here · 1 exercise
  4. 04
    Compare burst arrivals with sparse demand

    One batch policy can help one traffic pattern and hurt another.

    Read here · 1 exercise
  5. 05
    Check the scheduler from a different direction

    A second implementation should advance time differently and agree on the observable trace.

    Read here · 1 exercise
  6. 06
    Turn the lesson into a measured capacity plan

    A deployable conclusion needs evidence from the actual workload and serving boundary.

    Read here · 1 exercise

Also inside: a 8-point checklist, a glossary of 10 terms and 10 questions and answers to test yourself. 6 hands-on exercises, each with a worked answer at the back where the workbook gives one.

No login, no card, no account. Before the download we ask you to follow Mickai (two quick links). Free to download and use for personal learning, study groups and inside your own team. Please do not resell the workbooks or republish them as your own. Link people to trust-agent.ai instead.

Test yourself

Questions and answers

Does this course benchmark a GPU?

No. Its integer-time simulator assumes a service duration of 8 + 2n milliseconds. Those teaching values were not measured on hardware.

Is throughput the same as response latency?

No. Throughput counts completed work per time window; latency is the duration experienced by one request. A batch policy can affect them differently.

Does a 3-millisecond batching delay cap queueing at 3 milliseconds?

No. A request may also wait behind a busy server. The busy-server fixture has an 11-millisecond queue delay despite a 3-millisecond batching allowance.

Does the simulator implement continuous batching?

No. Every member of a batch finishes together. It has no token iterations, prefill/decode scheduling or changing active sequence set.

Which p95 method is used?

Nearest rank: sort the observations and take rank ceiling(N times 95 divided by 100). With four samples it selects the maximum.

What time window defines the reported throughput?

First arrival to final completion, including deliberate wait, idle gaps and final drain. It is a finite-cohort rate, not a sustained-capacity claim.

Why does waiting hurt the sparse fixture?

No companion arrives during the allowance, so every request waits an extra 3 milliseconds and still pays for a one-request batch.

Does zero batching delay always produce one-request batches?

No. Requests accumulated while the server is busy can form a group when it becomes free, subject to the capacity limit.

What do the 972 comparisons establish?

The event scheduler agrees with an independently clocked reference on those bounded fixtures under the same service formula. They do not validate hardware performance.

What is required before choosing a real serving limit?

A measured experiment on the actual configuration and workload, with defined timing boundaries, quality and error accounting, resource observations and owner-agreed acceptance criteria.

When you have finished

Get your certificate of completion

Type your name and download a certificate for this workbook as a PDF, ready to print or to add to LinkedIn. It is made on your own device, so your name is never sent to us. It is a self-declared certificate, not an accredited qualification.

Learn the language

Key terms

Throughput
Completed work divided by a declared observation window, with both work unit and time unit stated.
Response latency
Time from a request arriving at the chosen boundary to its complete result.
Queue delay
Time between arrival and the start of service, including waiting behind other work.
Batch
A group of requests served together under a stated scheduling and service contract.
Batching delay
The allowance for collecting a group; it does not cap waiting behind a busy server.
Time to first token
Time from the declared submission boundary to the first output token event used by the measuring tool.

6 of the workbook's 10 terms. The complete glossary is in the workbook.

Follow the evidence

Sources and checks

Facts last checked: .

Examples in this workbook were run on: Python 3.12.10 on Windows; standard library only. Thirteen learner tests include 972 comparisons against an independent clock-based reference. Four exact demo lines and three detected logic faults. All times are simulated integer milliseconds with an assumed service formula; no model-server, GPU, streaming or continuous-batching benchmark was performed. (2026-09-27).

These workbooks use AI assistance. See how the workbooks are made.

  1. Batchers and batching strategiesNVIDIA Triton documentation
  2. LLM benchmarking metricsNVIDIA NIM documentation
  3. Streaming generated textHugging Face documentation
  4. Monitoring distributed systemsGoogle SRE
  5. Rational numbers with FractionPython Software Foundation
  6. Unit testing and assertionsPython Software Foundation

Created by Mickarle Wagstaff-Irons - Micky Irons with the Mickai team. Published by Mickai LTD. Last updated 27 September 2026.

NextKeep going

Where to go next