models · Level 4

Build a tiny language model from scratch

Train, evaluate, save and sample a character model whose entire learning rule you can inspect.

By Mickarle Wagstaff-Irons - Micky Irons

  • Level 4Frontier
  • 180 min
  • 6 chapters
  • Free PDF, no account
The Structuremodels / 04

Start with the essentials

The short answer

This complete CPU-only lab learns next-character probabilities from twelve fictional workshop notes. You implement the loss and gradient, train 812 parameters, evaluate held-out documents and generate text. The model remembers only the preceding character. It is a genuine small language model, but it is not a transformer, an LLM or a useful general assistant.

What you will learn

  • Build next-character targets without joining documents.
  • Calculate mean cross-entropy and its logit gradient.
  • Run a complete fixed training and evaluation cycle.
  • Interpret held-out perplexity within its token and data definition.
  • Generate bounded samples and reload a checked JSON checkpoint.
  • State what a one-character model cannot learn about language.

Who it is for

Python builders who understand a training loop and want a small working model before taking on a trainable transformer.

Before you start

  • Read Python lists, functions and unit tests.
  • Understand gradients, exponentials and a parameter update.
  • Recognise the difference between an attention block and a complete language model.

Read a sample · Chapter 01 of 06

01

Define a model small enough to understand

Make the data, context and missing capabilities explicit.

Create one folder containing data.py, model.py, train.py and test_model.py. Copy the complete files shown in this workbook, joining the consecutive model.py parts in order. Python 3.12 is sufficient; the tested interpreter was 3.12.10. Run python -m unittest -v test_model and then python train.py from that folder. No package installation or downloaded model is involved. The examples use original fictional text rather than private messages or scraped documents.

This is a character bigram model: after a character, one row of learned scores predicts the next character or the end of the document. It does not inspect the earlier sentence. The input-only beginning context predicts the first character. With V output tokens, the score table has V+1 rows and V columns. Our vocabulary has 28 output tokens, giving 29*28 = 812 stored parameters. Some rows never receive training examples, but they remain part of the stored table.

The model is deliberately simpler than the attention block in the prerequisite. It uses no attention, embeddings, positions, hidden layers or backpropagation library. Instead, a directly learned score table lets us inspect every derivative and complete the surrounding training, evaluation and generation workflow. An attention block with fixed weights teaches a forward calculation; this course adds a fully trained language-model baseline through a different, smaller architecture.

Save the following original corpus as data.py. The split is fixed before constructing the vocabulary or transition counts. Twelve documents train the model, three monitor validation loss and three form the final test set. Exact document overlap is checked. These closely related workshop notes are still a tiny, narrow distribution, not independent evidence about general language. The split does not establish topic diversity or eliminate shared phrases.

python · 25 lines
"""Original fictional lantern-workshop notes, split by document."""
TRAIN = (
    "amber lamps glow beside the quiet canal.",
    "a small boat carries paper lanterns home.",
    "mira folds blue paper in the warm room.",
    "the brass bell rings before the workshop opens.",
    "soft rain taps the roof of the lantern shed.",
    "a green ribbon rests beside a wooden frame.",
    "the keeper trims a wick and closes the door.",
    "we carry the light across the narrow bridge.",
    "each evening the maker checks the paper seams.",
    "the little boat returns beneath a silver moon.",
    "warm tea waits on a table by the window.",
    "a quiet breeze moves the lantern above the steps.",
)
VALID = (
    "the maker carries a blue lamp to the bridge.",
    "a warm light glows beside the workshop door.",
    "mira checks the ribbon before the boat returns.",
)
TEST = (
    "paper lamps move above the quiet water.",
    "the keeper carries tea across the narrow room.",
    "a silver bell rings beside the wooden door.",
)

The corpus is already lowercase and uses spaces and full stops as characters. We do not strip, lowercase or otherwise normalise it at runtime. Adding normalisation later changes the experiment and must be applied consistently to all splits. Unicode characters are Python string elements here; this is not a byte tokenizer or a general text-normalisation recipe. Keep the original split when reproducing the published numbers.

Try it yourself · Activity 01

15 min

Specify the experiment

Write a short model card before running the code.

  1. State the context length and output units.
  2. Calculate the parameter count for V=28.
  3. List the document counts and two claims this corpus cannot support.

Would the phrase language model alone cause a reader to overestimate what your experiment can do?

Worked answer

The context is one previous character, with an extra beginning context. Outputs are character tokens, <eos> and <unk>. There are (28+1)*28 = 812 stored scores. The split contains 12 training, 3 validation and 3 test documents. It cannot establish general language competence or frontier-model training costs; all notes come from one small fictional workshop setting.

Read a sample · Chapter 02 of 06

02

Turn separate documents into targets

Count boundaries before counting accuracy.

Begin model.py with the imports and vocabulary function below. The vocabulary is learned from TRAIN only: ID 0 means end of document, ID 1 means an unknown character, and the remaining characters are sorted for stable IDs. The beginning context has ID V and exists only as an input row. It can never be sampled as an output token. The limit of 126 ordinary characters keeps this teaching implementation small; it is not a universal tokenizer limit.

python · 13 lines
"""Bounded character bigram model. Standard library only."""
import json
import math
import random


def vocabulary(docs):
    if not docs or any(not isinstance(s, str) or not s for s in docs):
        raise ValueError("nonempty text documents required")
    chars = sorted(set("".join(docs)))
    if len(chars) > 126:
        raise ValueError("at most 126 training characters")
    return ["<eos>", "<unk>"] + chars

Append transitions and initialise. Every document begins with a fresh beginning context. For a document ab, the pairs are beginning-to-a, a-to-b and b-to-end. An empty held-out document contributes beginning-to-end. The end token never connects to the next document. The count table records how many times each context-target pair occurs; expanding those counts back into repeated examples would give the same full-batch objective.

python · 16 lines
def transitions(docs, vocab):
    size = len(vocab)
    ids = {char: i for i, char in enumerate(vocab)}
    counts = [[0] * size for _ in range(size + 1)]
    total = unknown = 0
    for doc in docs:
        previous = size  # Input-only beginning-of-document context.
        targets = [ids.get(char, 1) for char in doc] + [0]
        unknown += targets.count(1)
        for target in targets:
            counts[previous][target] += 1
            total += 1
            previous = target
    if not total:
        raise ValueError("at least one target required")
    return counts, total, unknown

python · 2 lines
def initialise(size):
    return [[0.0] * size for _ in range(size + 1)]

For training documents [ab, a], the vocabulary is [<eos>, <unk>, a, b]. Beginning has ID 4. There are five targets: beginning-to-a twice, a-to-b once, a-to-end once and b-to-end once. The beginning row is [0, 0, 2, 0], the a row is [1, 0, 0, 1] and the b row is [1, 0, 0, 0]. The end and unknown context rows are all zero. There is no b-to-a pair joining documents.

An unseen held-out character maps to <unk>; it does not enlarge the vocabulary. The following prediction then uses the unknown context row, which has no observations in this training corpus. That row stays at its initial uniform distribution. The supplied validation and test splits have zero unknown targets, but a separate unit test deliberately introduces z into the small ab vocabulary. Report unknown counts alongside loss when using different data.

The low-level arithmetic functions assume tables made by these helpers. They are not exposed as a service accepting arbitrary matrices, and zip would silently truncate inconsistent manual shapes. Use the validated restore function later for checkpoint input. The data size is intentionally small; transitions allocates one target list per document and does not bound document length. Do not treat this script as an ingestion pipeline for unlimited or hostile input.

Try it yourself · Activity 02

20 min

Audit the transition table

Use the two training documents ab and a.

  1. List all five target events.
  2. Identify the beginning, end and unknown IDs.
  3. Encode a held-out z and an empty document without changing the vocabulary.

Which loss value would become misleading if you accidentally built the vocabulary from all three splits?

Worked answer

The five events are beginning-to-a, a-to-b, b-to-end, beginning-to-a and a-to-end. Beginning is input ID 4, end is output ID 0 and unknown is ID 1. The held-out documents [z, empty] give beginning-to-unknown, unknown-to-end and beginning-to-end, for three targets and one unknown. No held-out character is added to the vocabulary.

Read a sample · Chapter 03 of 06

03

Derive and update the learning rule

Follow one gradient before trusting a training curve.

Watch one learning step

Two documents, ab and a, give five targets. One gradient-descent step at rate 1 changes the beginning-context probability of a from 0.250000 to 0.332120.
Follow the course’s two-document hand calculation. The graphic highlights the beginning-context row; the tables retain every observed transition and all four outputs. Open the full-size training-step diagram.

The documents are ab and a. Their separate paths are BOS → a → b → <eos> and BOS → a → <eos>. Reset to BOS for each document. End is a target; beginning is an input-only context with ID 4. The output IDs are 0 for <eos>, 1 for <unk>, 2 for a and 3 for b.

All observed transitions: five targets across two separate documents
ContextTarget Count
BOSa2
ab1
a<eos>1
b<eos>1

There is no transition from one document into the next. Initialise every logit to zero. The beginning context appears twice, both times followed by a. For this row, N=2; across all contexts, T=5. Its derivative for each output is (N*p - count) / T, using the whole training loss’s denominator, not just the number of beginning contexts.

Beginning row before the update: every logit is zero
OutputCount ProbabilityGradient
<eos>00.2500000.100000
<unk>00.2500000.100000
a20.250000-0.300000
b00.2500000.100000

With learning rate 1, subtract the gradient: new logit = old logit - gradient. The negative derivative raises the score for a; the other scores fall. Apply softmax to the updated row. These are four probabilities for one context, not probabilities that the generated text is true.

Beginning row after one update at learning rate 1
OutputNew logit New probability
<eos>-0.1000000.222627
<unk>-0.1000000.222627
a0.3000000.332120
b-0.1000000.222627

p(a | BOS) = exp(0.3) / (exp(0.3) + 3*exp(-0.1)) = 0.332120, rounded. The other three outputs each have probability 0.222627, rounded. Use unrounded numbers for calculations; displayed probabilities need not sum to exactly one. This figure isolates the beginning row, while a full update changes every observed context row. Unobserved end and unknown context rows have zero gradients.

This two-document, one-update fixture is separate from the twelve-document workshop training run, which uses 400 updates at rate 4.0. Both use the same published model functions. The figure does not depict a transformer, an attention layer or a language-quality benchmark.

Append probabilities, objective and update to model.py. A row contains logits, which are unrestricted scores before normalisation. Softmax converts the row into positive probabilities by exponentiating and dividing by their sum. Subtracting the largest logit first keeps the largest exponential at one. The same shift cancels in the normalisation. The stable log normaliser in objective computes loss without taking the logarithm of a tiny rounded probability.

python · 5 lines
def probabilities(row):
    largest = max(row)
    masses = [math.exp(value - largest) for value in row]
    total = math.fsum(masses)
    return [mass / total for mass in masses]

python · 14 lines
def objective(weights, counts, total):
    loss = 0.0
    gradient = []
    for row, observed in zip(weights, counts):
        number = sum(observed)
        largest = max(row)
        log_z = largest + math.log(math.fsum(
            math.exp(value - largest) for value in row))
        loss += math.fsum(n * (log_z - value)
                         for n, value in zip(observed, row))
        p = probabilities(row)
        gradient.append([(number * prob - n) / total
                         for prob, n in zip(p, observed)])
    return loss / total, gradient

python · 4 lines
def update(weights, gradient, rate):
    for row, changes in zip(weights, gradient):
        for j, change in enumerate(changes):
            row[j] -= rate * change

For context i and target j, let C[i,j] be its observed count, N[i] the total for context i, p[i,j] its current probability and T the total number of targets across documents. The mean negative log likelihood is the sum of C[i,j]*(logZ[i]-W[i,j]) over all pairs, divided by T. We use natural logarithms, so the unit is nats per target. Each document-end target counts once. Beginning contexts do not add a separate target of their own.

The derivative with respect to W[i,j] is (N[i]*p[i,j]-C[i,j])/T. The first part is the count predicted by the model; subtracting the observed count identifies the discrepancy. A parameter update subtracts learning_rate times this derivative. We use a fixed rate of 4.0 for the original full-batch corpus and 400 updates. That choice is part of this experiment, not a general recommendation for another model or loss normalisation.

At zero logits the four-token ab fixture has probability 1/4 everywhere and mean loss log(4), about 1.386294. Its beginning context appears twice, both times targeting a, so its gradient is [0.1, 0.1, -0.3, 0.1]. With learning rate 1, one update changes that row to [-0.1, -0.1, 0.3, -0.1]. The probability of a becomes exp(0.3)/(exp(0.3)+3*exp(-0.1)), about 0.332120. The correct next character gains probability.

An unobserved context has N[i]=0 and all counts zero, so its entire gradient is zero. Its initial row remains uniform. Observed rows can also contain target pairs that never occur; finite training reduces their probability without making it exactly zero in this small run. We do not add a count-smoothing rule, regularisation, momentum, clipping or early stopping. Initialising every logit to zero is adequate for this direct table; it is not an instruction to initialise every layer of a multilayer network identically.

The release checks compared the analytic gradient with central finite differences of a separately written per-target loss. Across 24 small corpora, all 624 parameter checks agreed within 1e-9; the largest absolute error was about 3.70e-11. That calculation uses unshifted exponentials only on logits bounded between -2 and 2. It provides a useful independent arithmetic path without pretending to cover every possible floating-point input.

Try it yourself · Activity 03

20 min

Calculate one descent step

Use the beginning row of the ab and a fixture.

  1. Compute N, T and the four probabilities at initialisation.
  2. Derive the gradient and updated row with rate 1.
  3. Explain why adding the gradient would be the wrong direction.

Have you checked a derivative independently, or only checked whether training appears to improve?

Worked answer

N=2, T=5 and every probability is 0.25. The gradient is [0.1, 0.1, -0.3, 0.1]. Subtracting it gives [-0.1, -0.1, 0.3, -0.1], raising the a probability to about 0.332120. Adding it would lower the score of the repeatedly observed target. The one-update test checks an actual reduction in the complete fixture loss.

Read a sample · Chapter 04 of 06

04

Sample and preserve the trained state

Keep randomness, stopping and saved-file validation visible.

Append sample to model.py. Each call makes its own random.Random instance, leaving global random state alone. Draw one uniform number, accumulate the next-token probabilities and select the first cumulative interval containing the draw. Start from the beginning context and feed each selected character back as the next context. This repeatedly applies the learned one-character rule; it does not extend the model context.

python · 19 lines
def sample(weights, vocab, seed=7, limit=100):
    if type(limit) is not int or not 1 <= limit <= 1000:
        raise ValueError("limit must be an integer from 1 to 1000")
    rng = random.Random(seed)
    previous = len(vocab)
    output = []
    for _ in range(limit):
        draw, cumulative = rng.random(), 0.0
        target = len(vocab) - 1
        for j, p in enumerate(probabilities(weights[previous])):
            cumulative += p
            if draw < cumulative:
                target = j
                break
        if target == 0:
            return "".join(output), "eos"
        output.append("?" if target == 1 else vocab[target])
        previous = target
    return "".join(output), "limit"

The sampler stops on <eos> or after the supplied character limit, and returns the reason separately. Unknown output is displayed as a question mark while its internal ID remains 1. A cap is necessary even when end has positive probability: a probabilistic stopping rule does not provide a fixed maximum length. The default is 100 characters and the helper accepts integer limits from 1 to 1000. There is no temperature adjustment, prompt prefix or top-k filter in this version.

Append restore and checkpoint to finish model.py. JSON stores the vocabulary order and weight table together with version 1. Reloading the same numbers under a reordered vocabulary would change their meaning. Validation rejects an unexpected schema, duplicate or malformed tokens, invalid dimensions and non-finite or excessive weights. It limits the checkpoint string to two million characters. The helpers return ordinary lists, with no executable model object or custom deserialisation hook.

python · 26 lines
def restore(text):
    if not isinstance(text, str) or len(text) > 2_000_000:
        raise ValueError("checkpoint exceeds text limit")
    obj = json.loads(text)
    if not isinstance(obj, dict) or set(obj) != {"version", "vocab", "weights"}:
        raise ValueError("unexpected checkpoint fields")
    vocab, weights = obj["vocab"], obj["weights"]
    if type(obj["version"]) is not int or obj["version"] != 1:
        raise ValueError("unsupported version")
    if not isinstance(vocab, list) or not 3 <= len(vocab) <= 128:
        raise ValueError("invalid vocabulary size")
    if vocab[:2] != ["<eos>", "<unk>"]:
        raise ValueError("invalid special tokens")
    if any(not isinstance(c, str) or len(c) != 1 for c in vocab[2:]):
        raise ValueError("invalid character token")
    if len(set(vocab)) != len(vocab):
        raise ValueError("duplicate token")
    if not isinstance(weights, list) or len(weights) != len(vocab) + 1:
        raise ValueError("invalid context count")
    for row in weights:
        if not isinstance(row, list) or len(row) != len(vocab):
            raise ValueError("invalid row size")
        if any(type(x) not in (int, float) or not -1e4 <= x <= 1e4
               or not math.isfinite(x) for x in row):
            raise ValueError("invalid weight")
    return vocab, weights

python · 5 lines
def checkpoint(vocab, weights):
    text = json.dumps({"version": 1, "vocab": vocab, "weights": weights},
                      allow_nan=False, separators=(",", ":"))
    restore(text)  # Validate before offering a reusable file.
    return text

The two-million-character check happens after a caller has supplied a Python string. It does not bound a file read or network request before that string is created, nor does it cover every adversarial JSON structure. Only load the small checkpoint produced by this lab unless you add an appropriately bounded input layer. Versioning and shape checks reduce accidental errors; they do not turn a JSON parser into a complete hostile-input security boundary.

The checkpoint intentionally contains inference state only. Training configuration and corpus provenance are recorded in the script and course rather than silently bundled into the weight file. To reproduce training, retain all four source files, the interpreter version, the fixed split, rate 4.0, 400 updates and the result log. For a resumable optimiser with momentum or adaptive statistics, additional optimiser state would be needed. This simple batch update has no such state.

Try it yourself · Activity 04

15 min

Design a reproducible sample

Describe what must remain fixed when comparing generated text.

  1. List the inference state and sampling settings.
  2. Explain the two stopping reasons.
  3. Identify one checkpoint check and one input boundary it does not protect.

Could someone reload the numbers correctly but still sample with a different vocabulary mapping?

Worked answer

Keep vocabulary order, weights, sampler code, seed, character cap and runtime fixed. The loop stops on end-of-document or the explicit cap; the latter is not evidence that a sentence finished. A row-length check rejects a mismatched weight table. The string-size check does not bound an earlier file read. Save the source and training configuration separately because this checkpoint stores only inference state.

Read a sample · Chapter 05 of 06

05

Run the experiment and read the result

Use held-out numbers and unedited samples together.

Save train.py below. The loop evaluates step zero, performs 400 gradient updates and reports progress after 100, 200 and 400 updates. Validation is monitored without selecting a new setting. The test set is evaluated after the fixed schedule; it is never used to build the vocabulary or compute an update. Once a test result has been inspected, repeated design changes based on it would turn it into development feedback and require a fresh final test set.

python · 45 lines
"""Run: python train.py [--checkpoint new-file.json]."""
import argparse
import math
from pathlib import Path
from data import TRAIN, VALID, TEST
from model import (vocabulary, transitions, initialise, objective,
                   update, sample, checkpoint, restore)


def main():
    parser = argparse.ArgumentParser(description=__doc__)
    parser.add_argument("--checkpoint", type=Path)
    args = parser.parse_args()
    vocab = vocabulary(TRAIN)
    train = transitions(TRAIN, vocab)
    valid = transitions(VALID, vocab)
    weights = initialise(len(vocab))
    print(f"vocabulary={len(vocab)} parameters={len(vocab)*(len(vocab)+1)}")
    print(f"train_targets={train[1]} valid_targets={valid[1]}")
    for step in range(401):
        loss, gradient = objective(weights, *train[:2])
        if step in (0, 100, 200, 400):
            val_loss, _ = objective(weights, *valid[:2])
            print(f"step={step} train_nll={loss:.6f} valid_nll={val_loss:.6f}")
        if step < 400:
            update(weights, gradient, 4.0)
    test = transitions(TEST, vocab)
    test_loss, _ = objective(weights, *test[:2])
    print(f"test_targets={test[1]} test_nll={test_loss:.6f}")
    print(f"test_perplexity={math.exp(test_loss):.6f}")
    print(f"unknown_targets train={train[2]} valid={valid[2]} test={test[2]}")
    saved = checkpoint(vocab, weights)
    restored_vocab, restored_weights = restore(saved)
    assert restored_vocab == vocab and restored_weights == weights
    for seed in (7, 11, 23):
        text, stopped = sample(weights, vocab, seed=seed)
        print(f"seed={seed} stopped={stopped} text={text!r}")
    if args.checkpoint:
        with args.checkpoint.open("x", encoding="utf-8") as file:
            file.write(saved + "\n")
        print("checkpoint written; existing files are never replaced")


if __name__ == "__main__":
    main()

Run python train.py after creating all files. The exact tested output follows, including all three predetermined sample seeds without selecting the nicest-looking string. Training has 535 targets, validation 138 and test 131, each including document ends. Initial mean loss is log(28), about 3.332205. The final training loss is 2.031755 and validation loss is 2.076519. The test result is 1.967426 nats per target, or perplexity 7.152244.

text · 12 lines
vocabulary=28 parameters=812
train_targets=535 valid_targets=138
step=0 train_nll=3.332205 valid_nll=3.332205
step=100 train_nll=2.278816 valid_nll=2.282750
step=200 train_nll=2.128103 valid_nll=2.157809
step=400 train_nll=2.031755 valid_nll=2.076519
test_targets=131 test_nll=1.967426
test_perplexity=7.152244
unknown_targets train=0 valid=0 test=0
seed=7 stopped=eos text='a r pe oal bhe bowheves be tanghe a s foomphelftheve.'
seed=11 stopped=eos text='lithen ops e th wuin am be mes.'
seed=23 stopped=eos text='wss ree blam apeere thenterella mooopa fr.'

Perplexity here is exp(mean negative log likelihood). It summarises next-token prediction under this vocabulary, split and end-token convention. Compare only evaluations with matching definitions; a word-level, byte-level or differently normalised score measures a different task. Test loss being lower than training loss is possible because these small document sets have different character-transition mixtures. It does not prove the model generalises better than it fits its training set.

The output includes fragments resembling the source text, but it also contains broken words. A one-character model cannot distinguish the next-step distribution after two histories that end with the same character. It cannot track an earlier subject, reason about a workshop or answer a question from these notes. Low loss on this narrow character distribution is not evidence of factual accuracy, useful instruction following or safety in a real application.

To save a checkpoint, run python train.py --checkpoint first-model.json. The file is created exclusively and an existing file with that name raises FileExistsError. Choose a new name for a later run. The release exercised a real disk write, reload and repeated-name failure, confirming that the first file stayed unchanged. A fresh sample after reloading can use the same sample function and seed; exact reproducibility is scoped to the tested code and runtime rather than all future Python versions.

One local CPU run took about 0.35 seconds including interpreter startup on the review machine. This is a single timing observation, not a hardware comparison, memory measurement or estimate of transformer-training cost. No GPU or network is used by these files. The count table and weight table each have 812 entries, but Python objects add overhead; multiplying 812 by a nominal numeric byte width would not accurately describe process memory.

Try it yourself · Activity 05

20 min

Write an honest result note

Report the fixed experiment without exaggerating its capability.

  1. Include train, validation and test target counts and final losses.
  2. Explain why the perplexity cannot be compared directly with a differently tokenised model.
  3. Include an unedited sample and one concrete context limitation.

Which part of your report is directly measured, and which part would need a separate evaluation?

Worked answer

The fixed run uses 535, 138 and 131 targets, with final losses 2.031755, 2.076519 and 1.967426 respectively. Test perplexity is 7.152244 under this exact character vocabulary and end-token rule. Different token units change the denominator and prediction task. Include one of the printed strings unchanged and explain that the model uses only the previous character; it cannot preserve a sentence-level subject or provide factual answers.

Read a sample · Chapter 06 of 06

06

Test failures before extending the model

Keep the baseline executable while deciding what comes next.

Save test_model.py below and run python -m unittest -v test_model. The ten methods check document boundaries, unknown targets, uniform loss, a hand-derived gradient, a descent step, softmax stability, split disjointness, checkpoint round trips and rejection, and bounded deterministic sampling. The expected runner result is Ran 10 tests followed by OK. These tests are available as complete source so a learner can run them without a repository-specific harness.

python · 96 lines
import json
import math
import unittest
from data import TRAIN, VALID, TEST
from model import (vocabulary, transitions, initialise, probabilities,
                   objective, update, sample, checkpoint, restore)


class ModelTests(unittest.TestCase):
    def setUp(self):
        self.vocab = vocabulary(["ab", "a"])
        self.counts, self.total, _ = transitions(["ab", "a"], self.vocab)
        self.weights = initialise(len(self.vocab))

    def test_documents_do_not_join(self):
        eos, a, b, bos = 0, 2, 3, 4
        self.assertEqual(self.total, 5)
        self.assertEqual(self.counts[bos][a], 2)
        self.assertEqual(self.counts[a][b], 1)
        self.assertEqual(self.counts[a][eos], 1)
        self.assertEqual(self.counts[b][eos], 1)
        self.assertEqual(sum(self.counts[eos]), 0)
        self.assertEqual(self.counts[b][a], 0)

    def test_unknown_and_empty_document(self):
        counts, total, unknown = transitions(["z", ""], self.vocab)
        self.assertNotIn("z", self.vocab)
        self.assertEqual((total, unknown), (3, 1))
        self.assertEqual(counts[4][1], 1)
        self.assertEqual(counts[4][0], 1)

    def test_uniform_loss(self):
        loss, _ = objective(self.weights, self.counts, self.total)
        self.assertAlmostEqual(loss, math.log(4))
        self.assertAlmostEqual(math.exp(loss), 4)

    def test_gradient_known_answer(self):
        _, grad = objective(self.weights, self.counts, self.total)
        self.assertEqual(grad[4], [0.1, 0.1, -0.3, 0.1])
        self.assertEqual(grad[0], [0.0] * 4)
        for row in grad:
            self.assertAlmostEqual(sum(row), 0)

    def test_one_update_reduces_loss(self):
        before, grad = objective(self.weights, self.counts, self.total)
        update(self.weights, grad, 1.0)
        after, _ = objective(self.weights, self.counts, self.total)
        self.assertLess(after, before)

    def test_stability_and_offset(self):
        p = probabilities([1000.0, 999.0])
        self.assertAlmostEqual(p[0], math.e / (math.e + 1))
        self.assertEqual(p, probabilities([0.0, -1.0]))

    def test_document_split(self):
        self.assertFalse(set(TRAIN) & set(VALID))
        self.assertFalse(set(TRAIN) & set(TEST))
        self.assertFalse(set(VALID) & set(TEST))
        self.assertEqual((len(TRAIN), len(VALID), len(TEST)), (12, 3, 3))

    def test_checkpoint_round_trip(self):
        _, grad = objective(self.weights, self.counts, self.total)
        update(self.weights, grad, 1.0)
        vocab, weights = restore(checkpoint(self.vocab, self.weights))
        self.assertEqual((vocab, weights), (self.vocab, self.weights))
        self.assertEqual(probabilities(weights[4]), probabilities(self.weights[4]))

    def test_checkpoint_rejects_malformed_data(self):
        base = json.loads(checkpoint(self.vocab, self.weights))
        variants = [dict(base, version=True), dict(base, weights=[]),
                    dict(base, vocab=["<eos>", "<unk>", "a", "a"])]
        bad = json.loads(json.dumps(base))
        bad["weights"][0][0] = float("nan")
        variants.append(bad)
        for obj in variants:
            with self.assertRaises(ValueError):
                restore(json.dumps(obj))
        with self.assertRaises(ValueError):
            restore(" " * 2_000_001)

    def test_sampling_repeats_and_stops(self):
        first = sample(self.weights, self.vocab, seed=11, limit=10)
        self.assertEqual(first, sample(self.weights, self.vocab, seed=11, limit=10))
        self.assertLessEqual(len(first[0]), 10)
        for row in self.weights:
            row[:] = [1000.0, -1000.0, -1000.0, -1000.0]
        self.assertEqual(sample(self.weights, self.vocab), ("", "eos"))
        for row in self.weights:
            row[:] = [-1000.0, -1000.0, 1000.0, -1000.0]
        self.assertEqual(sample(self.weights, self.vocab, limit=5), ("aaaaa", "limit"))
        with self.assertRaises(ValueError):
            sample(self.weights, self.vocab, limit=True)


if __name__ == "__main__":
    unittest.main()

The release also tested three isolated faults in copied source: adding instead of subtracting the gradient, dividing loss by T+1, and dropping the end target. Each produced a real assertion failure. A missing import or syntax error would not show that the learning rule was checked. After any learner mutation, restore the original implementation and rerun all tests before trusting the next training result.

A probability sum near one is necessary but weak evidence. Reversing the update sign still produces normalised probabilities. Dividing a loss by the wrong target count can preserve its downward trend while reporting the wrong scale. Omitting end targets can train plausible character transitions while teaching no proper stopping distribution. The known-answer, boundary and loss tests check different failure modes rather than all repeating a single implementation detail.

For a careful next experiment, hold the original data and test result as a recorded baseline. Change one design decision, document it before running and use validation data for iteration. A larger character context would distinguish more histories, but a direct table grows rapidly as context combinations increase. A trainable attention model needs embeddings, positions, projections, a vocabulary head, gradient propagation and an aligned next-token objective. None is secretly present in this score table.

This workbook completes a small-model workflow, not a frontier-training recipe. Moving to a framework should begin with tiny equivalence checks for loss, gradients, target alignment and sampling. New data requires provenance, document separation and a suitable held-out evaluation. Larger models require actual resource measurements. The previous attention course supplies a tested forward block; combining it with training remains a separate project with its own correctness checks.

Try it yourself · Activity 06

15 min

Catch and restore a defect

Work in a separate copy of the four-file lab.

  1. Run the ten tests before changing anything.
  2. Change the update subtraction to addition and observe the loss-reduction failure.
  3. Restore the code, rerun all tests and explain what the failure established.

What new test would become necessary if you replaced one-character context with a trainable attention model?

Worked answer

Changing row[j] -= rate*change to addition should fail test_one_update_reduces_loss. The function still returns finite, normalised distributions, but it takes the wrong direction on the known fixture. Restore subtraction and require all ten methods to pass. This establishes that the fixture detects this sign error; it does not prove every future architecture or dataset correct.

Keep learning

The complete workbook

Build a training-only vocabulary, keep document boundaries intact, derive the softmax cross-entropy gradient and run a fixed full-batch training schedule. Check gradients against numerical differences, compare held-out loss, sample with a local random generator and validate saved weights. Four complete Python files, original data, six worked activities and exact executed output make the experiment reproducible.

  1. 01
    Define a model small enough to understand

    Make the data, context and missing capabilities explicit.

    Read here · 1 exercise
  2. 02
    Turn separate documents into targets

    Count boundaries before counting accuracy.

    Read here · 1 exercise
  3. 03
    Derive and update the learning rule

    Follow one gradient before trusting a training curve.

    Read here · 1 exercise
  4. 04
    Sample and preserve the trained state

    Keep randomness, stopping and saved-file validation visible.

    Read here · 1 exercise
  5. 05
    Run the experiment and read the result

    Use held-out numbers and unedited samples together.

    Read here · 1 exercise
  6. 06
    Test failures before extending the model

    Keep the baseline executable while deciding what comes next.

    Read here · 1 exercise

Also inside: a 8-point checklist, a glossary of 8 terms and 10 questions and answers to test yourself. 6 hands-on exercises, each with a worked answer at the back where the workbook gives one.

No login, no card, no account. Before the download we ask you to follow Mickai (two quick links). Free to download and use for personal learning, study groups and inside your own team. Please do not resell the workbooks or republish them as your own. Link people to trust-agent.ai instead.

Test yourself

Questions and answers

Is this a transformer or an LLM?

No. It is a trained character bigram language model with a single preceding-character context and 812 stored scores. There is no attention or hidden layer.

Why does the weight table have an extra row?

The input-only beginning context predicts the first token. It is not part of the output vocabulary, so the table has V+1 rows and V columns.

Does end connect to the next document?

No. Every document resets to the beginning context. Its end is a target, and no transition joins it to the next document.

Can validation add characters to the vocabulary?

No. The vocabulary comes from training documents only. An unseen validation or test character maps to the existing unknown token.

What is the logit gradient?

For a pair count C, context count N, probability p and total target count T, the derivative is (N*p-C)/T.

Why do unseen context rows stay uniform?

Their context counts and target counts are zero, so their gradients are zero. They retain their initial zero logits.

Does the sampler always finish a sentence?

No. It stops on the end token or a hard character cap. Even an end-token stop does not establish a grammatical sentence.

Does the checkpoint contain the whole experiment?

No. It stores vocabulary, weights and a format version. Preserve the four source files, data split, fixed training schedule, runtime and result log separately.

Why is test loss lower than training loss?

The small sets have different transition mixtures. The result alone does not establish superior generalisation, data leakage or a meaningful broad-language capability.

What do the gradient checks establish?

All 624 finite-difference comparisons across 24 small corpora agreed within the stated tolerance. They verify bounded arithmetic examples, not every input, architecture or task.

When you have finished

Get your certificate of completion

Type your name and download a certificate for this workbook as a PDF, ready to print or to add to LinkedIn. It is made on your own device, so your name is never sent to us. It is a self-declared certificate, not an accredited qualification.

Learn the language

Key terms

Bigram model
A model that predicts the next token using only the preceding token.
Logit
A score before the softmax normalisation into probabilities.
Cross-entropy loss
Here, the mean negative log probability assigned to each observed next target.
Full-batch update
One parameter update using the gradient over the entire training set.
Held-out data
Documents excluded from parameter updates and vocabulary construction in this experiment.
Perplexity
The exponential of mean natural-log loss for the stated tokens and evaluation data.

6 of the workbook's 8 terms. The complete glossary is in the workbook.

Follow the evidence

Sources and checks

Facts last checked: .

Examples in this workbook were run on: Python 3.12.10 on Windows 11. Standard library only; fixed 400-update CPU training, held-out evaluation, seeded generation and checkpoint round trip executed. No dependency installation, GPU, account, provider, model download or network is required by the lab. (2026-09-27).

These workbooks use AI assistance. See how the workbooks are made.

  1. Speech and Language Processing, chapter 3: N-gram Language Models, draft 19 August 2026Dan Jurafsky and James H. Martin, Stanford University
  2. Python 3.12 mathematical functionsPython Software Foundation
  3. Python 3.12 random and reproducibilityPython Software Foundation
  4. Python 3.12 JSON encoding and decodingPython Software Foundation
  5. Python 3.12 unittestPython Software Foundation

Created by Mickarle Wagstaff-Irons - Micky Irons with the Mickai team. Published by Mickai LTD. Last updated 27 September 2026.

NextKeep going

Where to go next