models · Level 4

Transformers: attention explained and implemented

Trace the numbers from queries and keys to a causal, two-head transformer block.

By Mickarle Wagstaff-Irons - Micky Irons

  • Level 4Frontier
  • 180 min
  • 6 chapters
  • Free PDF, no account
The Structuremodels / 04

Start with the essentials

The short answer

Attention forms weights from queries and keys, then uses those weights to mix values. This course makes that calculation visible with small matrices, an explicit causal mask and a fixed two-head transformer block. All examples run locally without a GPU or account. The block has no training, gradients, tokenizer or language-model output head; it teaches the forward calculation, not a working LLM.

What you will learn

  • Track query, key, value and output shapes.
  • Calculate a scaled attention row and its weighted output.
  • Implement stable softmax and explicit visibility masks.
  • Check that future input cannot affect a causal prefix.
  • Assemble a fixed two-head transformer block with residual paths.
  • Distinguish a tested forward pass from a trained language model.

Who it is for

Builders comfortable with Python lists, dot products and the idea of a training loop who want to inspect attention before using a tensor framework.

Before you start

  • Read Python functions, nested lists and unit tests.
  • Understand vectors, dot products and matrix dimensions.
  • Recognise tokenisation, model parameters and a training loop.

Read a sample · Chapter 01 of 06

01

Read the shapes before the numbers

Each row has a role. Write it down before multiplying anything.

Create a folder containing attention.py, block.py, demo.py and test_attention.py. The complete files appear in this workbook. Run them with Python 3.12; the exercised version was 3.12.10. There are no dependencies to install. Keep these files in one folder so the imports resolve. Work on the tiny fixture first: increasing its size will make the arithmetic less visible without making it a useful language model.

A query is a vector used to compare against key vectors. Each key has a corresponding value vector. The comparison produces one score per key; the normalised scores determine how much of each value reaches the output. Q contains query rows, K key rows and V value rows. If Q is L by dk, K is S by dk and V is S by dv, the weights have shape L by S and the result has shape L by dv.

Our first fixture has two query rows, two key rows and two value features. Q and K are both [[1, 0], [0, 1]], while V is [[10, 0], [0, 20]]. These numbers are chosen for hand calculation, not obtained from text or training. The first query matches the first key more strongly, but the second key still receives positive weight when both are visible. An attention row is a weighted mixture rather than a mandatory single-item selection.

Self-attention uses representations from one sequence to form queries, keys and values. This does not mean their numbers must be identical: trainable projections usually give them different roles. Cross-attention can take queries from one sequence and keys and values from another. Our attention function permits different query and key lengths in noncausal mode. Its causal mode deliberately requires equal lengths, avoiding an ambiguous cache-offset convention in this small lab.

Start attention.py with the imports, input checker and matrix multiplication below, placing the second part after the first with blank lines between functions. The checker accepts only rectangular, bounded finite numeric matrices. The 64-row, 64-column and magnitude limits are teaching constraints, not transformer architecture limits. Booleans are rejected even though Python treats bool as a subclass of int. The checks keep a malformed shape from silently passing through zip.

python · 16 lines
"""Small forward-only attention lab; no training or tensor dependency."""
import math


def matrix(value, name):
    if type(value) not in (list, tuple) or not 1 <= len(value) <= 64:
        raise ValueError(f"{name}: expected 1 to 64 rows")
    if any(type(row) not in (list, tuple) for row in value):
        raise ValueError(f"{name}: expected rows")
    width = len(value[0])
    if not 1 <= width <= 64 or any(len(row) != width for row in value):
        raise ValueError(f"{name}: expected a non-empty rectangle")
    if any(type(x) not in (int, float) or abs(x) > 1e6
           or not math.isfinite(x) for row in value for x in row):
        raise ValueError(f"{name}: expected bounded finite numbers")
    return [[float(x) for x in row] for row in value]

python · 7 lines
def matmul(left, right):
    a, b = matrix(left, "left"), matrix(right, "right")
    if len(a[0]) != len(b):
        raise ValueError("Incompatible multiplication shapes")
    columns = list(zip(*b))
    return [[math.fsum(x*y for x, y in zip(row, col))
             for col in columns] for row in a]

For matmul, the left width must equal the right row count. Multiplying [[1, 2]] by [[3, 4], [5, 6]] gives [[13, 16]]: the first result is 1*3 + 2*5, and the second is 1*4 + 2*6. Our explicit loops prioritise traceability. They do not offer the execution speed or automatic differentiation of a tensor library. Matrix validation also does not establish that a chosen model architecture is useful.

Try it yourself · Activity 01

15 min

Write the shape contract

Use L=3, S=5, dk=2 and dv=4.

  1. Write Q, K, V, score and output dimensions.
  2. Explain which two dimensions must agree for comparing queries and keys.
  3. Calculate the one-row multiplication shown above.

Could you identify a transposed dimension from a shape error before inspecting individual numbers?

Worked answer

Q is 3 by 2, K is 5 by 2, V is 5 by 4, scores and weights are 3 by 5, and output is 3 by 4. Q and K share feature width 2; K and V share five source positions. The example multiplication returns [[13, 16]]. Different query and source lengths are valid for our noncausal function.

Read a sample · Chapter 02 of 06

02

Turn scores into a weighted mixture

Scale the comparisons, stabilise the exponential and preserve the row sum.

The scaled dot-product rule uses Q times the transpose of K, divides scores by the square root of dk, applies a row-wise softmax, and multiplies the resulting weights by V. The original transformer paper describes this operation and multi-head composition. Here we implement the arithmetic directly. A softmax row assigns nonnegative weights summing to one before any dropout; this lab implements no dropout.

In the first fixture, dk is two. Query [1, 0] has dot products 1 and 0 with the two keys. Dividing by sqrt(2) gives approximately 0.707107 and 0. Exponentiating and normalising gives approximately 0.669762 and 0.330238. The output is therefore [10*0.669762, 20*0.330238], approximately [6.697615, 6.604769]. Keep full precision in the program; rounding intermediate values changes the last digits.

Softmax is unchanged mathematically when the same constant is subtracted from every finite score in a row: the common exponential factor cancels from numerator and denominator. Subtracting the largest score makes the largest exponential one. Thus [10000, 10000] can yield [0.5, 0.5] without evaluating exp(10000). Small terms may still underflow towards zero. This routine is a controlled floating-point teaching implementation, not an arbitrary-precision calculator.

Append softmax and attention to attention.py below. Negative infinity represents a disallowed key; its exponential contribution is zero. An entirely masked row is rejected because there is no visible value to mix. NaN and positive infinity are invalid scores. The public attention function validates bounded finite matrices before constructing scores; softmax is a small internal helper with a narrower input contract.

python · 9 lines
def softmax(scores):
    if not scores or any(math.isnan(x) or x == math.inf for x in scores):
        raise ValueError("Invalid scores")
    peak = max(scores)
    if peak == -math.inf:
        raise ValueError("No visible key in this row")
    values = [math.exp(x - peak) for x in scores]
    total = math.fsum(values)
    return [x / total for x in values]

python · 25 lines
def attention(query, key, value, *, causal=False, allow=None):
    q, k, v = (matrix(x, name) for x, name in
               ((query, "query"), (key, "key"), (value, "value")))
    nq, nk, dk, dv = len(q), len(k), len(q[0]), len(v[0])
    if len(k[0]) != dk or len(v) != nk:
        raise ValueError("Q/K widths and K/V row counts must agree")
    if type(causal) is not bool or (causal and nq != nk):
        raise ValueError("Causal mode requires equal sequence lengths")
    if allow is None:
        allow = [[True] * nk for _ in range(nq)]
    if (type(allow) not in (list, tuple) or len(allow) != nq
            or any(type(row) not in (list, tuple) or len(row) != nk
                   or any(type(x) is not bool for x in row)
                   for row in allow)):
        raise ValueError("Expected a boolean allow matrix")
    weights, output = [], []
    for i, row in enumerate(q):
        scores = [math.fsum(a*b for a, b in zip(row, col)) / math.sqrt(dk)
                  if allow[i][j] and (not causal or j <= i) else -math.inf
                  for j, col in enumerate(k)]
        probs = softmax(scores)
        weights.append(probs)
        output.append([math.fsum(probs[j] * v[j][d] for j in range(nk))
                       for d in range(dv)])
    return output, weights

The division uses the query/key width dk, not the number of tokens, the value width or the number of heads. Changing the sequence length therefore does not change this scaling factor when dk stays fixed. For our first row, dropping the scale would produce about 0.731059 instead of 0.669762 for the first key. That is a different distribution even though it still sums to one. A row-sum test alone would miss the defect.

Each output feature is a weighted average of that feature across visible value rows. It must lie between their minimum and maximum, up to floating-point tolerance. This is a useful local invariant for the attention output. It need not hold after an output projection, residual addition or feed-forward transformation, where other arithmetic changes the values. A correct invariant must name the stage to which it applies.

Try it yourself · Activity 02

20 min

Calculate one attention row

Work out the first noncausal row using the two-position fixture.

  1. Compute its two scaled scores.
  2. Calculate the softmax weights and both output features.
  3. Repeat the weight calculation without scaling and explain why row sums cannot catch that error.

Which dimension would you inspect first if all your weights were consistently too concentrated?

Worked answer

The scores are 1/sqrt(2) and 0. The first weight is exp(1/sqrt(2))/(exp(1/sqrt(2))+1), approximately 0.669762; the second is approximately 0.330238. The output is about [6.697615, 6.604769]. Without scaling, the first weight is about 0.731059. Both versions sum to one, so comparison with a known calculation is necessary.

Read a sample · Chapter 03 of 06

03

Make visibility a tested contract

Mask the scores before normalisation and distinguish causal order from padding.

Hide the future before mixing values

The first query has weights 0.669762 and 0.330238 without a causal mask. Hiding future key 1 changes them to 1 and 0, so the output becomes [10, 0].
Follow query zero through the exact two-position course fixture. Numeric labels, position IDs and the word MASKED carry the meaning alongside the gold bars. Open the full-size attention diagram.

The query rows are [1, 0] and [0, 1], in that order. Keys and values are paired by position. These are original fixed teaching numbers, not learned word embeddings or measured model performance.

Exact input keys and paired value vectors
PositionKeyValue
0[1, 0][10, 0]
1[0, 1][0, 20]

For query zero, dot products are [1, 0]. Divide by sqrt(2), the square root of the query/key feature width, giving scores [0.707107, 0] rounded to six decimals. Noncausal softmax gives [0.669762, 0.330238]. In causal mode, only key zero is visible to this query; key one's score becomes negative infinity before softmax, giving weights [1, 0]. Zeroing the score would not hide a key because exp(0)=1.

Every attention weight: six-decimal display, unrounded calculation
Mode / queryKey 0 weight Key 1 weight
Noncausal, query 00.6697620.330238
Noncausal, query 10.3302380.669762
Causal, query 01.0000000.000000
Causal, query 10.3302380.669762
Every weighted output feature, before any projection or residual path
Mode / queryFeature 0 Feature 1
Noncausal, query 06.6976156.604769
Noncausal, query 13.30238513.395231
Causal, query 010.0000000.000000
Causal, query 13.30238513.395231

Each output is w0 * [10, 0] + w1 * [0, 20]. Use unrounded weights when calculating; rounding the displayed weights first changes the last digits. Query one's results are unchanged because both positions remain visible. Our causal rule is j <= i: the current position is included.

The explicit allow mask is a separate visibility condition, with True meaning visible. If its intersection with the causal rule removes every key in a row, the lab raises ValueError; it does not fabricate a probability distribution. Attention weights describe this numeric mixture, not the truth of an answer. The diagram covers the two-position attention primitive; the course's later three-position, two-head block is a separate example.

In our causal mode, query position i may see key positions j less than or equal to i. The diagonal is visible. This convention fits a next-token training arrangement where an input position predicts the following token: seeing the current input token does not expose the next target token. An incorrectly aligned target can still leak the answer even with a correct triangular mask. The mask and the input/target shift must be designed together.

For the two-position fixture, the first query sees only value [10, 0], so its weights become [1, 0] and its output is exactly [10, 0]. The second query sees both values, so its result matches the second row of the noncausal example. Setting a prohibited score to zero would be wrong: exp(0) is one, so the supposedly hidden position could still contribute. Mask before softmax, then normalise over the positions that remain.

allow is an explicit boolean matrix with True meaning visible. It can exclude a padding key or an application-defined position. Effective visibility combines allow with the causal condition. If their intersection removes every key in a query row, the function raises ValueError. Padding keys and padding queries are different issues: our function masks keys but does not silently discard query outputs or training-loss positions. Those policies belong in the surrounding pipeline.

Save demo.py below. It runs both visibility modes, then uses the fixed block introduced in the next chapter. Create block.py before running python demo.py. The output below is captured from those exact files. It records rounded values for inspection, while the tests compare unrounded calculations with numerical tolerance. Rows and columns always retain their original positional order.

python · 23 lines
from attention import attention
from block import X, transformer_block

Q = [[1, 0], [0, 1]]
K = [[1, 0], [0, 1]]
V = [[10, 0], [0, 20]]


def show(label, rows):
    print(label)
    for row in rows:
        print(" ".join(f"{x: .6f}" for x in row))


if __name__ == "__main__":
    for causal in (False, True):
        output, weights = attention(Q, K, V, causal=causal)
        show(f"causal={causal} weights", weights)
        show("output", output)
    output, maps = transformer_block(X)
    show("fixed block output", output)
    for index, weights in enumerate(maps, 1):
        show(f"head {index}", weights)

text · 24 lines
causal=False weights
 0.669762  0.330238
 0.330238  0.669762
output
 6.697615  6.604769
 3.302385  13.395231
causal=True weights
 1.000000  0.000000
 0.330238  0.669762
output
 10.000000  0.000000
 3.302385  13.395231
fixed block output
 0.999998 -0.999998  0.999998 -0.999998
-0.999998  0.999998 -0.999998  0.999998
 1.347147  0.577349 -0.962248 -0.962248
head 1
 1.000000  0.000000  0.000000
 0.330238  0.669762  0.000000
 0.248255  0.248255  0.503490
head 2
 1.000000  0.000000  0.000000
 0.330238  0.669762  0.000000
 0.333333  0.333333  0.333333

Framework mask conventions must be checked at the actual API boundary. The reviewed PyTorch scaled_dot_product_attention documentation uses True for participating elements in its boolean attention mask. MultiheadAttention uses True for excluded positions in its boolean masks. Those meanings are opposite. The functional API also applies the dropout probability supplied to it, so evaluation code must supply zero when dropout should be disabled. No PyTorch call is executed in this workbook.

A future-change experiment gives stronger evidence than a triangular-looking display. Change only a later input, run the causal block again, and compare the earlier outputs. They should remain unchanged for this deterministic fixture. Run the same idea without the causal constraint and the later input may affect earlier outputs. This test checks the direction of information flow; it does not prove that a trained model follows every application instruction.

Try it yourself · Activity 03

20 min

Exclude a key and perturb the future

Compare causal visibility with an explicit allow mask.

  1. Evaluate Q, K and V with allow=[[False, True]]*2 and causal=False.
  2. Explain why that mask is invalid for the first row when causal=True.
  3. Predict which outputs may change when only the last block input changes.

Does your application distinguish an intentionally padded query from an accidental all-masked row?

Worked answer

Without the causal constraint, both queries have weights [0, 1] and output [0, 20]. With causal=True, the first query cannot see position one and allow excludes position zero, leaving no visible key; the function raises ValueError. Changing only the last block input can change the last output, while the causal prefix stays identical in this fixed computation.

Read a sample · Chapter 04 of 06

04

Assemble a small transformer block

Keep the separate heads visible, then trace the two residual paths.

Save block.py using the four consecutive parts below. X has three positions and four features. Treat these rows as representations already supplied to the block, with any positional information handled upstream. They are not word embeddings learned from a corpus. The first fixed projection selects features zero and one; the second selects features two and three. For transparency, each head reuses its selected vectors as Q, K and V. A trainable implementation would normally provide distinct learned projections.

python · 13 lines
"""A fixed two-head, post-normalisation teaching block, not a trained model."""
import math
from attention import attention, matmul, matrix

X = [[1, 0, 1, 0], [0, 1, 0, 1], [1, 1, 0, 0]]
P1 = [[1, 0], [0, 1], [0, 0], [0, 0]]
P2 = [[0, 0], [0, 0], [1, 0], [0, 1]]
WO = [[1, 0, 0, 0], [0, 1, 0, 0],
      [0, 0, 1, 0], [0, 0, 0, 1]]
W1 = [[1, -1, 0, 0], [0, 1, -1, 0],
      [0, 0, 1, -1], [-1, 0, 0, 1]]
W2 = [[.5, 0, 0, 0], [0, .5, 0, 0],
      [0, 0, .5, 0], [0, 0, 0, .5]]

Each head returns three rows with two features. Concatenating corresponding rows gives three rows with four features. WO is an explicit identity output projection, so it preserves this fixture’s concatenated values; retaining the multiplication shows where a learned output projection would act. Changing WO is a separate experiment from changing attention. It should not be confused with adding another attention head.

The block adds the attention output to its input and normalises each position across its four features. It then applies W1, a ReLU activation and W2 independently to each row, adds that feed-forward result to the hidden representation, and normalises again. This is a post-normalisation teaching layout. The original transformer describes residual connections and layer normalisation; architectures can choose different normalisation placement. The fixed implementation here specifies its own exact sequence.

python · 7 lines
def normalise(rows):
    result = []
    for row in rows:
        mean = math.fsum(row) / len(row)
        variance = math.fsum((x-mean)**2 for x in row) / len(row)
        result.append([(x-mean) / math.sqrt(variance + 1e-5) for x in row])
    return result

python · 2 lines
def add(a, b):
    return [[x+y for x, y in zip(ra, rb)] for ra, rb in zip(a, b)]

python · 18 lines
def transformer_block(rows, *, causal=True):
    x = matrix(rows, "block input")
    if len(x[0]) != 4:
        raise ValueError("This fixed block expects four features")
    heads, maps = [], []
    for projection in (P1, P2):
        projected = matmul(x, projection)
        head, weights = attention(projected, projected, projected,
                                  causal=causal)
        heads.append(head)
        maps.append(weights)
    joined = [a+b for a, b in zip(*heads)]
    mixed = matmul(joined, WO)
    hidden = normalise(add(x, mixed))
    expanded = matmul(hidden, W1)
    activated = [[max(0.0, x) for x in row] for row in expanded]
    feed_forward = matmul(activated, W2)
    return normalise(add(hidden, feed_forward)), maps

normalise uses the population variance across a row and adds 1e-5 inside the square root. Its affine gain is fixed to one and bias to zero, so it is a simplified layer normalisation with no trainable affine parameters. A constant row maps to zeros rather than dividing by zero. Because epsilon is positive, a normalised row’s variance is generally slightly below one. Checking exact unit variance would be an inappropriate test for this implementation.

The feed-forward width remains four in this deliberately compact example. There are no bias vectors, dropout, batches, learned positions or gradients. ReLU retains a positive number and replaces a negative one with zero. W2 then halves each activated feature. Together with the residual path, these operations change the representation without mixing positions. The attention operation is the part of this block that moves information between positions.

The two head maps differ on the third row. Head one has weights about [0.248255, 0.248255, 0.503490]; head two has [1/3, 1/3, 1/3], because its last projected query is [0, 0]. A uniform row is a predictable result here, not a malfunction. The final third output is approximately [1.347147, 0.577349, -0.962248, -0.962248]. These results are from fixed numbers, not evidence of language understanding.

Try it yourself · Activity 04

20 min

Trace the two heads

Follow the third row of X through the projections.

  1. Write its query for head one and head two.
  2. Explain the uniform row in the second head.
  3. Name the operations after head concatenation, in order.

Which of those steps mixes positions, and which transforms features within a single position?

Worked answer

The third row projects to [1, 1] in head one and [0, 0] in head two. The latter has zero dot product with every key, giving equal causal weights over its three visible positions. Concatenate the two head outputs, multiply by WO, add X, normalise, multiply by W1, apply ReLU, multiply by W2, add the hidden representation and normalise again.

Read a sample · Chapter 05 of 06

05

Test the calculation and expose defects

Use known answers and information-flow checks as well as numerical invariants.

Save test_attention.py and run python -m unittest -v test_attention from the folder containing all four files. The twelve methods cover the hand calculation, causal direction, mask renormalisation, empty visibility, stable exponentials, matrix shapes, invalid inputs, a joint key/value permutation, constant values, multiplication, block shape and future independence. The runner should finish with OK. Keep the failing assertion when a test disagrees; do not round the program until a test happens to pass.

python · 94 lines
import math
import unittest
from attention import attention, matmul, softmax
from block import X, transformer_block
from demo import Q, K, V


class AttentionTests(unittest.TestCase):
    def test_hand_calculated_weights_and_values(self):
        out, weights = attention(Q, K, V)
        a = math.exp(1/math.sqrt(2)) / (math.exp(1/math.sqrt(2)) + 1)
        self.assertAlmostEqual(weights[0][0], a)
        self.assertAlmostEqual(out[0][0], 10*a)
        self.assertAlmostEqual(out[0][1], 20*(1-a))

    def test_causal_direction_and_diagonal(self):
        out, weights = attention(Q, K, V, causal=True)
        self.assertEqual(weights[0], [1, 0])
        self.assertEqual(out[0], [10, 0])
        self.assertGreater(weights[1][0], 0)
        self.assertGreater(weights[1][1], 0)

    def test_mask_renormalises_visible_keys(self):
        out, weights = attention(Q, K, V, allow=[[False, True]]*2)
        self.assertEqual(weights, [[0, 1], [0, 1]])
        self.assertEqual(out, [[0, 20], [0, 20]])

    def test_no_visible_key_is_an_error(self):
        for mask in ([[False, False]]*2, [[False, True], [True, True]]):
            with self.assertRaises(ValueError):
                attention(Q, K, V, causal=True, allow=mask)

    def test_softmax_shift_and_large_scores(self):
        self.assertEqual(softmax([10000, 10000]), [.5, .5])
        for a, b in zip(softmax([1, 2, 3]), softmax([1001, 1002, 1003])):
            self.assertAlmostEqual(a, b)

    def test_shape_contracts(self):
        with self.assertRaises(ValueError):
            attention(Q, [[1]], V)
        with self.assertRaises(ValueError):
            attention(Q, K, [[1]])
        out, _ = attention([[0, 0]], K, [[1, 2, 3], [3, 4, 5]])
        self.assertEqual(out, [[2, 3, 4]])
        with self.assertRaises(ValueError):
            attention([[0, 0]], K, V, causal=True)

    def test_invalid_numbers_and_masks(self):
        for bad in ([], [[True, 0]], [[math.nan, 0]], [[math.inf, 0]],
                    [[1e7, 0]], [[1, 2], [3]]):
            with self.assertRaises(ValueError):
                attention(bad, K, V)
        for mask in ([[1, 1]]*2, [[True]], "all"):
            with self.assertRaises(ValueError):
                attention(Q, K, V, allow=mask)

    def test_joint_key_value_permutation(self):
        out, _ = attention(Q, K, V)
        permuted, _ = attention(Q, K[::-1], V[::-1])
        for row, other in zip(out, permuted):
            for a, b in zip(row, other):
                self.assertAlmostEqual(a, b)

    def test_constant_values_remain_constant(self):
        out, _ = attention(Q, K, [[3, -2], [3, -2]], causal=True)
        for row in out:
            for a, b in zip(row, [3, -2]):
                self.assertAlmostEqual(a, b)

    def test_matrix_multiplication(self):
        self.assertEqual(matmul([[1, 2]], [[3, 4], [5, 6]]), [[13, 16]])
        with self.assertRaises(ValueError):
            matmul([[1, 2]], [[3, 4]])

    def test_block_shape_and_normalisation(self):
        out, maps = transformer_block(X)
        self.assertEqual((len(out), len(out[0]), len(maps)), (3, 4, 2))
        self.assertAlmostEqual(out[0][0], 1, places=5)
        self.assertAlmostEqual(out[0][1], -1, places=5)
        for row in out:
            self.assertAlmostEqual(sum(row), 0)
            self.assertTrue(all(math.isfinite(x) for x in row))

    def test_future_change_cannot_change_block_prefix(self):
        changed = [row[:] for row in X]
        changed[2] = [10, -9, 8, -7]
        original, _ = transformer_block(X)
        altered, _ = transformer_block(changed)
        self.assertEqual(original[:2], altered[:2])
        self.assertNotEqual(original[2], altered[2])


if __name__ == "__main__":
    unittest.main()

Reordering key rows and their corresponding value rows together should preserve a noncausal output, apart from small floating-point effects, because the same pairings contribute to the same query. Reordering keys alone changes those pairings and is not a valid invariance. A causal mask refers to positional order, so the same permutation argument cannot be transferred unchanged to causal attention. State the assumptions behind each test.

The fixed-value test checks another useful property. If every visible value row is [3, -2], their weighted mixture is [3, -2] regardless of the query scores. The future-independence test exercises the complete block, including both residual paths and normalisation. Per-row normalisation preserves that information-flow boundary; normalising across sequence positions could break it. A passing attention primitive alone would not detect every error in a larger block.

The release verification also compared 120 small cases against a separate 50-digit Decimal calculation, with 3,722 additional assertions. It varied lengths one to five, query/key widths one to four, value widths one to three, visibility patterns and causal mode. The reference used unshifted high-precision exponentials rather than the float implementation’s max-subtraction path. This checks forward arithmetic over a bounded fixture family; it is not a test of gradients, GPU kernels or model quality.

Three fault experiments ran in separate copies: remove the square-root scale, reverse the causal inequality, and ignore the allow mask. Each produced actual assertion failures. An import error or syntax error would not establish that the numerical rule was tested. Restore the baseline after an experiment and rerun the complete suite. The delivered source contains the passing implementation, with the failed mutation logs retained only as review evidence.

Try it yourself · Activity 05

15 min

Demonstrate a meaningful failure

Work in a copy of the four-file folder.

  1. Run all twelve tests before editing.
  2. Replace the scale divisor with 1.0 and identify the failed hand-calculation assertion.
  3. Restore attention.py, rerun the suite and record why a row-sum check was insufficient.

What wrong implementation could satisfy your current invariants while returning different outputs?

Worked answer

Removing the scale changes the fixture’s first weight from about 0.669762 to 0.731059, so the known-answer comparison fails. The probabilities still sum to one and remain nonnegative. Restore division by sqrt(dk) and verify all twelve methods pass. The useful evidence is an arithmetic assertion failure followed by a passing restored baseline.

Read a sample · Chapter 06 of 06

06

Connect the block to a language model

List the missing interfaces before adding data, gradients or a larger machine.

This block returns one four-feature representation per input position. It does not return token probabilities. To make a small language model, a surrounding system needs token IDs, an embedding lookup, a positional scheme, trainable parameters, an output projection to vocabulary logits, an appropriate next-token loss and a gradient-based update. A generation loop must feed the produced token back into the next step. Those components are planned follow-on work, not hidden capabilities of these four files.

Write the input/target alignment explicitly. For a toy sequence [a, b, c, d], one training pair could use inputs [a, b, c] and targets [b, c, d]. Position zero can see a while predicting b; it must not receive b through an input feature, leaked target or incorrect mask. Special start/end tokens, padding and loss masking need their own definitions. An attention mask does not automatically remove padding from the loss.

A visibility mask provides an order-dependent boundary, but it is not a substitute for a deliberate representation of position. Decide how position enters the model and keep that decision consistent between training and generation. This course’s X is an explicit representation fixture; there is no positional encoding algorithm to infer from its numbers. A future trainable implementation should test repeated tokens at different positions and document the selected scheme.

Our function constructs an L by S weight table. With equal sequence lengths n, this table has n*n entries per head: four times as many when n doubles. Two heads and a batch dimension multiply the count further. This arithmetic describes the explicit table in the teaching code, not every optimised implementation’s peak memory. Python object overhead is also very different from a packed tensor. Do not turn this small script into a GPU purchasing estimate.

Attention weights reveal how this forward operation mixes value vectors. They are not a probability that a source statement is true and do not establish a complete explanation of a model’s answer. In our fixture the values are numbers with no textual claim attached at all. A language model adds many transformations before and after attention. Keep numerical inspection, predictive evaluation and factual evaluation as separate questions.

For a framework port, record tensor dimension order, parameter initialisation, mask polarity, dropout behaviour, normalisation axes, epsilon and dtype. Compare a tiny fixed forward case first, then add gradient and optimisation tests. Use a held-out data split for model evaluation and record data provenance. This release ran no framework port, gradient check, training benchmark or text generation. Its completed result is a reproducible attention calculation and a precise specification for the next step.

Try it yourself · Activity 06

15 min

Define the next integration boundary

Write a one-page plan for extending the fixture into a tiny language model.

  1. Specify input IDs and shifted targets for a four-token example.
  2. List the trainable components missing from this block.
  3. Name a forward equivalence check, a gradient check and a held-out evaluation separately.

Could another builder tell exactly what has been executed and what remains a proposed experiment?

Worked answer

A simple alignment is inputs [a, b, c] and targets [b, c, d] with causal diagonal visibility. Add embeddings, a positional scheme, distinct trainable projections, normalisation parameters as chosen, a vocabulary output layer and a next-token loss with optimisation. First reproduce this fixed forward fixture in the chosen framework, then check gradients on a tiny example, then evaluate a trained model on held-out data. Each stage needs its own evidence before being reported complete.

Keep learning

The complete workbook

Work through matrix shapes, stable softmax, scaled dot products, boolean masks, two attention heads, residual paths and normalisation. Run four complete Python files, twelve learner tests and original numerical examples, then specify the missing pieces of a trainable language model.

  1. 01
    Read the shapes before the numbers

    Each row has a role. Write it down before multiplying anything.

    Read here · 1 exercise
  2. 02
    Turn scores into a weighted mixture

    Scale the comparisons, stabilise the exponential and preserve the row sum.

    Read here · 1 exercise
  3. 03
    Make visibility a tested contract

    Mask the scores before normalisation and distinguish causal order from padding.

    Read here · 1 exercise
  4. 04
    Assemble a small transformer block

    Keep the separate heads visible, then trace the two residual paths.

    Read here · 1 exercise
  5. 05
    Test the calculation and expose defects

    Use known answers and information-flow checks as well as numerical invariants.

    Read here · 1 exercise
  6. 06
    Connect the block to a language model

    List the missing interfaces before adding data, gradients or a larger machine.

    Read here · 1 exercise

Also inside: a 8-point checklist, a glossary of 8 terms and 10 questions and answers to test yourself. 6 hands-on exercises, each with a worked answer at the back where the workbook gives one.

No login, no card, no account. Before the download we ask you to follow Mickai (two quick links). Free to download and use for personal learning, study groups and inside your own team. Please do not resell the workbooks or republish them as your own. Link people to trust-agent.ai instead.

Test yourself

Questions and answers

Do queries, keys and values need the same shape?

Queries and keys need equal feature widths, and keys and values need equal source lengths. Value width can differ. This lab requires equal query and key lengths only when causal=True.

Why does the checker reject boolean matrix entries?

They are usually a mistaken mask or input representation in this numeric fixture. Python allows booleans to behave as integers, so the checker rejects them explicitly.

Which dimension sets the scale?

The shared query/key feature width dk. The scores are divided by sqrt(dk), not by sequence length or value width.

Why subtract the largest score?

It leaves softmax mathematically unchanged while keeping the largest exponential at one, avoiding overflow from large positive scores in this bounded lab.

Why not replace hidden scores with zero?

exp(0) contributes positive mass. A hidden position must receive no mass before normalisation; this implementation uses negative infinity.

What happens when every key is masked?

The function raises ValueError. It does not silently invent a uniform output or report an all-zero row as a valid probability distribution.

Are the two heads trained?

No. They select different pairs of input features using fixed matrices and reuse each selected representation for Q, K and V. They demonstrate forward composition only.

Does the block include full trainable layer normalisation?

No. It uses per-row population variance and epsilon 1e-5, with gain fixed to one and bias fixed to zero. There are no trainable affine parameters.

What does the future-change test establish?

Changing only the final input leaves earlier outputs unchanged in this deterministic causal block. It checks information flow, not language quality or gradient correctness.

Is this a complete language model?

No. It lacks tokenisation, learned embeddings and positions, a vocabulary output layer, loss, gradients, optimisation and generation. Those are separate implementation and evaluation steps.

When you have finished

Get your certificate of completion

Type your name and download a certificate for this workbook as a PDF, ready to print or to add to LinkedIn. It is made on your own device, so your name is never sent to us. It is a self-declared certificate, not an accredited qualification.

Learn the language

Key terms

Query
A vector compared against keys to determine one attention row.
Key
A vector paired with a value and scored against a query.
Value
A vector included in the output mixture according to its attention weight.
Softmax
An exponential normalisation producing nonnegative weights that sum to one.
Causal mask
A visibility rule that, here, permits a position to use itself and earlier positions.
Attention head
One projected query/key/value calculation before head outputs are combined.

6 of the workbook's 8 terms. The complete glossary is in the workbook.

Follow the evidence

Sources and checks

Facts last checked: .

Examples in this workbook were run on: Python 3.12.10 on Windows 11. Standard-library forward arithmetic with fixed original matrices; no package install, GPU, network, model download or training run. PyTorch documentation was consulted, but PyTorch was not executed. (2026-09-27).

These workbooks use AI assistance. See how the workbooks are made.

  1. Attention Is All You Need, sections 3.1-3.2Vaswani et al., arXiv HTML
  2. Scaled dot product attentionPyTorch 2.14 documentation
  3. MultiheadAttentionPyTorch 2.14 documentation
  4. Python 3.12 mathematical functionsPython Software Foundation
  5. Python 3.12 unittestPython Software Foundation

Created by Mickarle Wagstaff-Irons - Micky Irons with the Mickai team. Published by Mickai LTD. Last updated 27 September 2026.

NextKeep going

Where to go next