models · Level 3

Tensors, gradients and a first training loop

Train tiny models on a CPU in seconds, by hand and then with automatic differentiation, and see what training is and is not.

By Mickarle Wagstaff-Irons - Micky Irons

  • Level 3Building
  • 95 min
  • 6 chapters
  • Free PDF, no account
The Structuremodels / 03

Start with the essentials

The short answer

Training repeats one loop: run the model on examples, score how wrong it is with a single number called the loss, compute the gradient (which way each parameter should move to lower that loss), and take a small step that way. Tiny models train on an ordinary CPU in seconds. Published frontier-scale language-model runs use the same idea with vastly more data, hardware, engineering and cost.

What you will learn

  • You will be able to read a tensor's shape and dtype and predict when broadcasting will stretch it, including the cases that fail silently.
  • You will be able to compute a mean squared error loss and its gradient by hand, and check it by two independent routes.
  • You will be able to spot a learning rate that is too small or too large from the loss alone.
  • You will be able to write the PyTorch training loop with batches and epochs, and train a tiny network on non-linear data.
  • You will be able to separate training from validation, recognise overfitting and find the classic training bugs.
  • You will be able to explain, with sources, what changes when language models are pretrained at frontier scale.

Who it is for

People comfortable writing small Python programs who want to understand what training a model actually means by doing it. You need a computer that can run Python 3.12 (the version tested here) and install two packages; on a Mac, PyTorch 2.14.0 needs Apple silicon and macOS 14 or later. No GPU, no account and no dataset download.

Before you start

  • Comfort with Python basics: variables, loops, functions and running a script from a terminal, as taught in Python: from first script to a useful automation. School algebra is enough; the one idea from calculus you need, a slope, is explained as you go.

Read a sample · Chapter 02 of 06

02

The loss and the gradient, by hand

Training needs one number that says how wrong the model is, and a way to tell which direction makes it smaller.

toy.py invents 20 points near y = 3x + 2. Its seed fixes the random numbers, so every run makes the same data. The model is a line with two parameters, weight w and bias b: prediction = w * x + b.

python · 15 lines
# toy.py: data, loss and gradient
import numpy as np

rng = np.random.default_rng(seed=0)
x = rng.uniform(-1, 1, size=20)
y = 3 * x + 2 + rng.normal(0, 0.1, size=20)   # a line plus noise


def loss(w, b):
    return np.mean((w * x + b - y) ** 2)


def gradient(w, b):
    err = w * x + b - y
    return np.mean(2 * err * x), np.mean(2 * err)

The loss scores 'how wrong' as one number. Here it is the mean squared error: the average of (prediction minus target) squared, so big misses count most.

The gradient holds the loss's slope for each parameter. It points uphill, says the Deep Learning textbook, so its negative points downhill. Differentiating gives the slopes in gradient: the mean of 2 * err * x for w, and of 2 * err for b.

Always check a hand-derived gradient, for instance by nudging each parameter slightly both ways and measuring the change in loss: a finite difference.

python · 7 lines
# check.py
from toy import loss, gradient

w, b, h = 0.5, -1.0, 1e-6          # h is a tiny nudge
print("by hand %.6f %.6f" % gradient(w, b))
print("nudged  %.6f %.6f" % ((loss(w + h, b) - loss(w - h, b)) / (2 * h),
                             (loss(w, b + h) - loss(w, b - h)) / (2 * h)))

text · 2 lines
by hand -2.194225 -6.135977
nudged  -2.194225 -6.135977

They agree to six decimals. Both slopes are negative, so raising w and b lowers the loss: the true 3 and 2 are above 0.5 and -1.

Try it yourself · Activity 02

15 min

One step by hand

With pencil and paper, take the points x = 1, 2, 3 and y = 5, 8, 11 (exactly on y = 3x + 2), starting at w = 0, b = 0.

  1. Work out the errors (prediction minus target) and the loss.
  2. Work out the slopes for w and b.
  3. Take one step with learning rate 0.1 (subtract 0.1 times each slope), find the new loss, and check it in Python.

Why did one step help so much? Would it always?

Worked answer

Errors -5, -8, -11; loss (25 + 64 + 121) / 3 = 70. Slope for w: 2 x (-5 - 16 - 33) / 3 = -36. Slope for b: 2 x (-24) / 3 = -16. New w = 3.6, b = 1.6; new errors 0.2, 0.8, 1.4; loss (0.04 + 0.64 + 1.96) / 3 = 0.88. With x = np.array([1, 2, 3]) and y = np.array([5, 8, 11]), print(np.mean((3.6 * x + 1.6 - y) ** 2)) prints 0.8800000000000008, a floating-point rounding artefact.

Read a sample · Chapter 04 of 06

04

Automatic differentiation and the training loop

Hand-derived gradients do not scale. Let PyTorch compute them, check it, then write the standard training loop.

Automatic differentiation (autograd) records the forward calculation as a graph, then applies the chain rule back through it; that backward sweep is backpropagation. Mark a tensor requires_grad=True, call loss.backward(), and its slope appears in .grad. This uses check.py's point, in float64 to match NumPy.

python · 12 lines
# autograd_check.py
import torch
from toy import x, y, gradient

xt, yt = torch.tensor(x), torch.tensor(y)
w = torch.tensor(0.5, dtype=torch.float64, requires_grad=True)
b = torch.tensor(-1.0, dtype=torch.float64, requires_grad=True)
print("by hand    %.6f %.6f" % gradient(0.5, -1.0))
for run in ["autograd  ", "not zeroed"]:
    loss = ((w * xt + b - yt) ** 2).mean()      # forward and loss
    loss.backward()                              # backward
    print(run, "%.6f %.6f" % (w.grad.item(), b.grad.item()))

text · 3 lines
by hand    -2.194225 -6.135977
autograd   -2.194225 -6.135977
not zeroed -4.388451 -12.271953

Three routes agree; in full, autograd and hand differ by under 1e-15, float64 rounding. The last line is a classic bug: the backward documentation says gradients accumulate, so a second pass without clearing doubled them.

Now the real loop, in float32, PyTorch's default. Linear(1, 1) is the same line, MSELoss the same loss, SGD the same update. A batch is the examples behind one update; an epoch is one pass through the data, here 4 batches of 5.

python · 23 lines
# train_line.py
import torch
from toy import x, y

torch.manual_seed(0)
X = torch.tensor(x, dtype=torch.float32).reshape(20, 1)
Y = torch.tensor(y, dtype=torch.float32).reshape(20, 1)
model = torch.nn.Linear(1, 1)
loss_fn = torch.nn.MSELoss()
opt = torch.optim.SGD(model.parameters(), lr=0.1)

for epoch in range(1, 31):
    order = torch.randperm(20)                 # shuffle each epoch
    for start in range(0, 20, 5):              # batches of 5
        batch = order[start:start + 5]
        opt.zero_grad()                        # 1 clear
        loss = loss_fn(model(X[batch]), Y[batch])  # 2 forward, 3 loss
        loss.backward()                        # 4 backward
        opt.step()                             # 5 step
    if epoch % 10 == 0:
        with torch.no_grad():
            print(epoch, round(loss_fn(model(X), Y).item(), 5))
print(round(model.weight.item(), 3), round(model.bias.item(), 3))

text · 4 lines
10 0.00759
20 0.00452
30 0.0045
2.971 2.007

The numbered steps are the training loop; PyTorch's documentation clears gradients first, though clearing after the step works too. torch.no_grad() stops tracking while you only measure, and .reshape(20, 1) gives targets the predictions' shape.

Try it yourself · Activity 04

10 min

Change the batch size

In train_line.py, change the two 5s on the batch lines to 1, run it, then try 20.

  1. First work out how many updates each version makes in 30 epochs.
  2. Compare the losses at epochs 10, 20 and 30.

Which would you choose if each update were expensive?

Worked answer

Batches of 1: 600 updates; losses 0.0046, 0.00557, 0.00468. Fast but jittery: each step sees one noisy point. Batches of 20: 30 updates; losses 0.65896, 0.11635, 0.02439, ending 2.751 2.014. Smooth, but unfinished.

Keep learning

The complete workbook

This workbook teaches training by doing it at toy scale. You will compute a gradient by hand in NumPy, confirm it with PyTorch's automatic differentiation, then fit a straight line and a tiny network while watching the learning rate, validation loss and common bugs. It ends with a sourced account of what changes when language models are pretrained at frontier scale.

  1. 01
    Tensors: shape, dtype and broadcasting

    Everything a model reads, stores and learns is a grid of numbers. Learn to read one before trusting any result.

    In the workbook · 1 exercise
  2. 02
    The loss and the gradient, by hand

    Training needs one number that says how wrong the model is, and a way to tell which direction makes it smaller.

    Read here · 1 exercise
  3. 03
    Gradient descent and the learning rate

    Gradient descent repeats one move: a small step downhill, then look again. The step size decides everything.

    In the workbook · 1 exercise
  4. 04
    Automatic differentiation and the training loop

    Hand-derived gradients do not scale. Let PyTorch compute them, check it, then write the standard training loop.

    Read here · 1 exercise
  5. 05
    A tiny network, validation and overfitting

    A straight line cannot fit every pattern. Add a hidden layer, then learn to tell learning from memorising.

    In the workbook · 1 exercise
  6. 06
    Classic bugs, and what changes at frontier scale

    Most failed training runs fail quietly. Learn the usual suspects, then see what this toy leaves out.

    In the workbook · 1 exercise

Also inside: a 8-point checklist, a glossary of 12 terms and 10 questions and answers to test yourself. 6 hands-on exercises, each with a worked answer at the back where the workbook gives one.

No login, no card, no account. Before the download we ask you to follow Mickai (two quick links). Free to download and use for personal learning, study groups and inside your own team. Please do not resell the workbooks or republish them as your own. Link people to trust-agent.ai instead.

Test yourself

Questions and answers

Do I need a GPU for this workbook?

No. Every script runs on an ordinary CPU in seconds, using PyTorch's CPU-only build. The data is invented by seeded scripts, so there is nothing to download and no account to create.

What is the difference between a NumPy array and a PyTorch tensor?

Both are grids of numbers with a shape and a dtype, and both broadcast the same way. PyTorch adds automatic differentiation and support for accelerators. This workbook uses NumPy for the hand-worked gradient and PyTorch for autograd and the network, and torch.tensor(array) copies one into the other.

What exactly is a gradient?

For each parameter, it is the slope of the loss: how fast the loss rises if you nudge that parameter up. Together the slopes point uphill, so gradient descent moves each parameter a little in the opposite direction, then recomputes.

How do I choose a learning rate?

By experiment. Try a few values spread well apart and watch the loss. If it rises or swings wildly, the rate is too large; if it barely moves, too small. The best value depends on the data and the model, so it moves when either changes.

Why do I have to zero the gradients?

PyTorch adds each new gradient to whatever is already stored, so without clearing, every update uses the sum of old and new gradients. Call optimizer.zero_grad() once per update. The PyTorch documentation clears at the start of each step; clearing straight after the step also works.

What is the difference between a batch and an epoch?

A batch is the group of examples behind one parameter update. An epoch is one full pass through the training data. With 20 examples in batches of 5, one epoch makes 4 updates. Smaller batches give more, noisier updates per epoch.

Why does my validation loss rise while my training loss falls?

The model is starting to fit noise in the training data rather than the pattern: overfitting. Keep the weights from the epoch with the best validation loss, add data, simplify the model, and check several seeds before drawing conclusions from a small validation set.

Why do my numbers differ slightly from the workbook's?

The printed numbers came from Windows 11, and the NumPy and PyTorch documentation expect identical numbers only with the same versions on the same platform. Use the pinned versions and the same seeds, then compare the overall pattern and, as a rule of thumb, the first two significant figures, rather than every digit.

My model predicts almost the same value for every input. Why?

Check shapes first. If targets have shape (N,) and predictions (N, 1), broadcasting compares every prediction with every target, and under a squared-error loss the best a model can do is predict their mean. Also check for a model too simple for the data, such as a line on XOR, or a learning rate far too small.

Is this how large language models are trained?

The core loop is the same idea, but published reports show far more around it: huge curated datasets, hundreds or thousands of accelerators working in parallel, mixed precision, tuned optimisers and schedules, checkpoints and restarts, extensive evaluation and substantial cost. This workbook teaches the idea, not frontier pretraining.

When you have finished

Get your certificate of completion

Type your name and download a certificate for this workbook as a PDF, ready to print or to add to LinkedIn. It is made on your own device, so your name is never sent to us. It is a self-declared certificate, not an accredited qualification.

Learn the language

Key terms

Tensor
A grid of numbers with any number of dimensions. NumPy calls it an array.
Shape
The size of each dimension of a tensor, such as (20, 1) for 20 rows and 1 column.
Dtype
The kind of number a tensor stores, such as float32, float64 or int64. Gradients need floating-point dtypes.
Broadcasting
Automatically stretching tensors of different shapes so they can be combined. Shapes are compared from the right; sizes must be equal or 1.
Loss
A single number scoring how wrong a model's predictions are. Training tries to make it smaller.
Gradient
The slope of the loss with respect to each parameter. It points uphill, so training steps the opposite way.

6 of the workbook's 12 terms. The complete glossary is in the workbook.

Follow the evidence

Sources and checks

Facts last checked: .

Examples in this workbook were run on: Python 3.12.10 on Windows 11, in fresh virtual environments with numpy 2.5.3 and torch 2.14.0 (CPU build, reported as 2.14.0+cpu), installed by pip 25.0.1 (2026-09-26).

These workbooks use AI assistance. See how the workbooks are made.

  1. Start Locally (installation selector)PyTorch
  2. torch 2.14.0 release filesPython Package Index (PyPI)
  3. venv: creation of virtual environmentsPython Software Foundation
  4. Broadcasting (NumPy 2.5 user guide)NumPy
  5. Data types (NumPy 2.5 user guide)NumPy
  6. Random Generator (NumPy 2.5 reference)NumPy
  7. Random compatibility policy (NumPy 2.5 reference)NumPy
  8. numpy.polyfit (NumPy 2.5 reference)NumPy
  9. Broadcasting semantics (PyTorch 2.14)PyTorch
  10. Tensor attributes: dtype (PyTorch 2.14)PyTorch
  11. torch.tensor (PyTorch 2.14)PyTorch
  12. torch.set_default_dtype (PyTorch 2.14)PyTorch
  13. Autograd mechanics (PyTorch 2.14)PyTorch
  14. torch.Tensor.backward (PyTorch 2.14)PyTorch
  15. torch.optim (PyTorch 2.14)PyTorch
  16. torch.optim.Optimizer.zero_grad (PyTorch 2.14)PyTorch
  17. torch.no_grad (PyTorch 2.14)PyTorch
  18. MSELoss (PyTorch 2.14)PyTorch
  19. BCEWithLogitsLoss (PyTorch 2.14)PyTorch
  20. Optimizing model parameters (tutorial)PyTorch
  21. Reproducibility (PyTorch 2.14)PyTorch
  22. Deep Learning, chapter 4: Numerical computation (Goodfellow, Bengio and Courville)MIT Press, free online edition
  23. Deep Learning, chapter 5: Machine learning basics (Goodfellow, Bengio and Courville)MIT Press, free online edition
  24. Deep Learning, chapter 6: Deep feedforward networks (Goodfellow, Bengio and Courville)MIT Press, free online edition
  25. Scaling Laws for Neural Language ModelsarXiv, 2001.08361
  26. Training Compute-Optimal Large Language ModelsarXiv, 2203.15556
  27. OPT: Open Pre-trained Transformer Language ModelsarXiv, 2205.01068
  28. Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LMarXiv, 2104.04473
  29. Mixed Precision TrainingarXiv, 1710.03740

Created by Mickarle Wagstaff-Irons - Micky Irons with the Mickai team. Published by Mickai LTD. Last updated 26 September 2026.

NextKeep going

Where to go next