models · Level 3
Tensors, gradients and a first training loop
Train tiny models on a CPU in seconds, by hand and then with automatic differentiation, and see what training is and is not.

Start with the essentials
The short answer
Training repeats one loop: run the model on examples, score how wrong it is with a single number called the loss, compute the gradient (which way each parameter should move to lower that loss), and take a small step that way. Tiny models train on an ordinary CPU in seconds. Published frontier-scale language-model runs use the same idea with vastly more data, hardware, engineering and cost.
What you will learn
- You will be able to read a tensor's shape and dtype and predict when broadcasting will stretch it, including the cases that fail silently.
- You will be able to compute a mean squared error loss and its gradient by hand, and check it by two independent routes.
- You will be able to spot a learning rate that is too small or too large from the loss alone.
- You will be able to write the PyTorch training loop with batches and epochs, and train a tiny network on non-linear data.
- You will be able to separate training from validation, recognise overfitting and find the classic training bugs.
- You will be able to explain, with sources, what changes when language models are pretrained at frontier scale.
Who it is for
People comfortable writing small Python programs who want to understand what training a model actually means by doing it. You need a computer that can run Python 3.12 (the version tested here) and install two packages; on a Mac, PyTorch 2.14.0 needs Apple silicon and macOS 14 or later. No GPU, no account and no dataset download.
Before you start
- Comfort with Python basics: variables, loops, functions and running a script from a terminal, as taught in Python: from first script to a useful automation. School algebra is enough; the one idea from calculus you need, a slope, is explained as you go.
Read a sample · Chapter 02 of 06
The loss and the gradient, by hand
Training needs one number that says how wrong the model is, and a way to tell which direction makes it smaller.
toy.py invents 20 points near y = 3x + 2. Its seed fixes the random numbers, so every run makes the same data. The model is a line with two parameters, weight w and bias b: prediction = w * x + b.
# toy.py: data, loss and gradient
import numpy as np
rng = np.random.default_rng(seed=0)
x = rng.uniform(-1, 1, size=20)
y = 3 * x + 2 + rng.normal(0, 0.1, size=20) # a line plus noise
def loss(w, b):
return np.mean((w * x + b - y) ** 2)
def gradient(w, b):
err = w * x + b - y
return np.mean(2 * err * x), np.mean(2 * err)The loss scores 'how wrong' as one number. Here it is the mean squared error: the average of (prediction minus target) squared, so big misses count most.
The gradient holds the loss's slope for each parameter. It points uphill, says the Deep Learning textbook, so its negative points downhill. Differentiating gives the slopes in gradient: the mean of 2 * err * x for w, and of 2 * err for b.
Always check a hand-derived gradient, for instance by nudging each parameter slightly both ways and measuring the change in loss: a finite difference.
# check.py
from toy import loss, gradient
w, b, h = 0.5, -1.0, 1e-6 # h is a tiny nudge
print("by hand %.6f %.6f" % gradient(w, b))
print("nudged %.6f %.6f" % ((loss(w + h, b) - loss(w - h, b)) / (2 * h),
(loss(w, b + h) - loss(w, b - h)) / (2 * h)))by hand -2.194225 -6.135977
nudged -2.194225 -6.135977They agree to six decimals. Both slopes are negative, so raising w and b lowers the loss: the true 3 and 2 are above 0.5 and -1.
Try it yourself · Activity 02
15 minOne step by hand
With pencil and paper, take the points x = 1, 2, 3 and y = 5, 8, 11 (exactly on y = 3x + 2), starting at w = 0, b = 0.
- Work out the errors (prediction minus target) and the loss.
- Work out the slopes for
wandb. - Take one step with learning rate 0.1 (subtract 0.1 times each slope), find the new loss, and check it in Python.
Why did one step help so much? Would it always?
Worked answer
Errors -5, -8, -11; loss (25 + 64 + 121) / 3 = 70. Slope for w: 2 x (-5 - 16 - 33) / 3 = -36. Slope for b: 2 x (-24) / 3 = -16. New w = 3.6, b = 1.6; new errors 0.2, 0.8, 1.4; loss (0.04 + 0.64 + 1.96) / 3 = 0.88. With x = np.array([1, 2, 3]) and y = np.array([5, 8, 11]), print(np.mean((3.6 * x + 1.6 - y) ** 2)) prints 0.8800000000000008, a floating-point rounding artefact.
Read a sample · Chapter 04 of 06
Automatic differentiation and the training loop
Hand-derived gradients do not scale. Let PyTorch compute them, check it, then write the standard training loop.
Automatic differentiation (autograd) records the forward calculation as a graph, then applies the chain rule back through it; that backward sweep is backpropagation. Mark a tensor requires_grad=True, call loss.backward(), and its slope appears in .grad. This uses check.py's point, in float64 to match NumPy.
# autograd_check.py
import torch
from toy import x, y, gradient
xt, yt = torch.tensor(x), torch.tensor(y)
w = torch.tensor(0.5, dtype=torch.float64, requires_grad=True)
b = torch.tensor(-1.0, dtype=torch.float64, requires_grad=True)
print("by hand %.6f %.6f" % gradient(0.5, -1.0))
for run in ["autograd ", "not zeroed"]:
loss = ((w * xt + b - yt) ** 2).mean() # forward and loss
loss.backward() # backward
print(run, "%.6f %.6f" % (w.grad.item(), b.grad.item()))by hand -2.194225 -6.135977
autograd -2.194225 -6.135977
not zeroed -4.388451 -12.271953Three routes agree; in full, autograd and hand differ by under 1e-15, float64 rounding. The last line is a classic bug: the backward documentation says gradients accumulate, so a second pass without clearing doubled them.
Now the real loop, in float32, PyTorch's default. Linear(1, 1) is the same line, MSELoss the same loss, SGD the same update. A batch is the examples behind one update; an epoch is one pass through the data, here 4 batches of 5.
# train_line.py
import torch
from toy import x, y
torch.manual_seed(0)
X = torch.tensor(x, dtype=torch.float32).reshape(20, 1)
Y = torch.tensor(y, dtype=torch.float32).reshape(20, 1)
model = torch.nn.Linear(1, 1)
loss_fn = torch.nn.MSELoss()
opt = torch.optim.SGD(model.parameters(), lr=0.1)
for epoch in range(1, 31):
order = torch.randperm(20) # shuffle each epoch
for start in range(0, 20, 5): # batches of 5
batch = order[start:start + 5]
opt.zero_grad() # 1 clear
loss = loss_fn(model(X[batch]), Y[batch]) # 2 forward, 3 loss
loss.backward() # 4 backward
opt.step() # 5 step
if epoch % 10 == 0:
with torch.no_grad():
print(epoch, round(loss_fn(model(X), Y).item(), 5))
print(round(model.weight.item(), 3), round(model.bias.item(), 3))10 0.00759
20 0.00452
30 0.0045
2.971 2.007The numbered steps are the training loop; PyTorch's documentation clears gradients first, though clearing after the step works too. torch.no_grad() stops tracking while you only measure, and .reshape(20, 1) gives targets the predictions' shape.
Try it yourself · Activity 04
10 minChange the batch size
In train_line.py, change the two 5s on the batch lines to 1, run it, then try 20.
- First work out how many updates each version makes in 30 epochs.
- Compare the losses at epochs 10, 20 and 30.
Which would you choose if each update were expensive?
Worked answer
Batches of 1: 600 updates; losses 0.0046, 0.00557, 0.00468. Fast but jittery: each step sees one noisy point. Batches of 20: 30 updates; losses 0.65896, 0.11635, 0.02439, ending 2.751 2.014. Smooth, but unfinished.
Keep learning
The complete workbook
This workbook teaches training by doing it at toy scale. You will compute a gradient by hand in NumPy, confirm it with PyTorch's automatic differentiation, then fit a straight line and a tiny network while watching the learning rate, validation loss and common bugs. It ends with a sourced account of what changes when language models are pretrained at frontier scale.
- 01Tensors: shape, dtype and broadcastingIn the workbook · 1 exercise
Everything a model reads, stores and learns is a grid of numbers. Learn to read one before trusting any result.
- 02The loss and the gradient, by handRead here · 1 exercise
Training needs one number that says how wrong the model is, and a way to tell which direction makes it smaller.
- 03Gradient descent and the learning rateIn the workbook · 1 exercise
Gradient descent repeats one move: a small step downhill, then look again. The step size decides everything.
- 04Automatic differentiation and the training loopRead here · 1 exercise
Hand-derived gradients do not scale. Let PyTorch compute them, check it, then write the standard training loop.
- 05A tiny network, validation and overfittingIn the workbook · 1 exercise
A straight line cannot fit every pattern. Add a hidden layer, then learn to tell learning from memorising.
- 06Classic bugs, and what changes at frontier scaleIn the workbook · 1 exercise
Most failed training runs fail quietly. Learn the usual suspects, then see what this toy leaves out.
Also inside: a 8-point checklist, a glossary of 12 terms and 10 questions and answers to test yourself. 6 hands-on exercises, each with a worked answer at the back where the workbook gives one.
No login, no card, no account. Before the download we ask you to follow Mickai (two quick links). Free to download and use for personal learning, study groups and inside your own team. Please do not resell the workbooks or republish them as your own. Link people to trust-agent.ai instead.
Test yourself
Questions and answers
Do I need a GPU for this workbook?
No. Every script runs on an ordinary CPU in seconds, using PyTorch's CPU-only build. The data is invented by seeded scripts, so there is nothing to download and no account to create.
What is the difference between a NumPy array and a PyTorch tensor?
Both are grids of numbers with a shape and a dtype, and both broadcast the same way. PyTorch adds automatic differentiation and support for accelerators. This workbook uses NumPy for the hand-worked gradient and PyTorch for autograd and the network, and torch.tensor(array) copies one into the other.
What exactly is a gradient?
For each parameter, it is the slope of the loss: how fast the loss rises if you nudge that parameter up. Together the slopes point uphill, so gradient descent moves each parameter a little in the opposite direction, then recomputes.
How do I choose a learning rate?
By experiment. Try a few values spread well apart and watch the loss. If it rises or swings wildly, the rate is too large; if it barely moves, too small. The best value depends on the data and the model, so it moves when either changes.
Why do I have to zero the gradients?
PyTorch adds each new gradient to whatever is already stored, so without clearing, every update uses the sum of old and new gradients. Call optimizer.zero_grad() once per update. The PyTorch documentation clears at the start of each step; clearing straight after the step also works.
What is the difference between a batch and an epoch?
A batch is the group of examples behind one parameter update. An epoch is one full pass through the training data. With 20 examples in batches of 5, one epoch makes 4 updates. Smaller batches give more, noisier updates per epoch.
Why does my validation loss rise while my training loss falls?
The model is starting to fit noise in the training data rather than the pattern: overfitting. Keep the weights from the epoch with the best validation loss, add data, simplify the model, and check several seeds before drawing conclusions from a small validation set.
Why do my numbers differ slightly from the workbook's?
The printed numbers came from Windows 11, and the NumPy and PyTorch documentation expect identical numbers only with the same versions on the same platform. Use the pinned versions and the same seeds, then compare the overall pattern and, as a rule of thumb, the first two significant figures, rather than every digit.
My model predicts almost the same value for every input. Why?
Check shapes first. If targets have shape (N,) and predictions (N, 1), broadcasting compares every prediction with every target, and under a squared-error loss the best a model can do is predict their mean. Also check for a model too simple for the data, such as a line on XOR, or a learning rate far too small.
Is this how large language models are trained?
The core loop is the same idea, but published reports show far more around it: huge curated datasets, hundreds or thousands of accelerators working in parallel, mixed precision, tuned optimisers and schedules, checkpoints and restarts, extensive evaluation and substantial cost. This workbook teaches the idea, not frontier pretraining.
When you have finished
Get your certificate of completion
Type your name and download a certificate for this workbook as a PDF, ready to print or to add to LinkedIn. It is made on your own device, so your name is never sent to us. It is a self-declared certificate, not an accredited qualification.
Learn the language
Key terms
- Tensor
- A grid of numbers with any number of dimensions. NumPy calls it an array.
- Shape
- The size of each dimension of a tensor, such as (20, 1) for 20 rows and 1 column.
- Dtype
- The kind of number a tensor stores, such as float32, float64 or int64. Gradients need floating-point dtypes.
- Broadcasting
- Automatically stretching tensors of different shapes so they can be combined. Shapes are compared from the right; sizes must be equal or 1.
- Loss
- A single number scoring how wrong a model's predictions are. Training tries to make it smaller.
- Gradient
- The slope of the loss with respect to each parameter. It points uphill, so training steps the opposite way.
6 of the workbook's 12 terms. The complete glossary is in the workbook.
Follow the evidence
Sources and checks
Facts last checked: .
Examples in this workbook were run on: Python 3.12.10 on Windows 11, in fresh virtual environments with numpy 2.5.3 and torch 2.14.0 (CPU build, reported as 2.14.0+cpu), installed by pip 25.0.1 (2026-09-26).
These workbooks use AI assistance. See how the workbooks are made.
- Start Locally (installation selector)PyTorch
- torch 2.14.0 release filesPython Package Index (PyPI)
- venv: creation of virtual environmentsPython Software Foundation
- Broadcasting (NumPy 2.5 user guide)NumPy
- Data types (NumPy 2.5 user guide)NumPy
- Random Generator (NumPy 2.5 reference)NumPy
- Random compatibility policy (NumPy 2.5 reference)NumPy
- numpy.polyfit (NumPy 2.5 reference)NumPy
- Broadcasting semantics (PyTorch 2.14)PyTorch
- Tensor attributes: dtype (PyTorch 2.14)PyTorch
- torch.tensor (PyTorch 2.14)PyTorch
- torch.set_default_dtype (PyTorch 2.14)PyTorch
- Autograd mechanics (PyTorch 2.14)PyTorch
- torch.Tensor.backward (PyTorch 2.14)PyTorch
- torch.optim (PyTorch 2.14)PyTorch
- torch.optim.Optimizer.zero_grad (PyTorch 2.14)PyTorch
- torch.no_grad (PyTorch 2.14)PyTorch
- MSELoss (PyTorch 2.14)PyTorch
- BCEWithLogitsLoss (PyTorch 2.14)PyTorch
- Optimizing model parameters (tutorial)PyTorch
- Reproducibility (PyTorch 2.14)PyTorch
- Deep Learning, chapter 4: Numerical computation (Goodfellow, Bengio and Courville)MIT Press, free online edition
- Deep Learning, chapter 5: Machine learning basics (Goodfellow, Bengio and Courville)MIT Press, free online edition
- Deep Learning, chapter 6: Deep feedforward networks (Goodfellow, Bengio and Courville)MIT Press, free online edition
- Scaling Laws for Neural Language ModelsarXiv, 2001.08361
- Training Compute-Optimal Large Language ModelsarXiv, 2203.15556
- OPT: Open Pre-trained Transformer Language ModelsarXiv, 2205.01068
- Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LMarXiv, 2104.04473
- Mixed Precision TrainingarXiv, 1710.03740
Created by Mickarle Wagstaff-Irons - Micky Irons with the Mickai team. Published by Mickai LTD. Last updated 26 September 2026.
NextKeep going
Where to go next
Recommended for you
GPU memory maths for LLMs and SLMs
A repeatable method for checking whether a language model fits your GPU: weights, KV cache, training memory and adapters, plus a small calculator you can run.
Recommended for you
How to evaluate a new frontier model the day it lands
A ten-step method for judging a new model: primary sources, licence, architecture, hardware, benchmarks, your own tests, safety, jurisdiction and a decision record.
Recommended for you
Inference engines explained
How inference engines run a model: memory maths, quantisation, prefill and decode, batching, hardware, and serving a model safely over a local API.