hardware · Level 3

GPU memory maths for LLMs and SLMs

Work out whether a language model will fit and run, for inference, fine-tuning and training, before you buy or download anything.

By Mickarle Wagstaff-Irons - Micky Irons

  • Level 3Building
  • 75 min
  • 6 chapters
  • Free PDF, no account
The Enginehardware / 03

Start with the essentials

The short answer

A language model fits when its weights, its key-value cache and an overhead allowance add up to less than your device's memory, with a safety margin. Weights are parameters times bytes per parameter. Training needs far more: gradients, optimiser state and activations, starting at 16 bytes per parameter for mixed-precision Adam-style training. Lower precision, adapters and shorter context can bring a model within reach.

What you will learn

  • You will be able to work out the memory a model's weights need at any precision.
  • You will be able to size the key-value cache from a model card, and know when the standard formula does not apply.
  • You will be able to explain what the 16-bytes-per-parameter training figure includes and what it leaves out.
  • You will understand how activation checkpointing and adapters trade time or flexibility for memory.
  • You will be able to run a fit check with a safety margin and choose a fix when a model does not fit.
  • You will be able to say honestly what one GPU can fine-tune or train, and what it cannot.

Who it is for

Anyone comfortable with basic arithmetic who wants to know, before buying hardware or downloading a model, whether a language model will fit and run. You need multiplication and division, and a terminal for the calculator.

Before you start

  • Comfort with basic arithmetic (multiplying and dividing) and the ideas of a model's parameters and precision. The Which AI model workbook is helpful background, and Inference engines explained is the natural next step.

Read a sample · Chapter 01 of 06

01

Weights: parameters times bytes

Every fit check starts with the weights, the number you can compute exactly.

A model is a long list of numbers called parameters, each stored in a set number of bytes chosen by its precision. The rule never changes: weights in bytes = parameters x bytes per parameter. For a mixture-of-experts model (one split into many expert blocks, of which only a few run for each token), count every expert.

Here 1 GB is 1,000,000,000 bytes. Some tools report GiB (1,073,741,824 bytes, about 7.4% more), so check the unit. Card sizes here are examples only: read the vendor's specification page for yours.

Worked examples use hypothetical models invented for teaching. Model A has 7 billion parameters, 36 layers, 32 query heads, 4 key-value heads and head dimension 128.

Scroll sideways to see every column.

Weights: parameters times bytes · Table 1
PrecisionBytes per parameterModel A weights
32-bit float428.00 GB
16-bit (fp16 or bf16)214.00 GB
8-bit integer17.00 GB
4-bit0.53.50 GB

At 16-bit: 7 billion x 2 = 14 billion bytes = 14.00 GB. Real 4-bit files are a little larger than the simple sum, because each block of weights also stores a scale, so trust the file size on the download page.

Read a sample · Chapter 02 of 06

02

The key-value cache

While it generates, a model stores two vectors per layer for every earlier token. That store can rival the weights.

These are keys and values, so the store is the key-value cache (KV cache).

text · 2 lines
kv_per_token = 2 x layers x kv_heads x head_dim x bytes_per_value
kv_cache     = kv_per_token x tokens_in_context x simultaneous_chats

The 2 is for keys and values. Use key-value heads, not attention heads: grouped-query attention shares each key-value head across several query heads, and multi-query attention shares one across all. The paged-attention paper's example, 2 x 5,120 (hidden size) x 40 (layers) x 2 (bytes) per token, is this sum with full multi-head attention.

Worked example: Model A, 16-bit values, 16,384 tokens.

text · 2 lines
per token = 2 x 36 x 4 x 128 x 2 = 73,728 bytes
one chat  = 73,728 x 16,384 = 1,207,959,552 bytes = 1.21 GB

With full multi-head attention (32 key-value heads), each token costs 2 x 36 x 32 x 128 x 2 = 589,824 bytes, and 589,824 x 16,384 = 9.66 GB: eight times the cache, from one line of a model card.

Try it yourself · Activity 01

8 min

Read a cache off a model card

A hypothetical model has 24 layers, 8 key-value heads, head dimension 64 and 16-bit values. Size the cache for 2 chats of 8,192 tokens each.

  1. Work out the bytes per token, then the cache.
  2. Repeat with 24 key-value heads and note the ratio.

Which model-card line matters most here?

Worked answer

Per token: 2 x 24 x 8 x 64 x 2 = 49,152 bytes. Cache: 49,152 x 8,192 x 2 = 805,306,368 bytes = 0.81 GB. With 24 key-value heads: 147,456 bytes per token and 2.42 GB, three times as much (24 / 8 = 3).

Keep learning

The complete workbook

This workbook teaches one repeatable method for answering the question 'will it fit?'. You will work out weights at any precision, size the key-value cache from a model card, budget the extra memory that training and fine-tuning need, and run a small calculator with a safety margin. It ends with an honest account of what one GPU can and cannot train.

  1. 01
    Weights: parameters times bytes

    Every fit check starts with the weights, the number you can compute exactly.

    Read here · Reading
  2. 02
    The key-value cache

    While it generates, a model stores two vectors per layer for every earlier token. That store can rival the weights.

    Read here · 1 exercise
  3. 03
    Activations, overhead and the allowance

    Weights and cache are exact. The rest is not, so budget an allowance, then measure.

    In the workbook · 1 exercise
  4. 04
    Training memory: what learning adds

    Inference only reads the weights. Training changes them, and the model states alone need eight times the size of a 16-bit model file.

    In the workbook · 1 exercise
  5. 05
    Adapters: fine-tuning with far less state

    Full fine-tuning pays 16 bytes for every parameter. Adapters train a tiny fraction and skip most of that bill.

    In the workbook · 1 exercise
  6. 06
    The fit check, the fixes and the honest limits

    Turn the method into a repeatable procedure, then be straight about what one GPU can do.

    In the workbook · 1 exercise

Also inside: a 9-point checklist, a glossary of 12 terms and 10 questions and answers to test yourself. 5 hands-on exercises, each with a worked answer at the back where the workbook gives one.

No login, no card, no account. Before the download we ask you to follow Mickai (two quick links). Free to download and use for personal learning, study groups and inside your own team. Please do not resell the workbooks or republish them as your own. Link people to trust-agent.ai instead.

Test yourself

Questions and answers

How do I work out how much memory a model's weights need?

Multiply the parameter count by the bytes per parameter: 2 for 16-bit, 1 for 8-bit, 0.5 for 4-bit, plus the small extra real 4-bit formats add for scaling values. A 7-billion-parameter model is 14 GB at 16-bit, 7 GB at 8-bit and about 3.5 GB at 4-bit. Check whether your tool reports GB or GiB.

What is the KV cache and what sets its size?

It stores the keys and values of every token already in the conversation, so the model need not recompute them. Per token it is 2 x layers x key-value heads x head dimension x bytes per value, multiplied by the tokens held and the number of simultaneous chats. Read the key-value head count from the model card, not the attention head count.

Does the KV cache formula always apply?

No. It assumes every layer keeps every past token in full. Sliding-window attention looks only at a recent window, so older tokens may not need keeping, and compressed designs such as latent attention store a smaller vector per token, so the formula can overstate them. Some models mix layer types. Use the publisher's stated cache size where given, and measure your own setup.

What does 16 bytes per parameter for training mean?

It is the ZeRO paper's count of model states for mixed-precision training with an Adam-style optimiser: 2 bytes for 16-bit weights, 2 for 16-bit gradients, and 12 for the 32-bit master weights, momentum and variance. It counts the parameter-related tensors only.

Does 16 bytes per parameter include activations?

No. The paper counts activations, temporary buffers and fragmented memory separately, as residual states. Treat 16 bytes per parameter as a floor. Activations grow with batch size, sequence length and depth, so estimate them, or measure a small run, and add them.

What is activation checkpointing?

It keeps only some activations during the forward pass and recomputes the rest when the backward pass needs them. The original paper's method needs memory that grows with the square root of the layer count, for one extra forward pass, roughly a third more compute per step. It does not shrink weights, gradients or optimiser state.

Why does LoRA need less memory than full fine-tuning?

The pretrained weights are frozen, so they need no gradients and no optimiser state. Only the small adapter matrices are trained, and only they pay the full per-parameter training cost. Most activations are still needed, so checkpointing and smaller batches still matter.

How much safety margin should I leave?

There is no published constant. This workbook plans to use at most 90% of the device's memory, leaving 10% for other programs and errors, plus an allowance of the larger of 1 GB or 10% of everything else counted. Both are planning assumptions. Measure your real peak memory and adjust.

What should I do when a model does not fit?

Change the biggest term first. Lower the precision, choose a smaller model, shorten the context or serve fewer chats, offload some data to system memory at a cost in speed, or split the model across more devices. For training, use a smaller batch, activation checkpointing or adapters.

Can one GPU pretrain a frontier-scale model?

No. Memory is only one constraint. Model states alone need 16 bytes per parameter, compute time runs to years on one device at that scale by this workbook's estimates, and the data and engineering needs are large too. A single GPU can run models that fit its memory, fine-tune small ones with adapters and train toy models.

When you have finished

Get your certificate of completion

Type your name and download a certificate for this workbook as a PDF, ready to print or to add to LinkedIn. It is made on your own device, so your name is never sent to us. It is a self-declared certificate, not an accredited qualification.

Learn the language

Key terms

Parameter
One learned number inside a model. The parameter count times the bytes per parameter gives the size of the weights.
Precision
How many bits store each number, for example 32, 16, 8 or 4. Fewer bits mean less memory and usually some loss of quality.
KV cache
Stored keys and values for every token already in the conversation, so attention does not recompute them. It grows with context and with simultaneous chats.
Key-value heads
The number of separate key and value sets in each layer. Grouped-query and multi-query attention use fewer than the number of query heads, which shrinks the cache.
Activations
The temporary tensors each layer produces. In training they are kept for the backward pass and can dominate memory.
Gradient
The calculated slope showing how each trainable parameter should change to reduce the error. Training stores one per trainable parameter.

6 of the workbook's 12 terms. The complete glossary is in the workbook.

Follow the evidence

Sources and checks

Facts last checked: .

Examples in this workbook were run on: Python 3.12.10 on Windows 11 (2026-09-25).

These workbooks use AI assistance. See how the workbooks are made.

  1. ZeRO: Memory Optimizations Toward Training Trillion Parameter ModelsarXiv, 1910.02054
  2. LoRA: Low-Rank Adaptation of Large Language ModelsarXiv, 2106.09685
  3. QLoRA: Efficient Finetuning of Quantized LLMsarXiv, 2305.14314
  4. Training Deep Nets with Sublinear Memory CostarXiv, 1604.06174
  5. Mixed Precision TrainingarXiv, 1710.03740
  6. Reducing Activation Recomputation in Large Transformer ModelsarXiv, 2205.05198
  7. FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessarXiv, 2205.14135
  8. Efficient Memory Management for Large Language Model Serving with PagedAttentionarXiv, 2309.06180
  9. GQA: Training Generalized Multi-Query Transformer Models from Multi-Head CheckpointsarXiv, 2305.13245
  10. Fast Transformer Decoding: One Write-Head is All You NeedarXiv, 1911.02150
  11. Longformer: The Long-Document TransformerarXiv, 2004.05150
  12. DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language ModelarXiv, 2405.04434
  13. Scaling Laws for Neural Language ModelsarXiv, 2001.08361
  14. Training Compute-Optimal Large Language ModelsarXiv, 2203.15556

Created by Mickarle Wagstaff-Irons - Micky Irons with the Mickai team. Published by Mickai LTD. Last updated 25 September 2026.

NextKeep going

Where to go next