inference · Level 4

Inference engines explained

How a model actually runs on real hardware: memory, speed, batching, and how to serve it safely.

By Mickarle Wagstaff-Irons - Micky Irons

  • Level 4Frontier
  • 80 min
  • 7 chapters
  • Free PDF, no account
The Streaminference / 04

Start with the essentials

The short answer

An inference engine is the software that runs a trained model to produce output. It loads the weights into memory, turns your prompt into tokens, then generates the reply one token at a time. Whether a model runs, and how fast, is mostly a memory question: the weights plus the KV cache must fit, and quantisation shrinks them at some cost to quality.

What you will learn

  • You will be able to explain the difference between training and inference, and what is inside a model download.
  • You will be able to estimate the memory a model needs from its parameters, precision and context length.
  • You will understand quantisation well enough to weigh precision, memory and quality for your own use.
  • You will be able to explain prefill, decode, batching, paged attention and speculative decoding in plain terms.
  • You will know the trade-offs between local and hosted inference, and between GPU, CPU and unified memory.
  • You will be able to serve a model over a local HTTP API and lock it down sensibly.

Who it is for

Readers who already use AI tools or APIs and now want to know what happens when a model runs, so they can size hardware, choose an engine and serve models safely. There is some arithmetic, but nothing beyond multiplication.

Before you start

  • Comfort with tokens and context windows (see What is AI?).
  • The workbook GPU memory maths for LLMs and SLMs is the natural preparation: chapter 3 here recaps its memory sum only briefly.
  • A terminal is helpful for the exercises but not essential.

Keep learning

The complete workbook

This workbook goes inside the engine that runs a model. You will learn what is in a model download, how to work out the memory a model needs, why generation has two different speeds, and how batching, paged attention and speculative decoding make serving efficient. It ends with the trade-offs between local and hosted inference and a safe way to serve a model over a local HTTP API.

  1. 01
    Training, inference and the model file

    Training builds a model. Inference uses it. Almost everything a builder does with AI is inference, and it has an engineering craft of its own.

    In the workbook · 1 exercise
  2. 02
    Precision and quantisation

    Every weight is a stored number. How many bits you spend on each one is your biggest lever over memory.

    In the workbook · Reading
  3. 03
    Memory in brief: weights plus KV cache

    Whether a model runs comes down to one sum. Here it is in brief, from the engine's side.

    In the workbook · 1 exercise
  4. 04
    Prefill, decode and the two speeds

    When someone says a model is fast, ask: fast at what? Generation has two very different phases.

    In the workbook · 1 exercise
  5. 05
    Batching and three serving ideas

    One user's request leaves most of a GPU idle. Serving engines exist to keep it busy without wasting memory.

    In the workbook · Reading
  6. 06
    Hardware, engines, and local versus hosted

    The right hardware follows from the memory sum, not the other way round.

    In the workbook · Reading
  7. 07
    The local HTTP API, and running it safely

    Most engines expose the model as a small web server. That is convenient, and it is also an attack surface.

    In the workbook · 1 exercise

Also inside: a 9-point checklist, a glossary of 12 terms and 10 questions and answers to test yourself. 4 hands-on exercises, each with a worked answer at the back where the workbook gives one.

No login, no card, no account. Before the download we ask you to follow Mickai (two quick links). Free to download and use for personal learning, study groups and inside your own team. Please do not resell the workbooks or republish them as your own. Link people to trust-agent.ai instead.

Test yourself

Questions and answers

What is the difference between training and inference?

Training adjusts a model's weights so it predicts text well, and is done on large clusters. Inference runs the finished model to produce output. During inference the weights do not change, and the model remembers nothing between requests unless the earlier text is sent again.

What is quantisation, and does it ruin quality?

Quantisation stores weights with fewer bits, such as 8 or 4 instead of 16, so the model needs less memory. Lighter quantisation usually costs little quality and heavier costs more, unevenly across tasks. Test on your own work rather than trusting a rule.

Why does a serving engine claim most of the GPU's memory when it starts?

Some engines built for many users load the weights, then reserve much of the remaining memory as a KV cache pool before any request arrives, and hand out space from it as chats run. vLLM, for example, pre-allocates its cache using a configurable fraction of GPU memory. A nearly full card is therefore expected, not a leak. If other programs need room, lower that fraction, or the maximum context where the engine sizes its cache from it.

What happens when an engine runs out of KV cache memory?

It cannot keep every chat's cache in memory, so it must choose. Depending on the engine, new requests wait in a queue, or a running request is paused, its cache freed and later recomputed, which vLLM's documentation calls preemption. Some engines reject requests once their queue is full. So once the weights are loaded, the cache sets how many chats of a given length one machine can hold at once. Shorter contexts raise that number.

Why is generating a reply slower than reading the prompt?

Prompt tokens can be processed in parallel during prefill. Reply tokens are produced one at a time during decode, and each step must read the model's weights from memory. Decode is usually limited by memory bandwidth, not arithmetic.

What is continuous batching?

It is a scheduling method that re-forms the batch after every generation step. Finished requests leave immediately and waiting requests join, so the hardware stays busy. It raises total throughput compared with waiting for a whole batch to finish.

Does speculative decoding change the answers?

Not when implemented correctly. A fast draft proposes several tokens and the large model verifies them in one pass, accepting only those it would have produced. The output distribution is unchanged, so it affects speed rather than content.

Can I run a model without a GPU?

Yes, some engines run on the CPU alone, using system memory. It is slower, mainly because system memory moves data more slowly than graphics memory. Smaller or more heavily quantised models make CPU-only running more practical.

Is running a model locally more private?

It can be, because prompts need not leave your machine. But privacy then depends on your setup: who can reach the server, what it logs and where the weights came from. Local is not automatically safe.

Is it safe to expose a local model server to my network?

Not by default. Keep it bound to 127.0.0.1. If others need access, use an authenticating, encrypted reverse proxy, turn on the engine's API key, disable experimental tools, and decide deliberately what gets logged.

When you have finished

Get your certificate of completion

Type your name and download a certificate for this workbook as a PDF, ready to print or to add to LinkedIn. It is made on your own device, so your name is never sent to us. It is a self-declared certificate, not an accredited qualification.

Learn the language

Key terms

Inference
Running a trained model to produce output. The weights are fixed while it runs.
Weights (parameters)
The learned numbers inside a model. Their count and precision decide most of its memory needs.
Quantisation
Storing weights with fewer bits than the model was trained with, to save memory at some cost in quality.
KV cache
The keys and values stored for every token already in each conversation, so attention does not recompute them. An engine needs room for one per active chat.
Prefill
The phase in which the engine processes the whole prompt in parallel before the first output token.
Decode
The phase in which the engine writes the reply one token at a time, usually limited by memory bandwidth.

6 of the workbook's 12 terms. The complete glossary is in the workbook.

Follow the evidence

Sources and checks

Facts last checked: .

These workbooks use AI assistance. See how the workbooks are made.

  1. llama.cpp: LLM inference in C/C++ (README)ggml-org, GitHub
  2. llama.cpp server documentationggml-org, GitHub
  3. GGUF file format specificationggml-org, GitHub
  4. vLLM (README)vLLM project, GitHub
  5. vLLM serve command documentationvLLM project
  6. vLLM documentation: Optimization and Tuning (preemption and cache pre-allocation)vLLM project
  7. Ollama (README)Ollama, GitHub
  8. Ollama FAQ: network exposure and concurrent requestsOllama, GitHub
  9. Ollama API documentation: compatibility with a common chat-completions APIOllama
  10. safetensors (README)Hugging Face, GitHub
  11. Pickle scanning and securityHugging Face Hub documentation
  12. Efficient Memory Management for Large Language Model Serving with PagedAttentionarXiv, 2309.06180
  13. Orca: A Distributed Serving System for Transformer-Based Generative ModelsUSENIX OSDI 2022
  14. Fast Inference from Transformers via Speculative DecodingarXiv, 2211.17192
  15. Accelerating Large Language Model Decoding with Speculative SamplingarXiv, 2302.01318

Created by Mickarle Wagstaff-Irons - Micky Irons with the Mickai team. Published by Mickai LTD. Last updated 25 September 2026.

NextKeep going

Where to go next