inference · Level 4
Inference engines explained
How a model actually runs on real hardware: memory, speed, batching, and how to serve it safely.

Start with the essentials
The short answer
An inference engine is the software that runs a trained model to produce output. It loads the weights into memory, turns your prompt into tokens, then generates the reply one token at a time. Whether a model runs, and how fast, is mostly a memory question: the weights plus the KV cache must fit, and quantisation shrinks them at some cost to quality.
What you will learn
- You will be able to explain the difference between training and inference, and what is inside a model download.
- You will be able to estimate the memory a model needs from its parameters, precision and context length.
- You will understand quantisation well enough to weigh precision, memory and quality for your own use.
- You will be able to explain prefill, decode, batching, paged attention and speculative decoding in plain terms.
- You will know the trade-offs between local and hosted inference, and between GPU, CPU and unified memory.
- You will be able to serve a model over a local HTTP API and lock it down sensibly.
Who it is for
Readers who already use AI tools or APIs and now want to know what happens when a model runs, so they can size hardware, choose an engine and serve models safely. There is some arithmetic, but nothing beyond multiplication.
Before you start
- Comfort with tokens and context windows (see What is AI?).
- The workbook GPU memory maths for LLMs and SLMs is the natural preparation: chapter 3 here recaps its memory sum only briefly.
- A terminal is helpful for the exercises but not essential.
Keep learning
The complete workbook
This workbook goes inside the engine that runs a model. You will learn what is in a model download, how to work out the memory a model needs, why generation has two different speeds, and how batching, paged attention and speculative decoding make serving efficient. It ends with the trade-offs between local and hosted inference and a safe way to serve a model over a local HTTP API.
- 01Training, inference and the model fileIn the workbook · 1 exercise
Training builds a model. Inference uses it. Almost everything a builder does with AI is inference, and it has an engineering craft of its own.
- 02Precision and quantisationIn the workbook · Reading
Every weight is a stored number. How many bits you spend on each one is your biggest lever over memory.
- 03Memory in brief: weights plus KV cacheIn the workbook · 1 exercise
Whether a model runs comes down to one sum. Here it is in brief, from the engine's side.
- 04Prefill, decode and the two speedsIn the workbook · 1 exercise
When someone says a model is fast, ask: fast at what? Generation has two very different phases.
- 05Batching and three serving ideasIn the workbook · Reading
One user's request leaves most of a GPU idle. Serving engines exist to keep it busy without wasting memory.
- 06Hardware, engines, and local versus hostedIn the workbook · Reading
The right hardware follows from the memory sum, not the other way round.
- 07The local HTTP API, and running it safelyIn the workbook · 1 exercise
Most engines expose the model as a small web server. That is convenient, and it is also an attack surface.
Also inside: a 9-point checklist, a glossary of 12 terms and 10 questions and answers to test yourself. 4 hands-on exercises, each with a worked answer at the back where the workbook gives one.
No login, no card, no account. Before the download we ask you to follow Mickai (two quick links). Free to download and use for personal learning, study groups and inside your own team. Please do not resell the workbooks or republish them as your own. Link people to trust-agent.ai instead.
Test yourself
Questions and answers
What is the difference between training and inference?
Training adjusts a model's weights so it predicts text well, and is done on large clusters. Inference runs the finished model to produce output. During inference the weights do not change, and the model remembers nothing between requests unless the earlier text is sent again.
What is quantisation, and does it ruin quality?
Quantisation stores weights with fewer bits, such as 8 or 4 instead of 16, so the model needs less memory. Lighter quantisation usually costs little quality and heavier costs more, unevenly across tasks. Test on your own work rather than trusting a rule.
Why does a serving engine claim most of the GPU's memory when it starts?
Some engines built for many users load the weights, then reserve much of the remaining memory as a KV cache pool before any request arrives, and hand out space from it as chats run. vLLM, for example, pre-allocates its cache using a configurable fraction of GPU memory. A nearly full card is therefore expected, not a leak. If other programs need room, lower that fraction, or the maximum context where the engine sizes its cache from it.
What happens when an engine runs out of KV cache memory?
It cannot keep every chat's cache in memory, so it must choose. Depending on the engine, new requests wait in a queue, or a running request is paused, its cache freed and later recomputed, which vLLM's documentation calls preemption. Some engines reject requests once their queue is full. So once the weights are loaded, the cache sets how many chats of a given length one machine can hold at once. Shorter contexts raise that number.
Why is generating a reply slower than reading the prompt?
Prompt tokens can be processed in parallel during prefill. Reply tokens are produced one at a time during decode, and each step must read the model's weights from memory. Decode is usually limited by memory bandwidth, not arithmetic.
What is continuous batching?
It is a scheduling method that re-forms the batch after every generation step. Finished requests leave immediately and waiting requests join, so the hardware stays busy. It raises total throughput compared with waiting for a whole batch to finish.
Does speculative decoding change the answers?
Not when implemented correctly. A fast draft proposes several tokens and the large model verifies them in one pass, accepting only those it would have produced. The output distribution is unchanged, so it affects speed rather than content.
Can I run a model without a GPU?
Yes, some engines run on the CPU alone, using system memory. It is slower, mainly because system memory moves data more slowly than graphics memory. Smaller or more heavily quantised models make CPU-only running more practical.
Is running a model locally more private?
It can be, because prompts need not leave your machine. But privacy then depends on your setup: who can reach the server, what it logs and where the weights came from. Local is not automatically safe.
Is it safe to expose a local model server to my network?
Not by default. Keep it bound to 127.0.0.1. If others need access, use an authenticating, encrypted reverse proxy, turn on the engine's API key, disable experimental tools, and decide deliberately what gets logged.
When you have finished
Get your certificate of completion
Type your name and download a certificate for this workbook as a PDF, ready to print or to add to LinkedIn. It is made on your own device, so your name is never sent to us. It is a self-declared certificate, not an accredited qualification.
Learn the language
Key terms
- Inference
- Running a trained model to produce output. The weights are fixed while it runs.
- Weights (parameters)
- The learned numbers inside a model. Their count and precision decide most of its memory needs.
- Quantisation
- Storing weights with fewer bits than the model was trained with, to save memory at some cost in quality.
- KV cache
- The keys and values stored for every token already in each conversation, so attention does not recompute them. An engine needs room for one per active chat.
- Prefill
- The phase in which the engine processes the whole prompt in parallel before the first output token.
- Decode
- The phase in which the engine writes the reply one token at a time, usually limited by memory bandwidth.
6 of the workbook's 12 terms. The complete glossary is in the workbook.
Follow the evidence
Sources and checks
Facts last checked: .
These workbooks use AI assistance. See how the workbooks are made.
- llama.cpp: LLM inference in C/C++ (README)ggml-org, GitHub
- llama.cpp server documentationggml-org, GitHub
- GGUF file format specificationggml-org, GitHub
- vLLM (README)vLLM project, GitHub
- vLLM serve command documentationvLLM project
- vLLM documentation: Optimization and Tuning (preemption and cache pre-allocation)vLLM project
- Ollama (README)Ollama, GitHub
- Ollama FAQ: network exposure and concurrent requestsOllama, GitHub
- Ollama API documentation: compatibility with a common chat-completions APIOllama
- safetensors (README)Hugging Face, GitHub
- Pickle scanning and securityHugging Face Hub documentation
- Efficient Memory Management for Large Language Model Serving with PagedAttentionarXiv, 2309.06180
- Orca: A Distributed Serving System for Transformer-Based Generative ModelsUSENIX OSDI 2022
- Fast Inference from Transformers via Speculative DecodingarXiv, 2211.17192
- Accelerating Large Language Model Decoding with Speculative SamplingarXiv, 2302.01318
Created by Mickarle Wagstaff-Irons - Micky Irons with the Mickai team. Published by Mickai LTD. Last updated 25 September 2026.
NextKeep going
Where to go next
Recommended for you
Reasoning models and reasoning engines
What reasoning models are, when extra thinking pays off, how to control it, why a reasoning trace is not proof, and how classical reasoning engines differ.
Recommended for you
How to evaluate a new frontier model the day it lands
A ten-step method for judging a new model: primary sources, licence, architecture, hardware, benchmarks, your own tests, safety, jurisdiction and a decision record.
Recommended for you
Host a chatbot on your website, safely
Put a chatbot on your site without leaking keys or running up bills: server function, limits, CORS, privacy, accessibility, fallbacks, monitoring and a launch checklist.