hardware · Level 2
Running open-weight AI models on your own computer, for free
A hands-on guide to Ollama, LM Studio and llama.cpp: what your RAM or VRAM can realistically run, how quantisation trades size for quality, and where to find real models on Hugging Face.

Start with the essentials
The short answer
You can run real open-weight AI models on an ordinary laptop for no ongoing cost, using free tools such as Ollama, LM Studio or llama.cpp. What you can run depends mostly on memory: a model's weights need roughly its parameter count times the bytes each parameter is stored in, and quantisation lowers that at some cost to quality. Start small, quantised and local, then grow from there.
What you will learn
- You will be able to explain what 'open-weight' and 'free' actually mean in this context, and how they differ from a hosted AI subscription.
- You will be able to install and use at least one of Ollama, LM Studio or llama.cpp to run a model locally.
- You will be able to estimate, roughly, whether a given model will fit your machine's memory, and know why a mixture-of-experts model's memory need does not match its speed.
- You will be able to search Hugging Face for an open-weight model, read enough of its model card to judge its licence and size, and choose a sensible quantisation.
- You will have run at least one model end to end on your own machine, and know the first three things to check when it is slow or will not load.
Who it is for
Anyone comfortable installing ordinary desktop software who wants to run an AI model on their own computer, for privacy, for offline use, to avoid a subscription, or simply to learn. No programming is required for the main path; a short section for the technically curious uses the command line.
Before you start
- None beyond comfort installing and running desktop applications. Which AI model should I use? is helpful background on open-weight models, and GPU memory maths and Inference engines explained are natural next steps once you want the exact formulas.
Keep learning
The complete workbook
This workbook is a practical getting-started guide, not a marketing pitch. You will install and compare Ollama, LM Studio and llama.cpp, learn a rough but honest way to judge whether a model will fit your machine's RAM or VRAM, understand what quantisation actually trades away, and use Hugging Face to find and check a real open-weight model. It ends with a full first run, a short troubleshooting table, and an honest account of what local models are, and are not, good for.
- 01Why run a model on your own computerIn the workbook · 1 exercise
Before you install anything, it helps to know what you are choosing, and what 'free' does and does not mean here.
- 02The three tools you will meet: Ollama, LM Studio and llama.cppIn the workbook · 1 exercise
Most local-model tools are a friendly front end over a much smaller number of inference engines. Get to know the engine and the front ends make more sense.
- 03Will it fit? Size, RAM, VRAM and quantisation without the full mathsIn the workbook · 1 exercise
You do not need the exact formulas to get started. You need one honest rule of thumb, and to know what quantisation is actually trading away.
- 04Finding a real model on Hugging FaceIn the workbook · 1 exercise
Hugging Face is the model hosting site both Ollama and LM Studio point to, directly or indirectly, and the most direct way to see what is actually available today.
- 05Running your first model, safely, and what comes nextIn the workbook · 1 exercise
Time to put it together: download, load, prompt, and know what to check when something goes wrong.
Also inside: a 9-point checklist, a glossary of 12 terms and 10 questions and answers to test yourself. 5 hands-on exercises, each with a worked answer at the back where the workbook gives one.
No login, no card, no account. Before the download we ask you to follow Mickai (two quick links). Free to download and use for personal learning, study groups and inside your own team. Please do not resell the workbooks or republish them as your own. Link people to trust-agent.ai instead.
Test yourself
Questions and answers
Is running a model locally actually free?
There is no subscription or per-message charge, so day to day it costs nothing extra beyond electricity. You do spend hardware, disk space and setup time, and if you need to buy a machine to do it, that is a real cost. For a laptop you already own, the marginal cost of trying it is small.
What is the difference between open-weight and open source?
Open-weight means the trained parameters are published for download under some licence. Open source, in the Open Source Initiative's stricter definition, also requires the training code and real information about the training data, plus freedom to use the model for any purpose. Most well-known 'open' models are open-weight, not fully open source by that definition.
Do I need Ollama, LM Studio or llama.cpp, or all three?
Just one to get started. LM Studio suits people who prefer a graphical browser and downloader. Ollama suits people comfortable with a short list of terminal commands, or who want its local API. llama.cpp directly suits people who want the most control or an unusual setup. All three can use the same GGUF files, so trying a second one later mostly changes the interface, not the underlying model.
How do I know if a model will fit my computer?
As a rough rule, a model's weights need about its parameter count times the bytes each parameter is stored in: roughly 2 bytes at 16-bit, 1 at 8-bit, and half a byte at 4-bit, plus a little more in practice. Compare that against your RAM, or your GPU's VRAM, leaving headroom for your operating system and the context you plan to use. GPU memory maths for LLMs and SLMs gives the exact method.
What does quantisation actually cost me?
Some quality, unevenly across tasks, in exchange for a smaller file and often more speed. Q4_K_M or Q5_K_M is a reasonable starting point for most people. There is no single published number for how much quality you lose at any given level; the only reliable test is comparing outputs on your own real prompts.
Why can a 30-billion-parameter model run at a reasonable speed on a laptop?
If it is a mixture-of-experts model, such as one labelled '30B-A3B', only a fraction of its parameters, here about 3 billion, actually run for each word generated, even though all 30 billion must be held in memory. Memory need follows the total; speed follows the active count.
Where should I actually look for open-weight models?
Hugging Face, filtered by task, Text Generation for a chat model, and by format, GGUF for llama.cpp, Ollama or LM Studio. Read the model card for its licence, base model and any stated limitations before downloading, and prefer safetensors or GGUF files over older pickle-based formats.
Are these models safe to use commercially?
It depends entirely on the individual licence, which you should read on the model's own page. Open weights are not automatically open source, and some licences restrict commercial use, redistribution, or certain applications. Open model licences, datasets and deployment obligations covers this properly; this workbook only gets you to a working local setup.
Should I expose my local model to my home network so other devices can use it?
Only with authentication in front of it. By default, Ollama and LM Studio's local servers listen only on your own machine, which is the safe default. Opening it up to your network without adding authentication is a real, not theoretical, risk if anyone else can reach that address.
What can a local model not do as well as a large hosted assistant?
Generally, very long or complex reasoning chains, and broad general knowledge at the edges of what a smaller model was trained on. This follows directly from running something small enough to fit on your machine, not a flaw in the tools themselves. For many everyday tasks, drafting, summarising, and answering questions about your own documents, a well-chosen small local model is genuinely good enough.
When you have finished
Get your certificate of completion
Type your name and download a certificate for this workbook as a PDF, ready to print or to add to LinkedIn. It is made on your own device, so your name is never sent to us. It is a self-declared certificate, not an accredited qualification.
Learn the language
Key terms
- Open-weight model
- A model whose trained parameters are published for anyone to download and run, under some licence. It does not necessarily mean the training code or data is published too.
- Quantisation
- Storing a model's weights with fewer bits than it was trained in, to shrink the file and its memory need, at some cost to quality.
- GGUF
- A single-file format for a model's weights and metadata, used by llama.cpp and the tools built on it, including many quantised versions of popular models.
- Parameter
- One learned number inside a model. The parameter count, multiplied by the bytes each one is stored in, gives roughly the size of the weights.
- VRAM
- The dedicated memory on a graphics card. Fast, and usually the smallest, most limiting memory pool for running a model on a GPU.
- Unified memory
- A single memory pool shared between the CPU and GPU, used by Apple Silicon Macs, which lets the GPU reach a larger amount of memory than a typical separate VRAM pool.
6 of the workbook's 12 terms. The complete glossary is in the workbook.
Follow the evidence
Sources and checks
Facts last checked: .
These workbooks use AI assistance. See how the workbooks are made.
- Ollama (README)Ollama, GitHub
- Ollama FAQ: network exposure and concurrent requestsOllama, GitHub
- LM Studio documentationLM Studio
- GGUF usage with LM StudioHugging Face Hub documentation
- llama.cpp: LLM inference in C/C++ (README)ggml-org, GitHub
- llama.cpp quantize tool documentationggml-org, GitHub
- GGUF file format specificationggml-org, GitHub
- Hugging Face Hub documentation: model cardsHugging Face Hub documentation
- Hugging Face Hub documentation: pickle securityHugging Face Hub documentation
- Qwen3 (README)Qwen team, Alibaba Cloud, GitHub
- Introducing gpt-ossOpenAI
- The Open Source AI Definition 1.0Open Source Initiative
Created by Mickarle Wagstaff-Irons - Micky Irons with the Mickai team. Published by Mickai LTD. Last updated 29 September 2026.
NextKeep going
Where to go next
Recommended for you
GPU memory maths for LLMs and SLMs
A repeatable method for checking whether a language model fits your GPU: weights, KV cache, training memory and adapters, plus a small calculator you can run.
Recommended for you
Inference engines explained
How inference engines run a model: memory maths, quantisation, prefill and decode, batching, hardware, and serving a model safely over a local API.
Recommended for you
Open model licences, datasets and deployment obligations
Decide whether you may use, host and redistribute an 'open' model or dataset: licence families, dataset rights, model cards, AI-BOMs, hash checks and a checklist tool.