foundations · Level 2
Which AI model should I use? Picking the right AI for the job
A practical decision framework, a fair five-prompt bake-off, and a sceptical way to read leaderboards.

Start with the essentials
The short answer
The best AI model is the one that does your tasks well, at a speed, cost and privacy level you can accept. Start by naming the task, shortlist two or three options, run the same five prompts on each, and score the results yourself. Public leaderboards are a starting point only, because they measure other people's tasks, not yours.
What you will learn
- You will be able to describe your own work as a set of task types and say which factors matter most for each.
- You will understand the difference between an assistant and a model, and between closed and open-weight models.
- You will be able to run a fair five-prompt bake-off and score it consistently.
- You will be able to read a leaderboard or benchmark claim critically and ask the right questions of it.
- You will have a simple model log and a rule for when to test again.
Who it is for
Anyone who has used a chatbot and now faces a choice between several tools, plans or models, for work, study or a personal project. No coding is needed.
Before you start
- Some experience of using a chatbot. The workbooks What is AI and Your first prompts are ideal preparation.
Keep learning
The complete workbook
This workbook gives you a repeatable way to choose an AI. You will separate assistants from models, weigh closed models against open-weight ones, run a fair personal bake-off, learn why benchmark charts can mislead, and start a model log so your choice stays current.
- 01Start with the job, not the brandIn the workbook · 1 exercise
Many poor AI choices begin with the wrong question. 'Which is the best AI?' has no answer. 'Which AI is best for this task, for me, this month?' does.
- 02Closed models and open-weight modelsIn the workbook · 1 exercise
There are two broad ways to get a model. Both are useful. They differ in who runs it, who sees your text, and who controls change.
- 03What actually mattersIn the workbook · 1 exercise
Eight things can matter. Rarely do all eight matter for one task. Your skill is knowing which two or three do.
- 04The five-prompt bake-offIn the workbook · 1 exercise
A bake-off is a small, fair test you run yourself. Under an hour of it can tell you more about your own needs than any chart.
- 05Reading leaderboards and benchmarks criticallyIn the workbook · 1 exercise
Charts are persuasive, and often unhelpful for your decision. Learn to ask what they really measure.
- 06The landscape today, and your model logIn the workbook · 1 exercise
Any list of AI tools is out of date almost as soon as it is printed. Here is a dated snapshot, and then the habit that matters more.
Also inside: a 9-point checklist, a glossary of 12 terms and 10 questions and answers to test yourself. 6 hands-on exercises, each with a worked answer at the back where the workbook gives one.
No login, no card, no account. Before the download we ask you to follow Mickai (two quick links). Free to download and use for personal learning, study groups and inside your own team. Please do not resell the workbooks or republish them as your own. Link people to trust-agent.ai instead.
Test yourself
Questions and answers
Which AI model is the best?
There is no single best model. The right one depends on your task, your budget, how much you care about privacy and how much effort you will spend on setup. Shortlist two or three, then test them on your own work with the same prompts.
What is the difference between an assistant and a model?
The model is the trained engine. The assistant is the product built around it, with the interface, hidden instructions, memory, search, file upload and safety filters. The same model can feel different in different assistants.
What does open-weight mean?
The developer publishes the model's weights so anyone can download and run them, under a licence. It does not necessarily mean the training data or training code are published, and the licence may limit how you can use it.
Is an open-weight model better for privacy?
It can be, when you run it on your own hardware, because your text need not leave your control. A hosted copy run by someone else is different. You also take on security, updates and setup yourself.
Do I need to pay for a better model?
Not always. Test a free option against your five prompts first. Pay only if the paid option clearly saves you time or improves quality on work that matters. Plans and free allowances change often, so check current terms.
How many prompts make a fair test?
Five well-chosen prompts, each run twice, is a practical start. It is a small sample that can reveal a clear mismatch but not prove a universal winner. Use the same wording, fresh chats and blind scoring where possible.
What is benchmark contamination?
It is when a benchmark's questions or answers appear in a model's training text. The model may then have effectively seen the test, so a high score can overstate how well it handles new problems.
Can I trust AI leaderboards?
Treat them as a way to build a shortlist, not as the final answer. Ask who ran the test, what was measured, whether it resembles your work, how large the gap is, and which version and date it covers.
Does a longer context window mean a better model?
No. A longer window lets a model take in more text at once, which helps with big documents. It does not mean the answers are better, and details in long text can still be missed. Test with your own files.
How often should I re-test my choice?
Re-test when a candidate releases something new, when your plan or its terms change, when answers feel worse, or when your work changes. Otherwise, every three months is a sensible habit. Your model log shows when.
When you have finished
Get your certificate of completion
Type your name and download a certificate for this workbook as a PDF, ready to print or to add to LinkedIn. It is made on your own device, so your name is never sent to us. It is a self-declared certificate, not an accredited qualification.
Learn the language
Key terms
- Assistant
- A finished AI product built around a model, including the interface, hidden instructions, memory, tools and safety filters.
- Model
- The trained engine that turns your input into a reply. Different assistants can use the same model, and one assistant can offer several.
- Weights
- The numbers learned in training that make up a model. Whoever has the weights can run the model.
- Open-weight model
- A model whose weights are published for download under a licence. It does not mean training data or code are also published.
- Closed model
- A model that runs on its developer's computers and is used through their service. You cannot download its weights.
- API
- A way for software to talk to a model directly, without the chat app. Developers use it to build their own tools.
6 of the workbook's 12 terms. The complete glossary is in the workbook.
Follow the evidence
Sources and checks
Facts last checked: .
These workbooks use AI assistance. See how the workbooks are made.
- Claude by AnthropicAnthropic
- Gemini, your AI assistant from GoogleGoogle
- ChatGPT release notesOpenAI Help Center
- Introducing the new CopilotMicrosoft
- GemmaGoogle DeepMind
- Introducing gpt-ossOpenAI
- Qwen3Qwen team, Alibaba Cloud
- Mistral AIMistral AI
Created by Mickarle Wagstaff-Irons - Micky Irons with the Mickai team. Published by Mickai LTD. Last updated 25 September 2026.
NextKeep going
Where to go next
Recommended for you
The free AI toolkit: how to choose tools and stay safe
How free AI tools work, how to choose and test them, how to check their current terms and privacy, how to spot fake apps and scams, and how to build your own starter stack.
Recommended for you
Reasoning models and reasoning engines
What reasoning models are, when extra thinking pays off, how to control it, why a reasoning trace is not proof, and how classical reasoning engines differ.
Recommended for you
How to evaluate a new frontier model the day it lands
A ten-step method for judging a new model: primary sources, licence, architecture, hardware, benchmarks, your own tests, safety, jurisdiction and a decision record.