assurance · Level 4
How to evaluate a new frontier model the day it lands
A repeatable ten-step method, your own evaluation set, and a worked example on the DeepSeek V4 family.

Start with the essentials
The short answer
To evaluate a new model on the day it lands, start from primary sources: the model card and release notes. Check the licence, architecture, context length, hardware needs, cost and data jurisdiction, and read benchmarks sceptically. Then test about twenty of your own real tasks, scored blind against a rubric, verify any downloaded weights, and record the decision with its date.
What you will learn
- You will be able to run a repeatable ten-step evaluation of any new model, starting from primary sources.
- You will know the difference between open-weight and open source, and how to read a licence.
- You will be able to explain dense versus mixture-of-experts models, and total versus active parameters.
- You will be able to read benchmark tables sceptically and spot contamination and cherry-picking.
- You will be able to build a twenty-task evaluation set with a rubric and blind side-by-side scoring.
- You will be able to check where a download came from and write a dated decision record.
Who it is for
Builders, technical leads and analysts who choose or recommend models and need a method that does not depend on hype. The memory sums in the workbook on inference engines are useful background.
Before you start
- Comfort with tokens, context windows and the memory sums in Inference engines explained.
- Access to two models you can compare, for the exercises.
Keep learning
The complete workbook
New models arrive constantly, and the announcement is a claim, not evidence. This workbook gives you a repeatable ten-step checklist, from reading the model card to verifying downloaded weights, and a method for building your own small evaluation set. It ends with a worked example on the DeepSeek V4 family, using only facts checked against official pages and model cards.
- 01The method in ten stepsIn the workbook · Reading
Models land every week, each with a confident announcement. Your job is to turn the claim into a decision you can defend.
- 02Steps 1 to 3: primary sources, licence and architectureIn the workbook · Reading
Most launch-day mistakes come from secondhand facts. Go to the source, then read what it actually says.
- 03Steps 4 to 6: hardware, cost and benchmarksIn the workbook · Reading
Now turn the claims into numbers you can plan with.
- 04Step 7: build your own evaluation setIn the workbook · 1 exercise
This is the step that matters most, and the one people skip. About twenty real tasks and an hour of scoring will tell you more about your use than any leaderboard can.
- 05Steps 8 to 10: safety, jurisdiction and supply chainIn the workbook · 1 exercise
A model can win your tests and still be the wrong choice. These steps catch what accuracy scores do not.
- 06Worked example: the DeepSeek V4 familyIn the workbook · 1 exercise
This is an educational analysis of a public model family, not advice or an endorsement, and it says nothing about any Mickai product. Every fact below was verified on 25 September 2026 from official pages and model cards, listed under sources. Anything not confirmed is marked verify before relying on this.
- 07The decision recordIn the workbook · 1 exercise
A decision you cannot reconstruct is one you will have to make again. Write it down on one page.
Also inside: a 10-point checklist, a glossary of 12 terms and 10 questions and answers to test yourself. 4 hands-on exercises, each with a worked answer at the back where the workbook gives one.
No login, no card, no account. Before the download we ask you to follow Mickai (two quick links). Free to download and use for personal learning, study groups and inside your own team. Please do not resell the workbooks or republish them as your own. Link people to trust-agent.ai instead.
Test yourself
Questions and answers
What is the first thing to do when a new model is released?
Read the publisher's own model card, release notes and technical report, reaching the repository by the link on the publisher's page. Save the URLs and dates. Everything else, including social posts and videos, is secondhand and can be wrong or out of date.
Is an open-weight model the same as open source?
No. Open-weight means the weights can be downloaded. Open source, as defined by the Open Source Initiative's AI definition, also needs information about training data and code, and freedoms to use, study, modify and share. Read the actual licence.
What are total and active parameters?
In a mixture-of-experts model, total parameters are everything stored, which sets memory needs. Active parameters are those used for each token, which sets computation per token. Check the card, because the active count can differ between reading a prompt and writing a reply.
Can I trust benchmark scores in a launch announcement?
Treat them as claims. Ask who ran them, with what settings, against which comparison models, and whether the test data could have leaked into training. Scores are a reason to test the model yourself, not a substitute for doing so.
What is benchmark contamination?
It happens when test questions or their answers appear in a model's training data, so the model has effectively seen the exam. Scores then overstate real ability. It is one reason to keep your own evaluation set private.
Why build my own evaluation set?
Public benchmarks measure general skills, not your work. A private set of about twenty real tasks, with a rubric and blind scoring, shows whether a model does your tasks better than the one you use now, and because it is private it is much harder for publishers to game.
How many tasks do I need?
About twenty is enough to see large differences and to catch obvious failures, and small enough to score in an hour. It will not resolve small gaps. Add tasks over time, especially cases where a model failed you, and version the set.
How do I run a blind comparison?
Give both models the same prompts at the same settings. Have the outputs shuffled and labelled A and B, ideally by someone else, then score them against your rubric without knowing which is which. Repeat runs to see how much outputs vary.
How do I check that downloaded weights are genuine?
Reach the repository from the publisher's announcement, prefer safetensors, note the exact revision, and compare each file's checksum with the hash the hub shows. Read any repository code before allowing it to run. Treat third-party conversions separately.
Why does jurisdiction matter for a hosted model?
Prompts sent to a hosted service are processed and stored under that provider's rules and the laws where it operates. Read its privacy policy and API terms, and involve your data protection lead. Running open weights yourself keeps prompts on your hardware.
When you have finished
Get your certificate of completion
Type your name and download a certificate for this workbook as a PDF, ready to print or to add to LinkedIn. It is made on your own device, so your name is never sent to us. It is a self-declared certificate, not an accredited qualification.
Learn the language
Key terms
- Open-weight
- A model whose trained weights can be downloaded, under whatever licence the publisher chooses.
- Open source (AI)
- Under the Open Source AI Definition, freedoms to use, study, modify and share, with information about data, code and parameters. Stricter than open-weight.
- Mixture of experts
- A design that stores many expert blocks and uses only a few for each token.
- Total parameters
- All the parameters a model stores. They set the memory needed to hold it.
- Active parameters
- The parameters used for each token. They set the computation per token.
- Context length
- The most text a model can take into account at once, measured in tokens.
6 of the workbook's 12 terms. The complete glossary is in the workbook.
Follow the evidence
Sources and checks
Facts last checked: .
These workbooks use AI assistance. See how the workbooks are made.
- DeepSeek V4 Preview Release (2026/04/24)DeepSeek API Docs
- DeepSeek-V4-Pro GA Release (2026/08/13)DeepSeek API Docs
- DeepSeek-V4-Flash-Vision-Exp Release (2026/08/21)DeepSeek API Docs
- DeepSeek-V4.1-Flash: Smarter, Faster, More Efficient (2026/09/10)DeepSeek API Docs
- Change LogDeepSeek API Docs
- Models and PricingDeepSeek API Docs
- DeepSeek Privacy Policy (last update 10 February 2026)DeepSeek
- Model card: DeepSeek-V4-ProDeepSeek on Hugging Face
- Model card: DeepSeek-V4-FlashDeepSeek on Hugging Face
- Model card: DeepSeek-V4-Flash-0731DeepSeek on Hugging Face
- Model card: DeepSeek-V4-Pro-0813DeepSeek on Hugging Face
- Model card: DeepSeek-V4.1-FlashDeepSeek on Hugging Face
- Files listing: DeepSeek-V4-Pro (used for size sums)DeepSeek on Hugging Face
- Files listing: DeepSeek-V4-Pro-Base (used for size sums)DeepSeek on Hugging Face
- Files listing: DeepSeek-V4-Flash (used for size sums)DeepSeek on Hugging Face
- Files listing: DeepSeek-V4.1-Flash (used for size sums)DeepSeek on Hugging Face
- The Open Source AI Definition 1.0Open Source Initiative
- Pickle scanning and securityHugging Face Hub documentation
Created by Mickarle Wagstaff-Irons - Micky Irons with the Mickai team. Published by Mickai LTD. Last updated 25 September 2026.
NextKeep going
Where to go next
Recommended for you
Inference engines explained
How inference engines run a model: memory maths, quantisation, prefill and decode, batching, hardware, and serving a model safely over a local API.
Recommended for you
Reasoning models and reasoning engines
What reasoning models are, when extra thinking pays off, how to control it, why a reasoning trace is not proof, and how classical reasoning engines differ.
Recommended for you
Which AI model should I use? Picking the right AI for the job
How to choose between AI assistants and models: what to weigh, how to run a fair five-prompt bake-off, and how to read leaderboards and benchmarks without being fooled.