assurance · Level 4

How to evaluate a new frontier model the day it lands

A repeatable ten-step method, your own evaluation set, and a worked example on the DeepSeek V4 family.

By Mickarle Wagstaff-Irons - Micky Irons

  • Level 4Frontier
  • 150 min
  • 7 chapters
  • Free PDF, no account
The Boundaryassurance / 04

Start with the essentials

The short answer

To evaluate a new model on the day it lands, start from primary sources: the model card and release notes. Check the licence, architecture, context length, hardware needs, cost and data jurisdiction, and read benchmarks sceptically. Then test about twenty of your own real tasks, scored blind against a rubric, verify any downloaded weights, and record the decision with its date.

What you will learn

  • You will be able to run a repeatable ten-step evaluation of any new model, starting from primary sources.
  • You will know the difference between open-weight and open source, and how to read a licence.
  • You will be able to explain dense versus mixture-of-experts models, and total versus active parameters.
  • You will be able to read benchmark tables sceptically and spot contamination and cherry-picking.
  • You will be able to build a twenty-task evaluation set with a rubric and blind side-by-side scoring.
  • You will be able to check where a download came from and write a dated decision record.

Who it is for

Builders, technical leads and analysts who choose or recommend models and need a method that does not depend on hype. The memory sums in the workbook on inference engines are useful background.

Before you start

  • Comfort with tokens, context windows and the memory sums in Inference engines explained.
  • Access to two models you can compare, for the exercises.

Keep learning

The complete workbook

New models arrive constantly, and the announcement is a claim, not evidence. This workbook gives you a repeatable ten-step checklist, from reading the model card to verifying downloaded weights, and a method for building your own small evaluation set. It ends with a worked example on the DeepSeek V4 family, using only facts checked against official pages and model cards.

  1. 01
    The method in ten steps

    Models land every week, each with a confident announcement. Your job is to turn the claim into a decision you can defend.

    In the workbook · Reading
  2. 02
    Steps 1 to 3: primary sources, licence and architecture

    Most launch-day mistakes come from secondhand facts. Go to the source, then read what it actually says.

    In the workbook · Reading
  3. 03
    Steps 4 to 6: hardware, cost and benchmarks

    Now turn the claims into numbers you can plan with.

    In the workbook · Reading
  4. 04
    Step 7: build your own evaluation set

    This is the step that matters most, and the one people skip. About twenty real tasks and an hour of scoring will tell you more about your use than any leaderboard can.

    In the workbook · 1 exercise
  5. 05
    Steps 8 to 10: safety, jurisdiction and supply chain

    A model can win your tests and still be the wrong choice. These steps catch what accuracy scores do not.

    In the workbook · 1 exercise
  6. 06
    Worked example: the DeepSeek V4 family

    This is an educational analysis of a public model family, not advice or an endorsement, and it says nothing about any Mickai product. Every fact below was verified on 25 September 2026 from official pages and model cards, listed under sources. Anything not confirmed is marked verify before relying on this.

    In the workbook · 1 exercise
  7. 07
    The decision record

    A decision you cannot reconstruct is one you will have to make again. Write it down on one page.

    In the workbook · 1 exercise

Also inside: a 10-point checklist, a glossary of 12 terms and 10 questions and answers to test yourself. 4 hands-on exercises, each with a worked answer at the back where the workbook gives one.

No login, no card, no account. Before the download we ask you to follow Mickai (two quick links). Free to download and use for personal learning, study groups and inside your own team. Please do not resell the workbooks or republish them as your own. Link people to trust-agent.ai instead.

Test yourself

Questions and answers

What is the first thing to do when a new model is released?

Read the publisher's own model card, release notes and technical report, reaching the repository by the link on the publisher's page. Save the URLs and dates. Everything else, including social posts and videos, is secondhand and can be wrong or out of date.

Is an open-weight model the same as open source?

No. Open-weight means the weights can be downloaded. Open source, as defined by the Open Source Initiative's AI definition, also needs information about training data and code, and freedoms to use, study, modify and share. Read the actual licence.

What are total and active parameters?

In a mixture-of-experts model, total parameters are everything stored, which sets memory needs. Active parameters are those used for each token, which sets computation per token. Check the card, because the active count can differ between reading a prompt and writing a reply.

Can I trust benchmark scores in a launch announcement?

Treat them as claims. Ask who ran them, with what settings, against which comparison models, and whether the test data could have leaked into training. Scores are a reason to test the model yourself, not a substitute for doing so.

What is benchmark contamination?

It happens when test questions or their answers appear in a model's training data, so the model has effectively seen the exam. Scores then overstate real ability. It is one reason to keep your own evaluation set private.

Why build my own evaluation set?

Public benchmarks measure general skills, not your work. A private set of about twenty real tasks, with a rubric and blind scoring, shows whether a model does your tasks better than the one you use now, and because it is private it is much harder for publishers to game.

How many tasks do I need?

About twenty is enough to see large differences and to catch obvious failures, and small enough to score in an hour. It will not resolve small gaps. Add tasks over time, especially cases where a model failed you, and version the set.

How do I run a blind comparison?

Give both models the same prompts at the same settings. Have the outputs shuffled and labelled A and B, ideally by someone else, then score them against your rubric without knowing which is which. Repeat runs to see how much outputs vary.

How do I check that downloaded weights are genuine?

Reach the repository from the publisher's announcement, prefer safetensors, note the exact revision, and compare each file's checksum with the hash the hub shows. Read any repository code before allowing it to run. Treat third-party conversions separately.

Why does jurisdiction matter for a hosted model?

Prompts sent to a hosted service are processed and stored under that provider's rules and the laws where it operates. Read its privacy policy and API terms, and involve your data protection lead. Running open weights yourself keeps prompts on your hardware.

When you have finished

Get your certificate of completion

Type your name and download a certificate for this workbook as a PDF, ready to print or to add to LinkedIn. It is made on your own device, so your name is never sent to us. It is a self-declared certificate, not an accredited qualification.

Learn the language

Key terms

Open-weight
A model whose trained weights can be downloaded, under whatever licence the publisher chooses.
Open source (AI)
Under the Open Source AI Definition, freedoms to use, study, modify and share, with information about data, code and parameters. Stricter than open-weight.
Mixture of experts
A design that stores many expert blocks and uses only a few for each token.
Total parameters
All the parameters a model stores. They set the memory needed to hold it.
Active parameters
The parameters used for each token. They set the computation per token.
Context length
The most text a model can take into account at once, measured in tokens.

6 of the workbook's 12 terms. The complete glossary is in the workbook.

Follow the evidence

Sources and checks

Facts last checked: .

These workbooks use AI assistance. See how the workbooks are made.

  1. DeepSeek V4 Preview Release (2026/04/24)DeepSeek API Docs
  2. DeepSeek-V4-Pro GA Release (2026/08/13)DeepSeek API Docs
  3. DeepSeek-V4-Flash-Vision-Exp Release (2026/08/21)DeepSeek API Docs
  4. DeepSeek-V4.1-Flash: Smarter, Faster, More Efficient (2026/09/10)DeepSeek API Docs
  5. Change LogDeepSeek API Docs
  6. Models and PricingDeepSeek API Docs
  7. DeepSeek Privacy Policy (last update 10 February 2026)DeepSeek
  8. Model card: DeepSeek-V4-ProDeepSeek on Hugging Face
  9. Model card: DeepSeek-V4-FlashDeepSeek on Hugging Face
  10. Model card: DeepSeek-V4-Flash-0731DeepSeek on Hugging Face
  11. Model card: DeepSeek-V4-Pro-0813DeepSeek on Hugging Face
  12. Model card: DeepSeek-V4.1-FlashDeepSeek on Hugging Face
  13. Files listing: DeepSeek-V4-Pro (used for size sums)DeepSeek on Hugging Face
  14. Files listing: DeepSeek-V4-Pro-Base (used for size sums)DeepSeek on Hugging Face
  15. Files listing: DeepSeek-V4-Flash (used for size sums)DeepSeek on Hugging Face
  16. Files listing: DeepSeek-V4.1-Flash (used for size sums)DeepSeek on Hugging Face
  17. The Open Source AI Definition 1.0Open Source Initiative
  18. Pickle scanning and securityHugging Face Hub documentation

Created by Mickarle Wagstaff-Irons - Micky Irons with the Mickai team. Published by Mickai LTD. Last updated 25 September 2026.

NextKeep going

Where to go next