Mickai / Learning by making

A small change.
A clearer understanding.

Build a useful interface. Watch text become tokens. See how a model chooses its next word. Inspect a proposed agent action. Four free experiments, with the controls in your hands.

No account · No installation · Your inputs stay in this tab

01 / Make it work

Your coding studio

HTML + CSS + JavaScript

Start with an evidence checklist you can use as a learning exercise. Edit the source, run it and test the result. Each project has a short brief and a clear check.

    How to check your work

    Ready to run

    Tab moves between controls; it is never trapped in an editor. Run applies all three panels. Reset restores the loaded project. Loading another project or leaving this page discards your edits. Download keeps a local, runnable copy.

    Live preview

    Run your project to open its preview.

    Console 0 / 100 messages

    Nothing logged yet.
    What this studio supports

    JavaScript runs in a separate worker. A small DOM bridge supports getElementById, simple tag, class and ID selectors, querySelectorAll (an array), textContent, form value, checked, disabled, className, and click, input, change, submit and keyboard listeners. Use CSS for styling. Full browser APIs, packages, scripts inside HTML and external resources are unavailable.

    The preview uses an opaque sandbox, sanitised HTML and a restrictive network policy. Stop terminates its JavaScript worker; an unresponsive worker is stopped automatically. This does not guarantee recovery from extreme memory use. These labs keep inputs in memory with no upload, telemetry or browser storage. Downloads contain your source and the same runner.

    Limits: 50,000 characters per editor, 400 HTML elements, 100 console messages and 1,000 DOM updates per run. This is a teaching studio, not a compliance assessment.

    02 / Inspect the pieces

    The token workshop

    Step through every merge

    This small character-pair BPE model starts with Unicode code points and repeatedly joins the most frequent neighbouring pair. Spaces and punctuation count too. It illustrates merge learning; it is not a production model tokenizer and cannot estimate its token bill.

    Training text as tokens

    Visible spaces: ␠ · Newlines: ↵ · Tabs: ⇥. These are display labels; the original whitespace is preserved. Up to 40 merges are learned.

    Learned merge rules

      03 / Change the odds

      The sampling bench

      Invented scores / real maths

      Inspect how eight invented next-token scores become a choice: temperature, then top-k, then top-p. Each draw samples the same next-token distribution; this is not an autoregressive conversation.

      0 picks the highest score. Higher values spread probability more evenly.

      Keep at most k candidates, then normalise their probabilities.

      Keep the smallest prefix reaching p of the post-top-k probability mass.

      Inspect each stage. Probabilities are rounded for display.
      TokenScoreTemperatureAfter top-kFinal chanceObserved

      The same settings and seed reproduce the same batch. Change the seed for a different batch. Observed frequencies need not match the exact probabilities.

      No draws yet.

      Two experiments and the exact rules
      1. Choose the creative prompt. Compare temperature 0.3 with 1.8 at k = 8 and p = 1. Which distribution is more concentrated?
      2. Choose equal scores. Set k = 4 and p = 0.5. Two tokens remain: top-p is calculated after normalising the four retained tokens.

      For T > 0, probability is proportional to exp((score − maximum score) / T). At T = 0 this lab uses greedy selection. Equal scores preserve the listed order, including the single winner at T = 0. Top-p includes the token crossing its threshold. Retained probabilities are normalised to sum to one. Real systems may differ. Changing settings clears the observed batch.

      04 / Inspect the boundary

      Would your agent be allowed?

      Synthetic requests / explicit rules

      An agent proposes an action. Your gate decides whether it is allowed, needs approval or is denied. Change one rule, predict the result, then inspect every reason.

      A decision worksheet, not an agent runner. Every actor, approval and dependency is fictional. No action, package, network request or model is executed. Nothing is uploaded or saved automatically. Download record saves your worksheet locally. A passing case does not establish real security.

      1. Choose a proposed action

      Load replaces only the request. Your rule changes stay in place.

      Edit values, keeping the documented fields. Text is data, never code. The approval ledger and principal permissions below are fixed fixtures outside the request.

      2. Edit the gate

      Action allowlist

      Allowlisting an action does not grant document access.

      Review rules

      Try this / 10 minutes

      1. Load the missing-approval case. Predict the result, then run.
      2. Disable exact approval. Run again. What changed, and why does that not prove the action is safe?
      3. Restore approval. Compare a moving dependency version with an unknown source. Which rule catches each?
      4. Run all challenges. Restore the baseline before the final comparison.

      Ready when JavaScript is available.

      3. Inspect the decision

      Not run

        DENY takes precedence over NEEDS APPROVAL. ALLOW means these teaching rules passed. No real action follows.

        Compare all 12 challenge outcomes

        Expected outcomes follow the baseline rules. Deliberately weakening a rule can change an outcome. Matching all 12 is a worksheet result, not a release approval.

        Run the challenge set to compare your current rules with the baseline.
        CaseBaselineYour rulesComparison
        The fixed world, approval ledger and exact rules

        analyst may read brief-alpha, create tasks and send messages in project-alpha. reader may only read brief-alpha. suspended has no grants. No actor may delete a document. Reads have audience self; project writes have audience team-alpha. These fictional permissions are independent of your editable allowlist.

        APP-104 binds request REQ-TASK-01, actor analyst, action create_task, resource project-alpha, text Confirm the evidence owner, audience team-alpha and one item at permission epoch 7. APP-OLD has the same binding at old epoch 6. APP-USED has already been consumed. Other IDs are unknown. The ledger is fixed and is not consumed by this repeatable exercise; a real executor would atomically validate and consume an approval.

        A dependency pin here is exactly three whole numbers, such as 1.4.2. Each component is 0 to 9999 with no leading zeroes. This deliberately narrow teaching rule does not model every package format. The only reviewed source label is reviewed-registry. A real service must establish this from trusted provenance, not accept a caller's label. An exact version or a known source does not prove integrity, compatibility, licence suitability or safe behaviour. No entries means no declared dependencies, not verified absence.

        Limits: eight dependency records, unique dependency names, 50 items, 1,000 characters of action text and 1,200 characters of notes. Unknown JSON fields and malformed inputs deny. Request notes cannot grant permissions. Approval binds the action fields, not the note or dependency labels, which are reviewed separately. There is no cryptographic signature, clock, production authorisation service, malware scan, real approval workflow or model evaluation in this exercise.