models · Level 3
Tokenisation: train and inspect a small tokenizer
Build a byte-pair encoding tokeniser from scratch in plain Python, check exact round trips, and see why tokens shape cost, context and behaviour.

Start with the essentials
The short answer
A tokeniser turns text into a list of numbered pieces called tokens, and back again. Byte-pair encoding builds the vocabulary by repeatedly merging the most frequent neighbouring pair, and starting from single bytes means valid Unicode text has a byte-level fallback. Tokens matter because models, context windows and many prices count them, and the same meaning can take very different numbers of tokens.
What you will learn
- You will be able to explain characters, words, bytes and subwords as units of text, and what each one costs.
- You will be able to run byte-pair encoding merges by hand and in plain Python, and encode words the merges never saw.
- You will be able to test a tokeniser's round trip and explain when it is lossless, and handle unknown text and special tokens safely.
- You will be able to explain, with worked numbers, how vocabulary size trades sequence length against embedding-table size.
- You will be able to measure how languages, numbers, capitals, spaces and code change token counts, and say what a toy measurement can and cannot show.
Who it is for
People who can run a small Python script and want to see exactly what a language model reads. Everything uses Python's standard library: nothing to install, download or sign up for, and no network. The code was tested with Python 3.12.10 on Windows 11; other versions and systems were not tested.
Before you start
- Enough Python to save and run a script from a terminal and to read a loop and a function, as taught in Python: from first script to a useful automation. A rough idea of what a language model does, as in What is AI, helps. The only maths is counting, multiplying and dividing.
Read a sample · Chapter 01 of 06
Characters, words, bytes and subwords
A text language model typically receives token ids that select learned embeddings. The tokeniser determines how the text becomes those ids.
A tokeniser (often spelt tokenizer) turns text into a list of whole numbers called token ids, and turns ids back into text. Each id stands for one token, a piece of text from a fixed list called the vocabulary. This workbook concerns a text model's token interface. The choice of pieces shapes its input and output sequences.
There are four common choices of piece: characters, words, bytes and subwords. Start by counting the first three. Everything in this workbook uses only Python's standard library: nothing to install, no downloads, no network. Make a folder called tokeniser, save this as units.py in it, open a terminal there and run python units.py (on macOS and Linux, type python3 wherever this workbook says python).
# units.py: count the same text as characters, words and bytes
print("text chars words bytes byte values")
for text in ["tokens", "New York", "caf\u00e9", "\u00a35", "\u4e66", "\U0001F600"]:
data = text.encode("utf-8")
print(f"{ascii(text):13} {len(text):5} {len(text.split()):5} {len(data):5} {list(data)}")
print(" ".join("New York".split()) == "New York")text chars words bytes byte values
'tokens' 6 1 6 [116, 111, 107, 101, 110, 115]
'New York' 9 2 9 [78, 101, 119, 32, 32, 89, 111, 114, 107]
'caf\xe9' 4 1 5 [99, 97, 102, 195, 169]
'\xa35' 2 1 3 [194, 163, 53]
'\u4e66' 1 1 3 [228, 185, 166]
'\U0001f600' 1 1 4 [240, 159, 152, 128]
Falseascii() prints any character outside plain ASCII as an escape, such as \xe9 for é, so the output shows on any terminal. The fifth text is the Chinese character for 'book' and the sixth is a smiling-face emoji. This workbook's fonts cannot print either, so the code spells them as escapes.
- Characters. Python counts Unicode characters (code points), so 'café' is 4.
- Words. Splitting on spaces is simple but loses information: the last line is False because 'New York' had two spaces, and joining the words gives back only one.
- Bytes. Computers store text as bytes, whole numbers from 0 to 255. UTF-8, the usual encoding, uses 1 byte for each ASCII character and 2 to 4 bytes for others (RFC 3629): 'é' takes 2, the Chinese character 3 and the emoji 4.
Scroll sideways to see every column.
| Unit | Vocabulary size | Tokens per text | Text it cannot handle |
|---|---|---|---|
| Characters | Over 130,000 in the 2019 GPT-2 paper; a historical example | Many | Any character left out of the vocabulary |
| Words | Every word seen, and still growing | Few | Any new word, name or typo |
| Bytes | 256 | Most: 1 to 4 per character | None |
| Subwords (BPE) | A size you choose: GPT-2 used 50,257 | In between | None, if built on bytes |
Subwords are the middle ground the GPT-2 paper describes: frequent words become single tokens, while rare words fall back to smaller pieces. Byte-pair encoding, the subject of the next chapter, is the usual way to choose them.
Try it yourself · Activity 01
10 minCount before you run
For each text, predict its characters, words and bytes. Then add the texts to the list in units.py and run it.
"na\u00efve"(naïve) and"\u00a310"(£10)."Stra\u00dfe"(Straße, German for street) and"\u65e5\u672c"(the two-character Japanese name for Japan)."a b ": two spaces in the middle and one at the end.
Which units give back exactly the text you started with, and which does it with a fixed vocabulary?
Worked answer
naïve: 5 characters, 1 word, 6 bytes (ï takes 2). £10: 3, 1 and 4 (£ takes 2). Straße: 6, 1 and 7 (ß takes 2). The Japanese name: 2 characters, 1 word, 6 bytes (3 each). 'a b ': 5 characters, 2 words, 5 bytes; splitting into words throws the extra spaces away. Characters and bytes both keep everything, but only bytes do it with a fixed vocabulary of 256.
Read a sample · Chapter 02 of 06
Byte-pair encoding by hand
BPE builds a vocabulary by gluing together the most frequent neighbouring pair, again and again. Do it once on paper and the code will hold no surprises.
Byte-pair encoding (BPE) began as a data compression method. Sennrich, Haddow and Birch adapted it to split words into subwords for machine translation, so that rare and unseen words could be built from smaller, known pieces. Training has four steps.
- Split every word into single symbols.
- Count every pair of neighbouring symbols, weighting each word by how often it occurs.
- Merge the most frequent pair into one new symbol, and write that merge down.
- Repeat until you have made the number of merges you chose.
The paper notes that the final vocabulary is the starting symbols plus one per merge, and that pairs never cross word boundaries. The ordered list of merges is the trained tokeniser.
Worked example. A text contains 'and' 5 times, 'band' 3, 'sand' 2, 'hand' 4 and 'bank' once. Spaces separate the symbols of each word.
Scroll sideways to see every column.
| Merge | Highest pair counts | Winner | Words afterwards |
|---|---|---|---|
| 1 | a+n 15, n+d 14, b+a 4, h+a 4 | an | an d, b an d, s an d, h an d, b an k |
| 2 | an+d 14, b+an 4, h+an 4 | and | and, b and, s and, h and, b an k |
| 3 | h+and 4, b+and 3, s+and 2 | hand | and, b and, s and, hand, b an k |
a+n scores 15 because every word contains it: 5 + 3 + 2 + 4 + 1.
Encoding a new word replays the merges in the order they were learned. 'brand' starts as b r a n d. Merge 1 gives b r an d, merge 2 gives b r and, and merge 3 finds no h+and. The result is three tokens: b, r and 'and'. 'handstand' becomes hand, s, t, and. Any word made of known symbols can be encoded, because the worst case is single symbols. The paper notes that characters never seen in training can still be unknown; chapter 3 avoids that by starting from bytes.
A companion lab. The free Tokeniser lab on the trust-agent.ai Labs page runs the same idea in your browser: type a training text and step through the merges. It starts from characters and marks the end of each word, whereas this workbook starts from bytes and attaches each space to the start of the next word. You do not need it to finish this workbook.
Try it yourself · Activity 02
15 minThree merges on paper
A text contains 'ten' 4 times, 'tent' 3, 'tents' 2, 'sent' 6 and 'set' once.
- Count the pairs and make three merges, writing down each winner and its count.
- Encode 'tense' and 'scent' with your three merges.
Why does 'scent' end in a token that 'tense' cannot use?
Worked answer
Merge 1: e+n, 15 (4 + 3 + 2 + 6), ahead of n+t at 11. Merge 2: en+t, 11 (3 + 2 + 6), ahead of t+en at 9. Merge 3: s+ent, 6, ahead of t+ent at 5. 'tense' becomes t, en, s, e: 'ent' needs a t straight after 'en'. 'scent' becomes s, c, ent: 'sent' cannot form because c sits between s and ent. Chapter 3 checks this by machine.
Read a sample · Chapter 03 of 06
Build a byte-level BPE in plain Python
About sixty lines of standard-library Python train a real, if tiny, BPE tokeniser. Starting from bytes means valid Unicode text has a byte-level fallback.
Save this as bpe.py in your tokeniser folder. It is the whole tokeniser, and every later script imports it.
# bpe.py: a tiny byte-level BPE tokeniser, standard library only
import re
from collections import Counter
# Pre-tokenise into chunks: a word with its leading space, a run of
# punctuation with its leading space, or a run of whitespace.
# Merges never cross from one chunk into the next.
CHUNK = re.compile(r" ?\w+| ?[^\w\s]+|\s+")
def merge(ids, pair, new_id):
"""Replace each occurrence of `pair` in `ids`, left to right, with `new_id`."""
out, i = [], 0
while i < len(ids):
if i + 1 < len(ids) and (ids[i], ids[i + 1]) == pair:
out.append(new_id)
i += 2
else:
out.append(ids[i])
i += 1
return tuple(out)
def train(text, n_merges):
"""Learn up to n_merges merges. Returns {(left id, right id): new id}."""
chunks = Counter(tuple(c.encode("utf-8")) for c in CHUNK.findall(text))
merges = {}
for new_id in range(256, 256 + n_merges):
pairs = Counter()
for chunk, freq in chunks.items():
for pair in zip(chunk, chunk[1:]):
pairs[pair] += freq
if not pairs:
break # nothing left to merge
best = max(pairs, key=pairs.get) # ties: the pair counted first
merges[best] = new_id
chunks = Counter({merge(c, best, new_id): f for c, f in chunks.items()})
return merges
def vocab(merges):
"""Map every token id to its bytes. Ids 0 to 255 are the raw bytes."""
table = {i: bytes([i]) for i in range(256)}
for (a, b), new_id in merges.items():
table[new_id] = table[a] + table[b]
return table
def encode(text, merges):
ids = []
for c in CHUNK.findall(text):
chunk = tuple(c.encode("utf-8"))
for pair, new_id in merges.items(): # replay merges in learned order
chunk = merge(chunk, pair, new_id)
ids.extend(chunk)
return ids
def decode(ids, table):
return b"".join(table[i] for i in ids).decode("utf-8", errors="replace")
def pieces(ids, table):
"""Show each token in [brackets], non-ASCII bytes as \\x escapes."""
return "".join("[" + str(table[i])[2:-1] + "]" for i in ids)CHUNKpre-tokenises: it cuts text into words, runs of punctuation and runs of whitespace, and each word or punctuation run keeps one space in front of it. Merges never cross chunks. This is a simplified cousin of the GPT-2 paper's rule, which stops merges across kinds of character, so that 'dog', 'dog.' and 'dog!' do not each take a token, but lets spaces join.trainturns each chunk into its UTF-8 bytes (ids 0 to 255), counts pairs weighted by how often each chunk occurs, and gives each winner the next free id from 256. Python'smaxreturns the first of several equal counts, and dictionaries keep insertion order, so every run learns the same merges.vocabspells out every id as bytes.encodereplays the merges in order, as in chapter 2.decodejoins the bytes and reads them as UTF-8.piecesprints each token in square brackets.
Now a training text. Save this as corpus.py: nine invented sentences (105 words), plus one sentence kept out of training to test on.
# corpus.py: an invented training text, and a sentence kept out of training
CORPUS = """The town library opens at 9am and closes at 6pm.
Readers come in from the rain and look for the books they want.
The children's corner is the busiest part of the library.
The children read picture books, and older readers read novels.
Readers can borrow 12 books at a time and keep them for 21 days.
If you return books late, you cannot borrow more until they are back.
The reading room has quiet desks, and the library lends laptops.
On Saturday the reading group talks about the books they are reading.
In the evening the library closes, and the last readers walk home."""
HELD_OUT = "The children are reading their new books in the quiet library."Save and run train_demo.py.
# train_demo.py: learn 60 merges and look at them
from bpe import train, vocab, encode, pieces
from corpus import CORPUS, HELD_OUT
merges = train(CORPUS, 60)
table = vocab(merges)
learned = [str(table[i])[2:-1] for i in merges.values()]
print(len(table), "tokens in the vocabulary")
print("first 10 merges:", learned[:10])
print("last 5 merges:", learned[-5:])
ids = encode(HELD_OUT, merges)
print(len(HELD_OUT.encode("utf-8")), "bytes ->", len(ids), "tokens")
print(pieces(ids, table))316 tokens in the vocabulary
first 10 merges: ['he', ' t', ' the', 're', ' a', ' l', 'ad', ' b', ' c', ' re']
last 5 merges: [' child', ' childre', ' children', ' p', ' readers']
62 bytes -> 24 tokens
[The][ children][ a][re][ reading][ the][i][r][ ][n][e][w][ books][ ][in][ the][ ][q][u][i][e][t][ library][.]Each merge adds one token, so 60 merges give 256 + 60 = 316. Early merges are common fragments: 'he', ' t', then ' the'. Later merges build whole words, with stepping stones on the way: ' childre' exists only to become ' children'.
The held-out sentence shrinks from 62 bytes to 24 tokens. Words that were frequent in training are single tokens, while 'their' and 'new' (absent from the training text) and 'quiet' (seen once) fall apart into single bytes. Real tokenisers do the same on vastly more text; this is the idea in miniature.
To check chapter 2 by machine, put each word on its own line so that no word gets a leading space. Save check_hand.py.
# check_hand.py: replay chapter 2 by machine
from bpe import train, vocab, encode, pieces
for words, tests in [
(["and"] * 5 + ["band"] * 3 + ["sand"] * 2 + ["hand"] * 4 + ["bank"], ["brand", "handstand"]),
(["ten"] * 4 + ["tent"] * 3 + ["tents"] * 2 + ["sent"] * 6 + ["set"], ["tense", "scent"]),
]:
merges = train("\n".join(words), 3) # one word per line: no spaces
table = vocab(merges)
print([str(table[i])[2:-1] for i in merges.values()])
for w in tests:
print(w, "->", pieces(encode(w, merges), table))Try it yourself · Activity 03
10 minCheck your hand merges
Use check_hand.py to test your answers from chapter 2.
- Predict its output from your paper answers, then run it.
- Change the 3 in the
traincall to 4. Predict the fourth merge for each word list, then run it again.
Did the fourth merge change how any of the four test words are encoded? Why not?
Worked answer
The output is ['an', 'and', 'hand'], then brand -> [b][r][and] and handstand -> [hand][s][t][and], then ['en', 'ent', 'sent'], then tense -> [t][en][s][e] and scent -> [s][c][ent]. With 4 merges the lists become ['an', 'and', 'hand', 'band'] and ['en', 'ent', 'sent', 'tent'] (b+and scores 3; t+ent scores 5). The encodings stay the same: none of the four test words contains b next to 'and' or t next to 'ent'.
Read a sample · Chapter 04 of 06
Encode, decode and check the round trip
A tokeniser must give back exactly the text it was given. Test that directly, then add a special token that typed text cannot forge.
The SentencePiece paper calls this lossless tokenisation: decoding the encoding gives back the original text (after any normalisation the tokeniser applies first). A simple word splitter that discards whitespace cannot do this. A closed word vocabulary also needs a policy for unseen words; a system with fallbacks can behave differently. roundtrip.py shows both.
# roundtrip.py: decoding the tokens must give back exactly the same text
from bpe import train, vocab, encode, decode, pieces
from corpus import CORPUS, HELD_OUT
known = set(CORPUS.lower().split()) # a word-level vocabulary
print([w if w in known else "<unk>" for w in HELD_OUT.lower().split()])
merges = train(CORPUS, 60)
table = vocab(merges)
for text in [
"Readers read.",
"Readers read. ", # two spaces, one trailing
"Tab\there,\nnew line",
"Caf\u00e9 cr\u00e8me for \u00a33.50", # accents and a pound sign
"Emoji \U0001F600 and \u4e66", # an emoji, a Chinese character
"Zyxwv qjk", # letters never seen together
]:
ids = encode(text, merges)
print(decode(ids, table) == text, len(ids), pieces(ids, table))
ids = encode("caf\u00e9", merges)
print(ascii([decode([i], table) for i in ids])) # one token at a time['the', 'children', 'are', 'reading', '<unk>', '<unk>', 'books', 'in', 'the', 'quiet', 'library.']
True 3 [Readers][ read][.]
True 7 [Readers][ ][ ][re][ad][.][ ]
True 14 [T][a][b][\t][he][re][,][\n][n][e][w][ l][in][e]
True 18 [C][a][f][\xc3][\xa9][ c][r][\xc3][\xa8][me][ for][ ][\xc2][\xa3][3][.][5][0]
True 15 [E][m][o][j][i][ ][\xf0][\x9f][\x98][\x80][ and][ ][\xe4][\xb9][\xa6]
True 9 [Z][y][x][w][v][ ][q][j][k]
['c', 'a', 'f', '\ufffd', '\ufffd']The first line uses this deliberately simple word-level vocabulary built from the training text: 'their' and 'new' become <unk>, an unknown token, and what they said is lost. Every BPE line prints True: double spaces, a tab, a newline, accents, an emoji, a Chinese character and nonsense letters all survive, because any text is bytes and every byte is in the vocabulary. The GPT-2 paper gives the same reason for starting from bytes: the model can then handle any Unicode string.
These tests provide evidence for these inputs, not a proof from testing alone. The round trip follows from preserving every UTF-8 byte and using only reversible merges. This applies to valid Unicode text supplied to encode, with the same vocabulary. Python strings containing unpaired surrogates fail UTF-8 encoding. Arbitrary token sequences may contain invalid UTF-8: use strict decoding to detect that, or a proper incremental decoder when streaming. Replacement decoding deliberately loses information for malformed sequences.
Look at the second line. The double space became its own chunk, so 'read' lost its leading space and split into [re][ad]. The trailing space is a token of its own.
Special tokens are ids with a job rather than text: marking the end of a document, padding a batch or standing in for unknown text. SentencePiece reserves ids for unknown, beginning-of-sentence, end-of-sentence and padding symbols. Save and run specials.py.
# specials.py: an end-of-document token that typed text cannot forge
from bpe import train, vocab, encode, decode, pieces
from corpus import CORPUS
merges = train(CORPUS, 60)
table = vocab(merges)
END = 256 + len(merges) # the next free id
table[END] = b"<|end|>" # how it looks when decoded
def encode_docs(docs):
ids = []
for doc in docs:
ids += encode(doc, merges) + [END] # only this line adds END
return ids
ids = encode_docs(["Readers read.", "Type <|end|> here."])
print(decode(ids, table))
print(ids.count(END), "real END tokens")
print(pieces(encode("<|end|>", merges), table))Readers read.<|end|>Type <|end|> here.<|end|>
2 real END tokens
[<][|][e][nd][|][>]The decoded text shows three <|end|> markers, but only two are real. The middle one was typed as part of a document, and encode treated it as six ordinary tokens. So: add special ids only in your own code, and check ids, never decoded text, which looks the same either way. One widely used open-source tokeniser library explains in its source code that, by default, it refuses to encode text that matches a special token, because special tokens can be used to trick a model.
Try it yourself · Activity 04
15 minTry to break the round trip
Add texts to the list in roundtrip.py and look for one that prints False.
- Try your name, a line of code, a sentence in another language (typed or pasted), and text full of spaces and tabs. Save the file as UTF-8, the usual default.
- In
bpe.py, changeerrors="replace"toerrors="strict", runroundtrip.pyagain, then change it back.
Why does replacement decoding hide malformed bytes, and when should strict decoding be used instead?
Worked answer
Every text prints True: whatever you type is stored as UTF-8 bytes, each byte is a token, and decode joins the same bytes back. With 'strict', the whole-text lines still print True, but the last line stops with UnicodeDecodeError: 'utf-8' codec can't decode byte 0xc3 in position 0: unexpected end of data. Half of 'é' is not valid UTF-8. 'replace' hides that by printing a replacement character; 'strict' makes it visible.
Read a sample · Chapter 05 of 06
Vocabulary size: shorter texts or a bigger table
More merges make texts shorter but the vocabulary bigger. Both sides cost something, so vocabulary size is a trade-off, not a setting to maximise.
Save and run sizes.py. It trains with more and more merges and counts the tokens needed for the training text and for the held-out sentence.
# sizes.py: more merges mean a bigger vocabulary and shorter sequences
from bpe import train, encode
from corpus import CORPUS, HELD_OUT
print("vocab corpus tokens held-out tokens")
for n in [0, 20, 60, 150, 300, 1000]:
merges = train(CORPUS, n)
print(f"{256 + len(merges):5} {len(encode(CORPUS, merges)):14} "
f"{len(encode(HELD_OUT, merges)):16}")vocab corpus tokens held-out tokens
256 571 62
276 391 39
316 277 24
406 174 22
452 128 17
452 128 17The training text falls from 571 tokens to 128, exactly one per chunk. Training then stops after 196 merges because no pairs are left, so asking for 1,000 changes nothing. The held-out sentence gains less and less: the first 60 merges save it 38 tokens, the next 136 save 7. Late merges fit the training text closely, much as a model can fit its training data better than new data.
Longer sequences cost. The 'Do All Languages Cost the Same?' paper describes commercial language-model services charging by the number of tokens processed or generated. A model's context window is counted in tokens too: GPT-2's held 1,024. Inside a transformer, the 'Attention Is All You Need' paper gives self-attention a cost per layer that grows with the square of the sequence length, so twice the tokens means about four times that work.
A bigger vocabulary costs too. Each token needs its own row of learned numbers, its embedding, as wide as the model (its width, d_model). The embedding table holds vocabulary size times width numbers. The output layer that scores every possible next token is the same size; the transformer paper shares one matrix between them. For the smallest GPT-2, 50,257 tokens times a width of 768 is 38,597,376 numbers, about a third of its reported 117 million parameters. And a rare token gets few training examples to learn its row from.
Scroll sideways to see every column.
| Effect | Smaller vocabulary | Larger vocabulary |
|---|---|---|
| Tokens per text | More | Fewer |
| Price, context used and attention work | Higher | Lower |
| Embedding and output tables | Smaller | Larger |
| Training examples per token | More | Fewer for rare tokens |
Real tokenisers make this trade on vastly more text. GPT-2 used 50,257 tokens. PaLM's 256k-token vocabulary was generated from its training data and chosen to support many languages without excessive tokenisation. SentencePiece asks for the final vocabulary size rather than the number of merges. Our 105-word corpus shows the same trade in miniature.
The comparisons above hold other choices fixed. Different model prices, architectures and tokenisers need separate measurement. Quadratic attention describes the full-attention term, not a prediction that a complete application will take exactly four times as long. Longer vocabularies can also increase output scoring work. Vocabulary size alone cannot determine which model is cheaper or better.
Try it yourself · Activity 05
10 minPrice the table
Work these out on paper, then check each with Python as a calculator.
- Embedding-table sizes for our toy vocabularies of 256 and 452 tokens at a width of 64.
- The largest GPT-2: 50,257 tokens at width 1,600, as a share of its reported 1,542 million parameters.
- The smallest GPT-2's table in bytes, stored at 4 bytes per number.
Why does the embedding table's share shrink as models grow?
Worked answer
256 x 64 = 16,384 and 452 x 64 = 28,928 numbers. 50,257 x 1,600 = 80,411,200, about 5.2% of 1,542 million, against about 33% for the smallest model. 38,597,376 x 4 = 154,389,504 bytes, about 154 MB. From the smallest to the largest GPT-2 the table grew only as fast as the width (1,600 / 768, about 2.1 times), while the whole model grew about 13 times, partly by adding layers (48 against 12).
Read a sample · Chapter 06 of 06
Test languages, numbers, spaces and code
A count belongs to a particular tokeniser and a particular text. Measure both before using that count to explain cost or model behaviour.
Save this as probes.py beside the earlier scripts, then run python probes.py. The examples use invented text only. No model is called and no account is needed. Each assertion checks the full round trip before printing the count and the pieces.
# probes.py: measure this toy, not a commercial model
from bpe import train, vocab, encode, decode, pieces
from corpus import CORPUS
merges = train(CORPUS, 60)
table = vocab(merges)
tests = [
" library", "library", "Library", "LIBRARY",
"the books", "the books", "the books ",
"12", "21", "1200", "12.00",
"books", "livres", "libros", "\u4e66",
"x=12", "x = 12", "", "e\u0301",
]
for text in tests:
ids = encode(text, merges)
assert decode(ids, table) == text
print(f"{ascii(text):16} {len(ids):2} {pieces(ids, table)}")' library' 1 [ library]
'library' 6 [l][i][b][ra][r][y]
'Library' 6 [L][i][b][ra][r][y]
'LIBRARY' 7 [L][I][B][R][A][R][Y]
'the books' 3 [t][he][ books]
'the books' 7 [t][he][ ][ ][b][oo][ks]
'the books ' 4 [t][he][ books][ ]
'12' 2 [1][2]
'21' 2 [2][1]
'1200' 4 [1][2][0][0]
'12.00' 5 [1][2][.][0][0]
'books' 3 [b][oo][ks]
'livres' 5 [l][i][v][re][s]
'libros' 5 [l][i][b][ro][s]
'\u4e66' 3 [\xe4][\xb9][\xa6]
'x=12' 4 [x][=][1][2]
'x = 12' 6 [x][ ][=][ ][1][2]
'' 0
'e\u0301' 3 [e][\xcc][\x81]The leading space in library belongs to a frequent chunk in the training text. Removing it changes the chunk and its merges. Capitalisation also changes the bytes. Two spaces create a different pre-tokenisation boundary from one space. These are properties of this small training corpus and splitting rule, not universal counts for these words.
The number examples show another limitation: this toy groups consecutive digits before learning merges. A model tokeniser may deliberately split digits differently. PaLM and LLaMA describe digit splitting in their papers. The arithmetic study in Sources investigates effects of tokenisation on particular models and tasks; this script does not test arithmetic ability.
The multilingual examples are short illustrations, not a fair benchmark. The training text is English, and the strings differ in meaning and grammatical number. Raw counts cannot establish that one language is harder or less capable. To compare access costs, use carefully checked parallel texts, identify the exact tokeniser version and report token counts alongside bytes and task outcomes. The study by Petrov and colleagues reports large disparities for equivalent texts across languages and tokenisers; our four words cannot reproduce that study.
For an actual application, count with the model's own tokeniser and account for its chat template, special tokens, retrieved material and tool descriptions. An API usage record may include tokens that are absent from a text-only local estimate. An English words-per-token shortcut is only a rough planning aid, especially for code, numbers and other scripts.
Try it yourself · Activity 06
15 minCreate a small tokenisation test set
Use the toy to investigate a change without making a claim about every model.
- Predict the difference between
x=12andx = 12, then compare the printed pieces. - Compare the empty string with a single space. Compare precomposed
"\u00e9"with"e\u0301"(e followed by a combining acute accent). - Write a four-column record: exact input, tokeniser and corpus, observed count, and one limit of the conclusion.
Why can two texts that look alike still have different token ids?
Worked answer
The printed code expressions have 4 and 6 tokens respectively: spaces change chunk boundaries and available merges. The empty string has 0 tokens and a single space has 1. Precomposed é is 2 UTF-8 bytes and 2 tokens here; e plus the combining accent is 3 bytes and 3 tokens. This toy applies no Unicode normalisation, so both strings round-trip exactly but are not identical. A valid conclusion is that these specific inputs differ for this corpus and implementation, not that every commercial model uses these counts.
Keep the training corpus, pre-tokenisation rule and merge order together as a versioned artefact. Changing any one can change ids. A model trained with one vocabulary cannot generally use a newly trained vocabulary as a drop-in replacement. Treat this exercise as a way to understand the boundary between text and a model, then use the documented tokenizer that belongs to the model you run.
Keep learning
The complete workbook
This workbook builds a byte-pair encoding tokeniser from scratch in about sixty lines of standard-library Python, trains it on an invented text and tests that decoding gives back exactly what was encoded. You will add a special token, measure how vocabulary size trades sequence length against table size, and see, in a clearly labelled toy measurement, why other languages, numbers, spaces and code can need more tokens.
- 01Characters, words, bytes and subwordsRead here · 1 exercise
A text language model typically receives token ids that select learned embeddings. The tokeniser determines how the text becomes those ids.
- 02Byte-pair encoding by handRead here · 1 exercise
BPE builds a vocabulary by gluing together the most frequent neighbouring pair, again and again. Do it once on paper and the code will hold no surprises.
- 03Build a byte-level BPE in plain PythonRead here · 1 exercise
About sixty lines of standard-library Python train a real, if tiny, BPE tokeniser. Starting from bytes means valid Unicode text has a byte-level fallback.
- 04Encode, decode and check the round tripRead here · 1 exercise
A tokeniser must give back exactly the text it was given. Test that directly, then add a special token that typed text cannot forge.
- 05Vocabulary size: shorter texts or a bigger tableRead here · 1 exercise
More merges make texts shorter but the vocabulary bigger. Both sides cost something, so vocabulary size is a trade-off, not a setting to maximise.
- 06Test languages, numbers, spaces and codeRead here · 1 exercise
A count belongs to a particular tokeniser and a particular text. Measure both before using that count to explain cost or model behaviour.
Also inside: a 8-point checklist, a glossary of 12 terms and 10 questions and answers to test yourself. 6 hands-on exercises, each with a worked answer at the back where the workbook gives one.
No login, no card, no account. Before the download we ask you to follow Mickai (two quick links). Free to download and use for personal learning, study groups and inside your own team. Please do not resell the workbooks or republish them as your own. Link people to trust-agent.ai instead.
Test yourself
Questions and answers
What is a token?
A unit represented by a whole-number token id. In a text tokeniser it can represent a word, part of a word, whitespace, punctuation or a byte. Text is encoded as ids before the model embeds and processes it. Models can also handle non-text inputs. Text context limits and prices are often counted in tokens rather than words.
Is a token the same as a word?
No. Common words are often one token and rare words several, and spaces, punctuation and even parts of one character can be tokens. In this workbook's toy tokeniser, ' library' with its leading space is one token, while 'library' at the start of a line is six. Token counts from one tokeniser do not transfer to another.
Why start from bytes rather than characters?
Valid Unicode text can be encoded as UTF-8 bytes. A vocabulary containing all 256 byte values can represent that encoding without an unknown token. The GPT-2 paper described a character vocabulary of over 130,000 at the time. UTF-8 uses two to four bytes for a non-ASCII character. This lab deliberately rejects unpaired surrogate values, which are not valid Unicode scalar values.
What does byte-pair encoding actually do?
Training counts every pair of neighbouring symbols in the training text, merges the most frequent pair into a new symbol, writes that merge down and repeats. Encoding new text replays the merges in the order they were learned. The final vocabulary is the starting symbols plus one token per merge, so the number of merges sets the vocabulary size.
What does lossless mean for a tokeniser?
Decoding the encoding gives back exactly the original text: every space, tab, accent and emoji. Test it directly by comparing decode(encode(text)) with text on awkward inputs. Some tokenisers first normalise text, for example by the Unicode NFKC rules that SentencePiece applies by default, so the round trip then returns the normalised text.
What are special tokens, and why can they be risky?
They are ids with a job, such as marking the end of a document or padding a batch. If typed text could turn into a special token, a user could forge one. Add special ids only in your own code, treat typed text as ordinary characters and check ids, not decoded text, which looks the same either way.
How big should a vocabulary be?
It depends on the corpus and model. Holding other choices fixed, a larger vocabulary can shorten sequences but enlarges embedding and output tables and leaves rare tokens fewer examples. Savings in attention work, context space or API charges depend on the actual model and pricing. GPT-2 used 50,257 tokens; PaLM used 256k. Measure the trade-off rather than assuming bigger is always better.
Why do some languages need more tokens?
Coverage of the training corpus, the script and the tokenisation algorithm affect sequence length. This byte-level lab falls back to bytes when no learned merge applies; not every tokeniser has that fallback. Research cited in the workbook found differences of up to 15 times for the same content across languages and tested tokenisers. That is an observed study result, not a universal multiplier.
Why would a trailing space or a capital letter matter?
They change the tokens. In this workbook's toy, 'the books' is 3 tokens but 'the books' with a double space is 7, a trailing space becomes a token of its own, and 'library', 'Library' and 'LIBRARY' share no tokens with ' library'. A model can only learn that such variants are related from its training data.
Is this toy tokeniser how real models tokenise?
This lab demonstrates byte-level BPE. Production tokenisers may use BPE, unigram models or other algorithms, with different normalisation, splitting and special-token rules. Many use byte fallbacks; some do not. The original PaLM and LLaMA papers describe splitting numbers into digits. Use the exact tokeniser and version for your model when measuring real inputs.
When you have finished
Get your certificate of completion
Type your name and download a certificate for this workbook as a PDF, ready to print or to add to LinkedIn. It is made on your own device, so your name is never sent to us. It is a self-declared certificate, not an accredited qualification.
Learn the language
Key terms
- Token
- One piece of text from a tokeniser's fixed list: a whole word, part of a word, a space, a punctuation mark or a single byte.
- Token id
- The whole number that stands for a token. Text models typically use token ids to select learned embeddings and score possible output tokens.
- Vocabulary
- The complete list of tokens a tokeniser can produce, each with its id.
- Tokeniser
- The program that turns text into token ids (encoding) and ids back into text (decoding). Often spelt tokenizer.
- UTF-8
- The usual way to store text as bytes: 1 byte for each ASCII character and up to 4 bytes for other characters.
- Byte-pair encoding (BPE)
- A way to build a vocabulary by repeatedly merging the most frequent pair of neighbouring symbols into a new symbol.
6 of the workbook's 12 terms. The complete glossary is in the workbook.
Follow the evidence
Sources and checks
Facts last checked: .
Examples in this workbook were run on: Python 3.12.10 on Windows 11; original scripts and printed outputs re-run by Codex in an isolated folder. Standard library only. No model or API calls. (2026-09-26).
These workbooks use AI assistance. See how the workbooks are made.
- Neural Machine Translation of Rare Words with Subword Units (Sennrich, Haddow and Birch)arXiv, 1508.07909
- Language Models are Unsupervised Multitask Learners (the GPT-2 paper, Radford et al.)OpenAI
- SentencePiece: A simple and language independent subword tokenizer and detokenizer for Neural Text Processing (Kudo and Richardson)arXiv, 1808.06226
- Attention Is All You Need (Vaswani et al.)arXiv, 1706.03762
- PaLM: Scaling Language Modeling with Pathways (Chowdhery et al.), section 2, VocabularyarXiv, 2204.02311
- LLaMA: Open and Efficient Foundation Language Models (Touvron et al.), section 2.1, TokenizerarXiv, 2302.13971
- Language Model Tokenizers Introduce Unfairness Between Languages (Petrov, La Malfa, Torr and Bibi)arXiv, 2305.15425
- Do All Languages Cost the Same? Tokenization in the Era of Commercial Language Models (Ahia et al.)arXiv, 2305.13707
- Tokenization counts: the impact of tokenization on arithmetic in frontier LLMs (Singh and Strouse)arXiv, 2402.14903
- tiktoken source code, core.py (the encode function's note on special tokens)OpenAI on GitHub (MIT licence)
- RFC 3629: UTF-8, a transformation format of ISO 10646RFC Editor (IETF)
- re: regular expression operations (the \w character class)Python Software Foundation
- Built-in functions: max (ties return the first item encountered)Python Software Foundation
- Built-in types: str.encode, bytes.decode and dictionary insertion orderPython Software Foundation
- codecs: error handlers (the 'replace' handler and U+FFFD)Python Software Foundation
- Tokeniser lab (a browser companion to this workbook)Trust Agent, Mickai LTD
Created by Mickarle Wagstaff-Irons - Micky Irons with the Mickai team. Published by Mickai LTD. Last updated 27 September 2026.
NextKeep going
Where to go next
Recommended for you
Tensors, gradients and a first training loop
Learn what training is by doing it on a CPU: tensors and shapes, the loss, gradients by hand and by autograd, the training loop, validation and the classic bugs.
Recommended for you
Embeddings and vector search from first principles
Build vector search from scratch: turn text into vectors, compare them three ways, search by brute force, score it with recall@k and MRR, and see where it fails.
Recommended for you
GPU memory maths for LLMs and SLMs
A repeatable method for checking whether a language model fits your GPU: weights, KV cache, training memory and adapters, plus a small calculator you can run.