Train the tokenizer
Trained on a mix of two built-in texts: a few pages of everyday English prose written for this page, and a Chinese translation of it. Both are shown in full below.
Compare a sentence
How it works
The tokenizer starts by treating every UTF-8 byte as its own token: 256 possible tokens, one per byte value. Training then repeatedly looks at the built-in corpus, finds whichever two adjacent tokens appear next to each other most often, and merges them into one new token. Each merge shrinks the corpus a little and grows the vocabulary by one. Do that a few dozen or a few hundred times and common patterns — English words like "the" or "ing", for instance — end up as a single token instead of several bytes.
Which patterns get merged depends entirely on the training text. A tokenizer trained only on English learns English-shaped merges and has nothing useful to offer a Chinese sentence, because the two languages barely share any byte patterns — that is what the corpus mix slider controls: how much of the training text is English versus Chinese.
Common Chinese characters take 3 bytes each in UTF-8, versus 1 byte per common English letter. Move the corpus mix toward Chinese and add a few merges, and Chinese sentences start collapsing into fewer tokens the same way English ones do — move it toward English instead, and Chinese stays close to 3 tokens per character while English keeps shrinking. A chip that is not valid UTF-8 on its own — one byte out of a 3-byte character, say — shows as hex bytes rather than a broken character, since browsers do not have a glyph for "a third of a character."
This is a toy tokenizer trained on a few pages of text, not a real production tokenizer — it is only meant to show the shape of the effect, not to match any specific model's numbers.