Videos · Transformers & LLMs, From Scratch
How LLMs Read Text: Tokenization Explained Visually (BPE, Special Tokens)
Deep dive · 6:26 · AI & ML · Watch on YouTube
ChatGPT never reads a single word. Before any language model can think, your text is cut into tokens, and every token becomes a number. This video builds tokenization from scratch, with real tokenizer outputs at every step.
You'll see
- why neural networks need tokens (and what a vocabulary is)
- why splitting into characters makes sequences too long
- why splitting into words makes the vocabulary explode, with typos becoming an unknown (UNK) token
- byte pair encoding (BPE), step by step on a real training run
- how learned merges tokenize words the model never saw, and how byte fallback handles emoji
- what GPT-4o's real tokenizer does with spaces, long words, numbers, code, and Azerbaijani vs English
- special tokens: end of text, and the chat format of OpenAI's open model gpt-oss
- why strawberry is a single token, and why that makes counting its r's hard
Sources & notes
Every token on screen is real: GPT-4o (o200k_base), GPT-4 (cl100k_base) and gpt-oss (o200k_harmony) tokenizer outputs via tiktoken, and the BPE walkthrough is a real run of the algorithm. Animations are made with Python and Manim.
📺 Episode 1 of "Transformers & LLMs, From Scratch". Next: embeddings, how token IDs become meaning.