Videos · Transformers & LLMs, From Scratch

How LLMs Read Text: Tokenization Explained Visually (BPE, Special Tokens)

Deep dive · 6:26 · AI & ML · Watch on YouTube

ChatGPT never reads a single word. Before any language model can think, your text is cut into tokens, and every token becomes a number. This video builds tokenization from scratch, with real tokenizer outputs at every step.

You'll see

  • why neural networks need tokens (and what a vocabulary is)
  • why splitting into characters makes sequences too long
  • why splitting into words makes the vocabulary explode, with typos becoming an unknown (UNK) token
  • byte pair encoding (BPE), step by step on a real training run
  • how learned merges tokenize words the model never saw, and how byte fallback handles emoji
  • what GPT-4o's real tokenizer does with spaces, long words, numbers, code, and Azerbaijani vs English
  • special tokens: end of text, and the chat format of OpenAI's open model gpt-oss
  • why strawberry is a single token, and why that makes counting its r's hard

Sources & notes

Every token on screen is real: GPT-4o (o200k_base), GPT-4 (cl100k_base) and gpt-oss (o200k_harmony) tokenizer outputs via tiktoken, and the BPE walkthrough is a real run of the algorithm. Animations are made with Python and Manim.

📺 Episode 1 of "Transformers & LLMs, From Scratch". Next: embeddings, how token IDs become meaning.

tokenizationtokenizertokensLLMlarge language modelstransformersbyte pair encodingBPEsubword tokenizationWordPieceSentencePiecetiktoken