Videos · Transformers & LLMs, From Scratch

ChatGPT Can't Read Words. It Sees This.

Short · 1:38 · AI & ML · Watch on YouTube

ChatGPT has never read a single word. Here's what it actually sees.

Before a language model can do anything, text is cut into tokens, and every token becomes a number from a fixed vocabulary. Cutting by letters makes sequences too long. Cutting by words makes the vocabulary explode, and typos become unknown. Real models use byte pair encoding: start from single letters and keep merging the most frequent pair, until common words are one token and new words split into familiar pieces.

Sources & notes

Every token in this video is real: GPT-4o and gpt-oss tokenizer outputs, plus a real BPE run.

tokenizationtokensLLMlarge language modelsChatGPTGPT-4otransformersbyte pair encodingBPEsubword tokenizationtiktokenvocabulary