◈AI Codex
Foundation Models & LLMsDevelopers

Tokenization

The step where AI breaks your text into small pieces — called tokens — before it can process anything. On current Claude models a token is a bit over half a word, following the tokenizer change introduced with Claude Opus 4.7; on older models it is about three-quarters of a word. This matters practically because API costs and usage limits are measured in tokens, not words or characters. The more text you send and receive, the more tokens you use.

◎

In practice

Before Claude can process your message, it gets broken into tokens. Code tokenizes less efficiently than prose because of its punctuation and unusual patterns, and non-English text often needs more tokens per word than English. Practically: a 10,000-word document is roughly 18,000 tokens on a current model, which affects both cost and how much room is left in the context window.

Related concepts