The unit everything in AI is priced and measured in
In brief
Tokens are how language models read and write text — and how every AI API charges you. Understanding them turns abstract pricing into something you can predict and control.
Contents
If you've ever wondered why AI APIs charge by "tokens" instead of words or characters, this is the explanation.
What a token is
A token is a chunk of text — somewhere between a character and a word. It's the atomic unit that language models work with.
English text breaks down roughly like this:
- Common short words are usually one token: "the," "is," "a," "in"
- Longer words often split into two or more tokens: "token" is one token, "tokenization" might be two or three
- Punctuation, spaces, and special characters each take tokens
- Numbers are chunked in various ways
The rule of thumb changed. Anthropic introduced a new tokenizer with Claude Opus 4.7, and on current models one token is a bit over half a word, or roughly 2.5 characters. A million tokens holds about 555,000 words. On models from before that change, the older ratio still applies: about ¾ of a word per token, or 750,000 words per million tokens.
If you are estimating costs, use the ratio for the model you are actually calling, and use the count_tokens endpoint when the number has to be right.
Why models use tokens instead of words
Words are inconsistent. "Run" and "running" are related but different strings. "Unbelievable" is one word but contains recognizable sub-units. "New York" is two words but often functions as one concept.
Tokens let the model work at a level that captures meaningful sub-units without being arbitrarily split at spaces. The tokenizer — the system that converts text to tokens before feeding it to the model — is trained to find cuts that preserve semantic meaning.
How this affects you as a builder
Pricing. Every AI API, including Anthropic's, charges per token — for input (what you send) and output (what Claude generates). Understanding token counts lets you estimate and control costs before they surprise you.
context window limits. Claude Fable 5, Opus 5 and Sonnet 5 carry a 1,000,000-token context window — combined input and output, roughly 555,000 words. Claude Haiku 4.5 carries 200,000. It is a lot, but it is finite, and long documents, conversation history and system prompts all count toward it.
Performance. Fewer input tokens means faster responses and lower latency. Verbose prompts cost more and process more slowly than concise ones.
The practical implications
A few things worth knowing for everyday use:
Code uses more tokens than prose. Programming languages have many special characters and unusual patterns that tokenize inefficiently.
Non-English text often uses more tokens per "word" than English. Languages with complex morphology or non-Latin scripts can require significantly more tokens to express the same content.
Whitespace and formatting add up. Excessive newlines, indentation, and markdown syntax all consume tokens. Clean, tight formatting uses context window more efficiently.
How to estimate your usage
For rough estimates on current models: take your word count and multiply by 1.8. On models older than Claude Opus 4.7, multiply by 1.3 instead. For precise counts before making API calls, use the count_tokens endpoint to get exact figures before committing to a request — worth doing whenever the estimate is driving a budget rather than a guess.
Understanding tokens turns the abstract "AI cost" into something predictable. Once you can estimate token counts reliably, you can design applications that are efficient by default rather than expensive by accident.
Further reading
- Pricing — current token pricing for all Claude models
- Token-saving updates on the Anthropic API — recent API changes that reduce token usage
- Prompt caching — how to reduce costs by caching repeated context