Comparison · verified August 21, 2026
Claude vs GPT-5.6 for Writing
Writing is the hardest thing to compare honestly, because quality here is taste and taste is not a benchmark. Anyone telling you one model is definitively the better writer is telling you about their own ear.
What can be stated plainly: the output ceilings, the context available for long documents, how each one lets you encode a house style, and what a heavy writing workload costs. Then a blind test, because on this particular question a twenty-minute experiment genuinely does beat any amount of reading.
The verified differences
What each one costs
Last verified August 21, 2026
USD per million tokens, standard rates, no discounts applied. Read the three caveats underneath before you put these numbers in a spreadsheet.
| Model | Tier | Input | Output | Notes |
|---|---|---|---|---|
| Claude Fable 5 | frontier | $10 | $50 | The frontier model. Twice the price of Opus 5 for the hardest work. |
| Claude Opus 5 | frontier | $5 | $25 | The default flagship. Thinking on by default. Fast mode available at $10/$50. |
| Claude Sonnet 5 | balanced | $2 | $10 | What most production agent fleets run on. The September 2026 increase to $3/$15 was cancelled. |
| Claude Haiku 4.5 | fast | $1 | $5 | High-volume and latency-sensitive work. |
| GPT-5.6 Sol | frontier | $5 | $30 | The flagship tier. Long-context requests meter at $10/$45. |
| GPT-5.6 Terra | balanced | $2 | $12 | The everyday tier, cut to this rate on July 30, 2026. Long-context meters at $4/$18. |
| GPT-5.6 Luna | fast | $0.20 | $1.2 | The cheapest frontier-family option on the market. Long-context meters at $0.40/$1.80. |
Anthropic changed tokenizers at Claude 4.7
Claude 4.7 and later — including Opus 5, Sonnet 5, and Fable 5 — use a newer tokenizer that produces roughly 30% more tokens for the same text than Sonnet 4.6 and earlier. A $2/MTok model on the new tokenizer is not directly comparable to a $2/MTok model on an older one, or to another vendor. Compare cost per task, not cost per token.
OpenAI meters long context separately
GPT-5.6 publishes a second, higher rate for long-context requests: Sol goes from $5/$30 to $10/$45, Terra from $2/$12 to $4/$18, Luna from $0.20/$1.20 to $0.40/$1.80. Anthropic includes the full 1M-token window at standard pricing on Claude 4.6 and later — a 900k-token request costs the same per token as a 9k one. If your workload is context-heavy, that difference is larger than the headline gap.
Two of Google’s current rates are introductory
Gemini 3.7 Flash and 3.6 Flash are priced at $0.75/$3.75 only through December 31, 2026. Both double to $1.50/$7.50 in 2027. If you are building a twelve-month cost model on Gemini Flash, model the 2027 number, not the one on the page today.
Caching and batching move the number more than model choice
A Claude cache hit costs 10% of the standard input price, and the Batch API takes 50% off both directions; the two stack. A workload that reuses a large system prompt can land below a nominally cheaper model that you are calling uncached. Model selection is usually the third-largest lever, after caching and after not sending the context at all.
What the evidence says, and who is saying it
Almost every performance number in public circulation traces back to a vendor running its own harness. That does not make the numbers useless, but the attribution belongs next to the claim.
What writers consistently report
- Claude is described as more restrained by default — less inclined to open with a summary of the question or close with an offer to help further.
- GPT is described as more eager to produce structure: headers, bullets, and framing scaffolds even when the prompt did not ask for them.
- Both are reported to converge substantially once given a strong style guide and a few examples.
This is aggregated practitioner sentiment, not measurement, and it is the kind of claim that ages badly across model versions. It is here because it is the most common thing people ask, not because it is evidence. That third point is the one to take seriously: the prompt matters more than the model.
Run a blind test, because taste is the whole question
Twenty minutes settles this better than any comparison page, including this one. The critical detail is blind — knowing which model produced which draft contaminates the judgement completely, and people who skip this step reliably pick the vendor they already preferred.
- Take five pieces you have already written and were happy with. Real work, not test prompts. Your own published writing is the only benchmark that encodes your taste.
- Write one brief per piece and give both models the identical brief. Include your style guide if you have one. If you do not have one, write it first — it will improve both outputs more than switching models would.
- Strip the labels before you read. Have someone else shuffle them, or paste them into a document with the sources removed. This step is the entire experiment.
- Score on edit distance, not first impression. How much would you have to change before you would publish it under your name? A draft that reads well but needs a full rewrite of the argument is worse than a plain one that is structurally right.
- Test the second turn as well as the first. Most writing work is revision. "Cut this by a third and lose the hedging" is a more revealing prompt than the initial brief, and the models differ more on it.
Claude for writing and editing covers the prompting patterns, and building Skills for your team covers packaging a house style so everyone gets the same voice.
The honest bottom line
Pick Claude if you work on long documents in one piece rather than in fragments, if you want a style guide encoded once and applied automatically across a team, or if you edit in Word and want tracked changes rather than copy-paste. The 1M-token window at flat pricing means an entire manuscript plus your style guide fits in context, which changes what is possible on a book-length edit.
Pick GPT-5.6 if your writing workflow already lives in the OpenAI ecosystem, or if you write in short high-volume bursts where Luna’s pricing makes a difference at scale. On a single 800-word draft with a good brief, you will struggle to reliably pick the winner blind — which is itself the finding.
Neither, first if you have not written down your style guide. The gap between a briefed model and an unbriefed one is far larger than the gap between these two vendors, and it costs nothing to close. Switching models to fix a prompt problem is the most common wasted migration in this category.