Opus 4.7 uses 58% more input tokens than GPT-5.5
A controlled comparison of identical fixtures reports higher input-token counts for Opus 4.7, separating input accounting from output efficiency and task quality.
TL;DR
-
Opus 4.7 used about 58–59% more input tokens than GPT-5.5 on the same fixtures.
-
My earlier Opus 4.7 article asked whether Opus quality justifies higher token cost.
-
Public comparisons now put GPT-5.5 quality broadly on par with Opus 4.7.
-
That makes the trade-off concrete: Opus carries a visible input-token tax.
-
With quality on par, OpenAI’s input-efficiency advantage becomes hard to ignore.
-
This article focuses on input-token accounting; output-token efficiency is a separate comparison.
-
You can reproduce the benchmark with token-count-compare.
🚨 The part people are missing
My controlled 21-fixture comparison between Claude Opus 4.7 and GPT-5.5 did not surprise me by picking a winner. It confirmed the same trade-off I raised in the earlier article: Opus 4.7 reported about 58–59% more input tokens than GPT-5.5 across the suite.
There are separate discussions about output-token efficiency; this article is deliberately narrower. It focuses on the input side: how the same task material counts before the model starts generating. If you want the broader output/pricing angle, see GPT 5.5 vs. Claude Opus 4.7: Benchmarks, Pricing, and What Actually Matters.
The quality premise matters. MindStudio’s coding comparison says both models are “close enough” on most benchmark tasks that there is no clean overall winner, while Mashable’s benchmark roundup says Opus 4.7 has an edge in advanced and agentic coding but GPT-5.5 performs better on most benchmarks (MindStudio, Mashable). That is enough for my conclusion: with quality broadly on par, input-token efficiency becomes a product decision, not a trivia point.
🛠️ Try the benchmark yourself
I used token-count-compare to run the comparison. The project includes the harness and fixtures, so you can swap in your own tasks and see how they count.
📊 The comparison at a glance
-
Full benchmark prompt totals: Opus 21,374 vs GPT-5.5 13,504 (+7,870 tokens, +58.28% for Opus).
-
Raw text totals: Opus 18,644 vs GPT-5.5 11,749 (+6,895 tokens, +58.69% for Opus).
-
Plain prose sample: Opus 1,519 vs GPT-5.5 993 (+526 tokens, +52.97% for Opus).
-
Long-form sample: Opus 2,290 vs GPT-5.5 1,386 (+904 tokens, +65.22% for Opus).
-
Code-heavy comments sample: Opus 428 vs GPT-5.5 240 (+188 tokens, +78.33% for Opus).
-
Java sample: Opus 1,305 vs GPT-5.5 802 (+503 tokens, +62.72% for Opus).
That does not prove identical tokenizers or a perfect apples-to-apples segmentation story. It does show something more practical: on the input side, Opus 4.7 is materially more token-hungry.
🔍 What the public docs actually say
Anthropic’s Claude Opus 4.7 announcement says the model ships with an updated tokenizer. Simon Willison measured the effect directly and found 1.46x on a text-heavy system prompt and 1.08x on a text-heavy PDF (Claude Token Counter, now with model comparisons).
OpenAI’s GPT-5.5 latest-model guide tells a different story. It says GPT-5.5 reaches strong results with fewer reasoning tokens than prior models, even at the same effort level. OpenAI’s launch post adds that GPT-5.5 uses significantly fewer tokens to complete the same Codex tasks, while matching GPT-5.4 per-token latency in real-world serving.
For the earlier tokenizer-focused piece, I also published Tokenizer Inflation in Opus 4.7 Is Real, but So Are the Quality Trade-Offs.
That earlier question was whether Anthropic’s more granular tokenizer was worth the trade-off. This test makes the answer look a lot less ambiguous because GPT-5.5 quality is on par for the job.
That is the key split:
-
Anthropic is making the input accounting more explicit, and sometimes more expensive.
-
OpenAI is making the task execution more efficient.
Those are related, but not the same.
⚙️ What the experts are noticing
Simon Willison’s take was practical, not tribal: if a task is well-specified and he wants the right answer, he leans GPT-5.5; if he wants a conversation or something “Claude Code-shaped,” he leans Opus 4.7.
Zvi Mowshowitz made a similar point in his GPT-5.5 commentary: GPT-5.5 looks like the better choice for “just the facts” or straightforward requests, while Opus 4.7 still feels stronger for open-ended or interpretive work.
That’s the part I’d emphasize most: serious users are not asking “which brand won?” They’re asking “which model is the better fit for this task shape — and what did it cost to get there?”
🧪 What my tests changed in my own thinking
My own benchmark suite wasn’t public commentary; it was a controlled comparison built to see how the models behaved on the same fixtures. The result made me less interested in tokenizer theater and more interested in the end-to-end workflow.
If Opus 4.7 needs materially more input tokens to do the same work in a controlled suite, then the migration question changes.
Not: “Which model has the cleverer tokenizer?”
But: “Which model gives me the best total task economics for my actual workload?”
That matters because prompt cost, retry cost, and context headroom all compound.
A model can be better on capability and still be worse for your product if it bloats prompt consumption. A model can also be more efficient without being the best fit for messy, open-ended work.
🧗 The takeaway
I don’t think Anthropic’s path is wrong.
I do think OpenAI has shown another credible path: keep the input side much more efficient while staying competitive on quality. Because GPT-5.5 quality is on par for the job, that is a real product win.
For builders, that is the real lesson.
The winning model is not the one that makes the prettiest tokenizer story. It is the one that gets more useful work done per dollar, per retry, and per context window.
That’s the metric I trust now.
🧪 Appendix: how I ran the experiment
I ran the github.com/williamliu-ai/token-count-compare project against 21 controlled synthetic fixtures, using only the project-local .env file.
The raw-text comparison below uses provider-side token-count endpoints on the fixture text only — no generation request, no benchmark wrapper. That makes it the cleanest apples-to-apples read on the accounting difference.
-
21 fixtures in one flat
fixtures/folder -
Claude Opus 4.7 vs GPT-5.5 on the same inputs
-
benchmark usage verification: 42 pass / 0 fail / 0 skip at ±0 token tolerance
-
Provider-side raw-text counts were verified across the same fixtures

This is the section I trust most: raw counts, same fixtures, repeatable harness, and a repo anyone can inspect.