← Blog

Anthropic is on a fast lane to token bankruptcy

Richer HTML artifacts can increase generation and revision costs, making measured token accounting essential before treating more compute as the solution.

TL;DR

  • ⚠️ Anthropic’s HTML-output push makes the token problem worse, not smaller.

  • 📊 Opus 4.7 already used 58–59% more input tokens than GPT-5.5 in my fixture test.

  • 📉 At 1.35x token inflation, prefill compute rises ~82% and raw-text context drops ~26%.

  • 🔁 HTML compounds the bill: generate more tokens, then resend them for every revision.

  • 🧨 Leasing xAI compute is a capacity patch, not a token-economics strategy.

  • 🎯 Anthropic needs to fix the cost architecture before Claude Code scales into a margin trap.

⚠️ The HTML argument is attractive — and dangerous

Thariq Shihipar’s piece on “The Unreasonable Effectiveness of HTML” makes a genuinely good product point: Claude Code can produce artifacts that are easier to read, navigate, and act on when they are HTML instead of Markdown.

The examples are compelling. The companion gallery shows twenty self-contained HTML artifacts across planning, code review, design, prototyping, diagrams, reports, research, and custom editors.

I agree with the user-experience argument.

But I think the economics are being under-discussed.

HTML is not just “a better document.” It is also more surface area: tags, CSS, JavaScript, SVG, component scaffolding, export buttons, layout code, and interaction state. Every one of those tokens has to be generated. Then, if you want to change the artifact, many workflows send the artifact back into the model and pay again.

That is the hidden loop:

  • generate a richer artifact,

  • review the richer artifact,

  • ask for changes,

  • resend a large chunk of the artifact,

  • generate another large artifact,

  • repeat.

For a model line already fighting token inflation, that is not a small tax. It is compounding interest.

📊 The numbers from my earlier tests already point in the wrong direction

In my earlier article, “Opus 4.7 uses 58% more input tokens than GPT-5.5”, I ran a controlled 21-fixture comparison using the same inputs across Claude Opus 4.7 and GPT-5.5.

The headline result was blunt:

  • Full benchmark prompt totals: Opus 21,374 vs GPT-5.5 13,504.

  • Difference: +7,870 input tokens for Opus.

  • Relative tax: +58.28% for Opus.

  • Raw-text totals: Opus 18,644 vs GPT-5.5 11,749.

  • Raw-text tax: +58.69% for Opus.

The pattern held across normal text and code-heavy material:

  • Plain prose: Opus 1,519 vs GPT-5.5 993 (+52.97%).

  • Long-form sample: Opus 2,290 vs GPT-5.5 1,386 (+65.22%).

  • Code-heavy comments: Opus 428 vs GPT-5.5 240 (+78.33%).

  • Java sample: Opus 1,305 vs GPT-5.5 802 (+62.72%).

The benchmark is reproducible in the public token-count-compare repo.

This matters because HTML artifacts are often closer to code-heavy or mixed-structure documents than plain prose. They are punctuation-heavy. They contain nested syntax. They often include CSS and JavaScript. They are exactly the kind of artifact where token accounting and context headroom become product constraints.

📉 The earlier Opus 4.7 tokenizer piece was the warning shot

In “Tokenizer Inflation in Opus 4.7 Is Real, but So Are the Quality Trade-Offs”, I framed Opus 4.7 as a serving-system migration, not just a model upgrade.

The concrete math was already uncomfortable:

  • Opus 4.7 can retokenize the same input to roughly 1.0x–1.35x tokens.

  • At 1.35x inflation, prefill compute rises by about 82% because attention work scales approximately with sequence length squared.

  • KV-cache memory rises about 35%.

  • Effective raw-text context drops about 26%.

That was before adding a recommendation that agents should produce richer HTML artifacts by default.

Now combine the two ideas:

  1. Opus already carries a visible input-token tax.

  2. HTML increases artifact size relative to Markdown or prose.

  3. Revision workflows often re-ingest the artifact.

  4. Agentic coding loops already carry tool calls, retries, diffs, logs, and memory.

That is how you get token bankruptcy: not from one bad prompt, but from a thousand “better UX” decisions that each expand the loop.

🔁 The feedback loop problem is the missing section

The biggest missing practical question in the HTML argument is not “Can Claude produce beautiful HTML?”

It can.

The question is: how do humans efficiently give feedback without paying the full artifact tax over and over?

With Markdown, feedback is usually cheap:

Change the second section. Cut the intro. Move this bullet up. Rewrite the conclusion.

With HTML, feedback can become structurally heavier:

  • The user may need to refer to visual regions that are not cleanly represented in text.

  • The model may need the full DOM or large sections of HTML to make safe edits.

  • CSS and layout changes can touch distant parts of the file.

  • Interactive elements need JavaScript state and behavior, not just prose.

  • Small visual changes can require large regenerated outputs.

The right answer may be “use HTML as the rendered artifact, but maintain a smaller source-of-truth representation underneath.”

For example:

  • keep semantic content in Markdown or JSON,

  • generate HTML as a disposable view,

  • accept patch-style edits instead of full-file rewrites,

  • preserve stable element IDs for targeted feedback,

  • ask the model to output diffs, not entire documents,

  • separate content changes from layout changes.

Without that architecture, HTML turns every feedback cycle into a larger token invoice.

🧨 Leasing xAI compute does not fix bad token economics

This is why the recent xAI/Anthropic compute news matters.

Simon Willison noted that Anthropic struck a deal with SpaceX/xAI to use “all of the capacity of their Colossus data center” in his notes on the xAI/Anthropic data center deal. Tom’s Hardware framed the same story as SpaceX renting access to a supercomputer with 220,000 Nvidia GPUs and 300 megawatts of AI compute power.

More compute helps capacity.

It does not automatically fix unit economics.

If the product pattern encourages users to generate bigger artifacts, send bigger artifacts back for feedback, run longer agent loops, and tolerate higher tokenizer overhead, then rented compute becomes a pressure valve. It is not a cure.

The uncomfortable possibility is that Anthropic is solving demand with supply while leaving the token architecture too expensive.

That is not sustainable at Claude Code scale.

🎯 What Anthropic should fix ASAP

I do not think the answer is “do not use HTML.”

HTML is useful. The UX upside is real.

The answer is that Anthropic needs to make rich artifacts token-efficient by design.

But the fix list should be evidence-based. I would narrow it to three defensible asks:

  • Publish artifact-token accounting. Show the token cost of the same Claude Code task as Markdown, plain HTML, styled HTML, and interactive HTML. Then show the cost over two or three revision passes. Without that comparison, “HTML is better UX” is only half the product story.

  • Reduce the Opus token tax on markup- and code-like inputs. This is the part my prior benchmark already supports: Opus 4.7 used 58–59% more input tokens than GPT-5.5 across the same fixtures, with code-heavy comments at +78.33%.

  • Prefer patch-style revisions only if the artifact benchmark proves full rewrites dominate cost. I should not assume Claude Code is always re-ingesting or regenerating whole files. The right ask is simpler: measure the loop, then optimize the expensive part.

I would remove the more speculative prescriptions — DOM-aware compression, stable anchors, canonical source models, and a special HTML revision protocol — until there is direct evidence that those are the actual bottlenecks.

If Anthropic wants Claude Code to be the agentic coding interface, this is not a polish issue. It is core infrastructure.

The winning coding agent will not be the one that makes the prettiest artifact once.

It will be the one that keeps the human in the loop without bankrupting the token budget every time the human says, “change that part.”

🧭 My view

Thariq is right that HTML can make Claude Code feel dramatically more useful.

But if Anthropic pushes HTML artifacts without solving revision economics, they are accelerating into the exact cost problem Opus 4.7 already exposed.

The product experience is moving toward richer, more visual, more interactive agent outputs.

The unanswered question is whether the revision loop can stay cheap as those artifacts get bigger.

That is the problem Anthropic should measure publicly.

Fix it now — before “beautiful Claude artifacts” become another way to hide a massive token tax.