โ† Blog

Where 6.26 billion tokens went in my AI-built site redesign

Part 6: finishing and launching the redesign brought the total to 6.26 billion tokens across Codex and Claude Code, $3,691 at API prices. 97% were cached context read again. The breakdown by kind, phase, role and detour, how I nearly undercounted it by 2 billion, and what the last four days added.

Part 6 of 6 in "Rebuilding williamliu.ai with agents." Start with the overview.

TL;DR

  • ๐Ÿ” 6.26 billion tokens in total, 97% of them cached context re-read; only about 1 token in 34 was new.
  • ๐Ÿ’ต At API prices that is $3,691: $2,473 for Codex, $1,218 for Claude Code. I use both on subscriptions.
  • ๐Ÿง  The biggest single consumer was still coordination: the Codex root agent alone used 1.39 billion tokens, 30% of Codex's total.
  • ๐Ÿ›ก๏ธ Review runs were 63% of the Codex runs but 22% of its tokens; building used 48% and the root's coordination 30%.
  • ๐Ÿš€ Finishing and launching added 1.12 billion tokens, 18% of the final total.
  • ๐Ÿงพ My first count missed 2 billion Codex tokens: resumed runs restart their usage counters.

Almost all of the tokens were agents re-reading their own context, and the agent that re-read the most was the one coordinating everyone else. I counted usage across 683 Codex and Claude Code agent runs, from the build through the September 16 launch: every request of every Codex run, and each Claude message's usage counted once. Two short Codex runs recorded no usage at all, so the totals are a lower bound.

Where 6.26 billion tokens went by phase: the Codex build, 3.30 billion (53%); aligning the site with the design, 1.02 billion (16%); the deploy path and related infrastructure, 0.82 billion (13%); and the launch extension, 1.12 billion (18%).
fig 01 Where 6.26 billion tokens went by phase: the Codex build, 3.30 billion (53%); aligning the site with the design, 1.02 billion (16%); the deploy path and related infrastructure, 0.82 billion (13%); and the launch extension, 1.12 billion (18%).

๐Ÿ” 97% of the tokens were re-reads

The agents read 6.07 billion tokens of context they had already seen and wrote 19.1 million new ones, a ratio of roughly 318 to 1. Each time an agent takes a step, it re-sends its whole conversation to the model. Providers serve the repeated part from a cache at a large discount, so a cached token is cheap but not free, and it still fills the context window.

Here is the total, split into those three kinds:

The 6.26 billion tokens by kind: 6.07 billion cached re-reads (97.0%), 167 million fresh input tokens (2.7%), and 19.1 million output tokens written by the models (0.3%).
fig 02 The 6.26 billion tokens by kind: 6.07 billion cached re-reads (97.0%), 167 million fresh input tokens (2.7%), and 19.1 million output tokens written by the models (0.3%).
  • Re-read from cache: 6.07 billion (97%). Codex 4.44 billion, Claude 1.63 billion.
  • Fresh input: 167 million (3%). Input not served from cache: new files, tool output and instructions, plus anything whose cache had expired.
  • Written: 19.1 million (0.3%). Code, reviews, plans and messages. About 6.66 million of Codex's 16.18 million were hidden reasoning.

Only about 1 token in 34 was new. That shapes everything else in this part: repeated context, not output, is the larger token pool to optimize. These logs don't show which session strategy would shrink it.

๐Ÿง  The coordinator used 30% of Codex's tokens

The single largest consumer in the whole project was still the Codex root agent, which used 1.39 billion tokens coordinating the build. It planned the milestones, coordinated a tree of 332 agents, read their reports, integrated their work and answered me. Every one of those steps re-sent its conversation, and the conversation only grew. Across the whole effort, Codex tokens split into building (48%), the root's coordination (30%) and review (22%).

Codex tokens across the whole effort by role, with tokens per run: implementation, repair and research 1.49 billion (8.5 million per run); the goal root 1.39 billion; implementation runs launched by Claude Code 695 million (11.6 million per run); Finders, Adversaries and Referees 790 million (2.3 million per run); review runs launched by Claude Code 225 million (4.8 million per run).
fig 03 Codex tokens across the whole effort by role, with tokens per run: implementation, repair and research 1.49 billion (8.5 million per run); the goal root 1.39 billion; implementation runs launched by Claude Code 695 million (11.6 million per run); Finders, Adversaries and Referees 790 million (2.3 million per run); review runs launched by Claude Code 225 million (4.8 million per run).
  • Building: 48%. Implementation, repair and research runs used 1.49 billion tokens; implementation runs launched by Claude Code used another 695 million.
  • Coordination: 30%. The root agent, 1.39 billion.
  • Review: 22%. Finders, Adversaries and Referees used 790 million tokens; the review and audit runs Claude Code launched used another 225 million.

The per-run numbers still point the same way. Implementation, repair and research runs averaged 8.5 million tokens. Implementation runs launched by Claude Code averaged 11.6 million; named reviewers averaged 2.3 million, and Claude-launched review and audit runs averaged 4.8 million. A long-lived coordinator and cold-started workers both pay to re-read; they just pay in different places.

๐Ÿงฑ Just over half went to the build; launching took another 18%

The build took 53% of the final total, aligning the site with its design 16%, the deploy path and related infrastructure 13%, and finishing and launching 18%.

  • Build, 3.30 billion (53%). The Codex goal session: spec, plan, sixteen milestones, 333 agents.
  • Alignment, 1.02 billion (16%). Staging the redesign, the design audit and twelve fix rounds: 0.66 billion Codex and 0.36 billion Claude.
  • Deploy path and infrastructure, 0.82 billion (13%). Memory debugging, the clean-checkout deploy, deploy timing, and a deferred plan to restructure how the site stores media: 0.20 billion Codex and 0.62 billion Claude.
  • Launch extension, 1.12 billion (18%). Finishing the blog support, correcting the staged pages, checking the series, planning the next search work and launching the redesign.

The original three phases did not change; the denominator did. The 1.12 billion-token extension is what took the work from the September 12 draft through launch. Part 5 covers what went wrong on the deploy path.

๐Ÿ’ต The dollars: $3,691 at API prices

At published API rates, the whole effort would have cost $3,691: $2,473 for Codex and $1,218 for Claude Code. I use both on subscriptions, so read this as a price comparison. Codex is priced at OpenAI's standard rates for each model, request by request; Claude Code at Anthropic's, with cache reads and writes at their own rates.

API-equivalent cost by area and tool: the strict redesign $2,303 ($2,048 Codex, $256 Claude Code); the prelaunch deploy path and related infrastructure $561 ($114 Codex, $447 Claude Code); the launch extension $827 ($311 Codex, $516 Claude Code).
fig 04 API-equivalent cost by area and tool: the strict redesign $2,303 ($2,048 Codex, $256 Claude Code); the prelaunch deploy path and related infrastructure $561 ($114 Codex, $447 Claude Code); the launch extension $827 ($311 Codex, $516 Claude Code).
  • The strict redesign: $2,303. Codex $2,048, most of it GPT-5.6 Sol, and Claude Code $256.
  • The prelaunch deploy path and related infrastructure: $561. Codex $114 and Claude Code $447. The Claude share breaks down as $277 for a day of memory debugging and a deferred media-storage plan, $138 for the clean-checkout deploy fix that finally worked, and $31 for other deploy sessions and security reviews of deploy changes.
  • The launch extension: $827. Codex added $311; Claude Code added $516.
  • Codex by model: GPT-5.6 Sol $1,823, GPT-6 Astra $367, and $283 across five cheaper models.
  • Codex by role: the root agent $807, implementation, repair and research $1,010 including Claude-launched implementation, and review $656 including Claude-launched review and audit runs.

Parts are rounded, so they can differ from the totals by a dollar.

๐Ÿš€ What the last four days cost

From September 13 through 16, the agents used 848 million tokens, or $641 at API prices: $347 on September 13, $187 on September 14, $35 on September 15 and $72 on launch day. The work across those days included the inline-video feature and its 15 Codex review rounds, fixes after the real blog pages missed the design, the Writing-to-Blog rename, the AEO/SEO audit and plan review, and the final merge, checks and launch. The extension ledger also includes a late-September-12 fact-checking bucket that was absent from the snapshot; the aggregate ledger is the source of the full 1.12 billion-token, $827 extension.

Claude's share grew because it stayed in the orchestration seat: writing briefs, integrating and reviewing Codex work, and carrying the plan review and launch. Of the extension's API-equivalent cost, Claude Code added $516 while Codex added $311. That moved Claude from $702 of the first $2,864 to $1,218 of the final $3,691.

๐Ÿงพ How I nearly undercounted by 2 billion tokens

My first count said Codex used 2.17 billion tokens; the real figure is 4.16 billion, because a Codex run's usage counter starts over when a paused run is resumed. Each Codex log records a running total. The obvious method is to read the last one. But when I paused the build and resumed it, the counter restarted. 76 of the 514 runs had restarted counters, and the root agent restarted three times, which hid 1.23 billion of its 1.39 billion tokens.

Codex tokens counted two ways: 2.17 billion reading each run's final counter, 4.16 billion summing every request, a difference of about 2 billion tokens across 76 runs whose counters restarted.
fig 05 Codex tokens counted two ways: 2.17 billion reading each run's final counter, 4.16 billion summing every request, a difference of about 2 billion tokens across 76 runs whose counters restarted.

The draft of this series went through four rounds of independent fact-checking with the wrong number, because the checker used the same method. It came out when I priced Codex request by request and the sum didn't match. Two ways of counting, per-request usage and the peaks before each restart plus the final total, now agree on the same 4,157,346,610 tokens. If you measure your own agents, sum per-request usage, and check it against a second method.

๐Ÿ—‘๏ธ What I'd call waste

Three detours produced work that was later reverted or deferred: a player that never matched the design, a plan for the wrong week, and review rounds past the point of value.

  • The custom audio player: at least 114 million Codex tokens, about $28 at API prices. I liked the design's player, but the one built was far from it. Fifteen agents worked on it by name, Finders 4 through 12, repairs 6 through 9, and a final Adversary and Referee, for 86 million tokens; the run that took it out used 28 million more. The earlier rounds aren't in that count, and neither are the days of review.
  • The deferred media plan: part of that API-equivalent $277 day. Research, a spec and a plan, each through several rounds of Codex review, for an overhaul I then postponed so the deploy could be fixed first.
  • Review past the point of value. Until I changed the rule, medium findings also had to be fixed or explicitly waived before a milestone could move. After the change, only critical and high findings blocked, and the rest went to a backlog with a priority. I can't separate those tokens cleanly, which is itself a lesson: tag review rounds by severity and you can.

That severity rule put Encoded Judgment Becomes the New Moat into its smallest useful form: a blocking threshold left my head and changed every later review.

Not everything that looks like overhead was waste. The same review loop found a test helper that would have printed a real credential with debug logging on. The audit rounds took the site from 20 deviations to the last one fixed. I'd pay those again.

๐Ÿ“ About 39 million tokens per message I typed, on average

The project recorded 38.6 million tokens for each of my 162 messages. That is an average across the whole effort, not the cost of any one message, but it shows how much agent work a few words set in motion. Autonomous runs, reviews and deploys all happened between my messages. A precise brief, like the design reference, files in scope and tests to run that Claude Code wrote for each fix round, is where those few words matter most.

Tokens per unit: 38.6 million per human message, 67.3 million per commit, 9.2 million per agent run.
fig 06 Tokens per unit: 38.6 million per human message, 67.3 million per commit, 9.2 million per agent run.

๐Ÿงญ What I'd change next time

Watch the coordinator, stop review earlier, prove the deploy path first, and count tokens two ways.

  • Watch the coordinator's context. One long-lived root agent used 30% of the Codex tokens. Give it a way to hand work off without carrying every report in its own conversation.
  • Prove the deploy and launch path first. The prelaunch deploy path cost $561; finishing and launching added another $827.
  • Write the severity rule on day one. Blocking only on critical and high findings is the cheapest review policy that still catches what matters.
  • Look at components early, next to the design. Most of the player's fourteen review rounds checked behavior; only the last compared it with the design.
  • Count tokens per request, and cross-check. A running total that restarts will quietly halve your numbers.
  • Tag tokens by phase and purpose as you go. I reconstructed all of this afterwards. Next time the logs should say which milestone and which detour each run belonged to.

That's the series. The redesign is now the live site, and the reader-facing tour shows what changed.

What share of your agents' tokens goes to coordination rather than work, and how would you know?

How this post was made: drafted by Claude Code from my own session logs, checked against those logs by a separate Codex agent, and edited by me. The cover illustration is AI-generated; the charts are drawn from the numbers above.