683 agent runs, 6.26 billion tokens, 20 corrections: my site redesign
Part 1 of a six-part series on rebuilding williamliu.ai with Codex and Claude Code: 683 agent runs, 6.26 billion tokens, 20 human corrections, and the five detours that cost the most.
Part 1 of 6 in "Rebuilding williamliu.ai with agents." The reader-facing tour of the new site is here.
TL;DR
- ๐งฎ Fifteen days, 683 agent runs, 6.26 billion tokens (4.59 billion Codex, 1.66 billion Claude Code), 162 messages from me.
- ๐ 97% of those tokens were cached context read again; the agents wrote about 19 million tokens in total.
- ๐ก๏ธ More than half of the Codex runs were dedicated reviewers (347 of 630), yet they used only 17% of Codex's tokens.
- ๐งโโ๏ธ By my count I corrected the agents 20 times: 5 times on the product or how the site looked, 15 times on how they were working.
- ๐ต At API prices these tokens come to $3,691 ($2,473 Codex, $1,218 Claude Code); I use both on subscriptions.
- ๐ฅ The redesign stayed off production until it was ready, yet a staging mistake stopped production database writes for about two hours.
I redesigned my website in Claude Design and let AI coding agents build it, and the numbers show where agents are strong and where they still need a person. The site is live with a new editorial design, a Chinese version of every page, and 139 more pages than before. It took two agent stacks, fifteen days (September 2 to 16), and more review runs than build runs. This series is my accounting of it: every number below comes from the session logs on my machine, and a separate agent checked them against those logs before publication.
This is the overview. Five deep dives follow, each on one topic, not one day.

๐งฎ The scorecard, in plain numbers
A small personal site took 683 agent runs and 6.26 billion tokens, and nearly all of those tokens were the agents re-reading context they had already seen. Here is what each number means.
- 683 agent runs. One run is one agent session with its own context: a builder, a reviewer, an automatic security check. Codex accounts for 630 of them, Claude Code for 53.
- 6.26 billion tokens, 97% cached. Each time an agent takes a step, it re-sends its whole conversation. The provider serves most of it from cache at a steep discount. Across the whole effort the agents wrote about 19 million tokens and re-read about 6.07 billion, roughly 318 tokens read for every token written. Two short Codex runs recorded no usage at all, so every total here is a lower bound.
- Codex did nearly three quarters of the work by tokens: 4.59 billion, against 1.66 billion for Claude Code. Codex ran the build and most of the implementation; Claude Code orchestrated the second phase and did most of the deploy and launch work.
- $3,691 at API prices: $2,473 for Codex and $1,218 for Claude Code. That is what these tokens would have cost at OpenAI's and Anthropic's published API rates, cached tokens at their discounted rate. I use both tools on subscriptions, so read it as a price comparison, not what I paid.
- 162 messages from me, 20 of them corrections by my own classification. The rest were new requests, questions, approvals and three other messages.
- 93 commits, 226 files, about 56,000 lines added. At launch on September 16, the site had gone from 135 pages to 274, and 137 of them were Chinese.
These count the whole effort, including the deployment repairs the redesign forced. Counting only the redesign itself gives 441 runs, 4.32 billion tokens and $2,303 at API prices. Part 6 breaks the tokens down by phase, role and detour.
๐ Two phases, two agent stacks
Codex built the redesign in five days; Claude Code, directing Codex one fix at a time, got it deployed and made it match the design. The work split cleanly in two.

In the first phase I gave Codex one long-running goal: implement the design, keep production safe, use subagents in parallel, and review every change until no critical bugs remain. The root agent planned sixteen numbered milestones and, counting itself, ran 333 agents over the next five days, several at a time, each in its own git worktree. The rule was that each change passed a Finder, an Adversary and a Referee before it could be committed, and the few exceptions were written down as such. On September 7 the last milestone was committed.
Then it stalled. The staging deploy failed six times, for four different reasons. I told Codex to hand the deployment to Claude Code with everything it had found. In the second phase, Claude Code repaired the deploy path, put the redesign on a staging URL, audited it against the design, and ran twelve rounds of fixes. Codex runs implemented the fixes, Claude Code made some small corrections itself, and Claude Code reviewed each one, screenshotted it and ran the full test gate. Part 3 compares the two ways of working.
๐ก๏ธ Most of the agents were reviewers
More than half of the Codex runs existed to find fault with the rest, yet they used only 17% of Codex's tokens; the biggest single consumer was the agent that coordinated the original build. Of the 630 Codex runs, 347 were dedicated reviewers: 132 Finders, 111 Adversaries and 104 Referees. Together they used 790 million tokens, 17% of Codex's total, or $527 at API prices. A reviewer reads one frozen diff and the files around it, so each run is small. Another 47 review and audit runs launched by Claude Code used 225 million. The root agent of the Codex build used 1.39 billion tokens, 30% of everything Codex used, spent mostly re-reading its own growing conversation as it coordinated 332 other agents.

The reviewers earned their keep. They caught a browser test helper that could print a real credential if debug logging was switched on, and a shared cookie jar that could carry a preview access token from one test server to another. They also kept going long after the value dropped. The custom audio player went through 14 review rounds and still looked far from the design, so I gave it up. Midway through, I gave the loop a stopping rule: "non critical or high priority feedback do not block the milestone from moving forward."
๐งโโ๏ธ What I did: 162 messages, 20 corrections
My job turned out to be pacing and taste, not code review: by my classification, 15 of my 20 corrections were about how the agents were working, and 5 were about the product or how the site looked. In the Codex phase, 20 of my 43 messages were some version of "what is the progress". Most of the process corrections were about pace: "pause the execution", "never stop before you try again", "do not hang on the non-critical issues."
The first four visual corrections were things no automated check was looking for. Section titles ended with a period. The media player's style was still wrong. A series showing three episodes gave no hint that there were more. The "+" on the expand button was too small for its circle. Each was fixed and committed within about an hour and a half. The extended count adds my later correction that the implemented player was far from the design, plus two corrections to how the agents were working. Part 4 goes through all 20.
That is the human-side split I described in Hiring AI: The Assembly Line Worker Era Is Ending: agents handle execution; the person decides what to delegate, distrust, verify, ship and kill.
๐ง The five detours that cost the most
The expensive mistakes were mostly outside the design: environment, deployment, and building something that didn't match the design. Ranked roughly by what they cost:
- "Done" wasn't done. After every milestone had passed its gates, a design-conformance audit found 20 places where the site didn't match the design, 6 of them serious. Three re-audits took that from 20 to 6, 2 and 1, and the ninth fix round closed the last one. (Part 2)
- Previews that would not deploy. Codex's six attempts failed for four reasons, one of them a Python crash that made old CPU firmware a plausible suspect. Claude Code's first attempts then hit a deploy step that grew past 50 GB of memory until my machine ran out and I had to kill it. At API prices, the deploy path and related infrastructure came to $561 across both tools; Claude Code alone spent more there ($447) than on the redesign ($256). (Part 5)
- A player that never matched the design. Fourteen review rounds, and at least 114 million tokens to build and then remove it, went into a custom audio player that ended up far from the design's. One sentence from me deleted 1,850 lines of it. (Part 4)
- Staging broke production writes. About seven staging publishes filled a database that production shares, and every write failed for about two hours. (Part 5)
- The wrong design file. The first research pass analyzed a stale copy of the design, a 96 KB archive instead of the current 394 KB one. The review verdict had to be thrown out. (Part 2)
๐งญ What I'd tell someone about to do this
Give agents a spec with a precedence order, a stopping rule for review, a deploy path proven early, and a person who looks at the result. Concretely:
- Fingerprint your inputs. Record the exact design file by hash before anyone reviews anything.
- Deploy something trivial first. The riskiest part of this project was the path to a staging URL, and it was tested last.
- Cap review rounds by severity. Critical and high findings block. Everything else goes to a backlog with a priority.
- Audit against the design after the gates pass. Tests prove behavior; they do not prove the page looks like the picture.
- Isolate staging by capacity, not only by name. A separate key prefix in a shared database is not a separate database.
- Look at it yourself. Five of my corrections came from looking at the product the agents made.
The rest of the series:
- Part 2: What it took to turn a Claude Design prototype into a buildable spec
- Part 3: Codex goal mode built it; Claude Code directing Codex finished it
- Part 4: What only the human caught: 20 corrections across 162 prompts
- Part 5: Keeping production safe while staging a redesign, and where it failed
- Part 6: Where 6.26 billion tokens went in my AI-built site redesign
What would you cap first in your own agent setup: review rounds, parallel agents, or the deploy step?
๐ Four more days, then launch
From September 13 to 16, the work shifted from rebuilding the site to finishing the publication system, planning what came next, and launching it. First came inline video for the blog, so the design walkthrough could play inside the post. Then the real seven-post series exposed blog-index and article-page alignment problems that the short test fixtures had hidden; those were fixed against the design. The section called "Writing" became "Blog," with ๅๅฎข as the Chinese page title and ๆ็ซ in the Chinese navigation. The site's positioning sentence became: "Technical podcasts and a blog on language models, coding agents, and engineering practice."
I also had the agents research AEO and SEO and write a seven-milestone plan. That plan was reviewed and corrected, but it was not implemented before launch. On September 16, the final checks passed, the redesign merged, the production content artifacts were rebuilt, and williamliu.ai went live in the new design.
How this post was made: drafted by Claude Code from my own session logs, checked against those logs by a separate Codex agent, and edited by me. The cover illustration is AI-generated; the charts are drawn from the numbers above.