← Blog

Codex goal mode built it; Claude Code directing Codex finished it

Part 3: one long-running Codex goal ran 333 agents with a three-role review loop and built the redesign in five days, then stalled at deployment. Claude Code took over as the orchestrator, sending fresh Codex runs to implement and review the fixes and making small corrections itself. What each mode was good at here, with numbers.

Part 3 of 6 in "Rebuilding williamliu.ai with agents." Start with the overview.

TL;DR

  • 🏗️ Codex goal mode turned one instruction into 333 agents and 43 commits in five days.
  • 🔍 The rule: every change frozen as an exact diff and passed by a Finder, an Adversary and a Referee, with exceptions written down.
  • 🧱 It stalled on things outside the code: a sandbox that couldn't run and a deploy that failed six times.
  • 🔁 Claude Code then worked as the orchestrator: fresh Codex runs for the fixes, small corrections of its own, then its review, screenshots and the full test gate.
  • ⚖️ Codex was better at sustained parallel building; Claude Code at steering, looking at the result, and operations.
  • 🪞 Each caught the other's mistakes, and each made mistakes of its own.

The same redesign was built by two very different agent setups, and in this project they failed in different ways, which made each a check on the other. The first was a single Codex goal: one instruction from me, one root agent, hundreds of subagents. The second was Claude Code in my terminal, deciding what to fix next and handing the work to fresh Codex runs. This part compares them on the same codebase, the same week, with numbers from the logs.

The two setups side by side. Codex goal mode, September 2 to 8: 333 agents including the root, 43 commits, 3.30 billion Codex tokens, $1,671 at API prices. Claude Code directing Codex, September 10 to 12: one Claude Code session, 89 Codex agents, 16 commits, 0.66 billion Codex tokens plus 0.36 billion Claude tokens, $632 at API prices.
fig 01 The two setups side by side. Codex goal mode, September 2 to 8: 333 agents including the root, 43 commits, 3.30 billion Codex tokens, $1,671 at API prices. Claude Code directing Codex, September 10 to 12: one Claude Code session, 89 Codex agents, 16 commits, 0.66 billion Codex tokens plus 0.36 billion Claude tokens, $632 at API prices.

By the production launch on September 16, the inclusive ledger had grown from 514 to 630 Codex runs and from 36 to 53 Claude runs. The 333-run build and 89-run alignment comparison in this post did not change: those strict phases were already complete. The later work gave me more evidence about the same division of labor.

🏗️ Goal mode: one instruction, 333 agents

One paragraph from me became a five-day build that mostly ran on its own, stopping for my approval five times: research, a spec, a plan of milestones, and 43 commits. My instruction to Codex asked it to compare the design with the current site, break the work into small tasks that could run in parallel, avoid the data layer, keep production safe, use subagents, and review changes "until no critical bugs." Codex ran it as a goal: a long-lived objective that keeps going across turns until it is met or blocked.

The root agent researched, wrote the spec and the plan, and waited for my approval at each gate. It stopped five times to wait for me, for a total of almost six hours. The longest wait, about three hours, was for my approval of the corrected spec. Then it fanned out. Fonts, the blog pages, the podcast pages and the browser test harness each became a lane with its own agents and its own git worktree. One agent at a time owned the shared files, and the root integrated finished work one commit at a time.

I told it early to pick models by difficulty, and it did: seven models across the run, small ones for simple tasks and the strongest one as the final Referee.

🔍 Three reviewers per change, by rule

The rule was that nothing was committed until a Finder, an Adversary and a Referee had examined the exact same frozen diff. Each candidate change was staged with an explicit list of allowed paths and fingerprinted with a SHA-256 hash of its diff. The Finder looked for bugs, security problems and missing tests. The Adversary challenged each finding and hunted for what the Finder missed. The Referee, usually the strongest model, ruled on severity. Any upheld blocking finding meant a repair or an explicit decision, then a new hash and a fresh cycle. Late in the build, when independent reviewers weren't available, the root reviewed some changes itself and labeled that as a fallback, not as independent review.

Codex runs and tokens in the build phase by role: 202 reviewers (75 Finders, 62 Adversaries, 65 Referees) used 561 million tokens; 130 implementation, repair and research agents used 1.34 billion; the one root agent used 1.39 billion.
fig 02 Codex runs and tokens in the build phase by role: 202 reviewers (75 Finders, 62 Adversaries, 65 Referees) used 561 million tokens; 130 implementation, repair and research agents used 1.34 billion; the one root agent used 1.39 billion.

That is 202 reviewer runs in the build phase, against 130 building, repair and research runs. By tokens the picture flips: the reviewers used 561 million, 17% of the build's 3.30 billion, while the builders used 1.34 billion and the root agent that coordinated them all used 1.39 billion. A review reads one frozen diff; a builder reads the code it changes; the root re-reads its whole growing conversation at every step. The frozen hash mattered more than I expected. A reviewer cannot approve a change that is still moving, and a repair cannot quietly widen its scope.

The frozen-diff loop made Dynamic Workflows Are the First AI Coding Feature That Looks Like a Process Tool concrete: the process itself was inspectable, repeatable and separate from any one agent's answer.

It found real problems. A browser-test helper would have printed a real access credential if someone ran the tests with debug logging on. A shared cookie jar could carry a preview access token between two test servers. Both were found with fake secrets and fixed before commit.

After this draft, the inline-video feature tested the same idea under Claude Code's orchestration. Codex implemented it, then Codex reviewed it for 15 rounds. After 11 clean-looking review rounds, the repository's full gate—not another reviewer—caught INV-7: publish-side code must not import content-type modules. Reviews 12 and 13 followed that fix; then the full gate caught a second invariant: a build given a snapshot must not reopen its inputs. Reviews 14 and 15 closed it, with round 15 READY TO COMMIT. That did not make the reviews useless. It showed their boundary: frozen-diff review could pressure-test the change, while the repository's own invariants remained a separate source of truth.

🧱 Where goal mode stalled

The build finished on September 7; getting it onto a staging URL is where goal mode stopped making progress. Two things outside the code slowed it down.

  • The sandbox didn't work on my machine. Codex's sandbox couldn't create the network namespace it needs, so even reading an existing file through the patch tool failed. It worked around this by writing whole replacement files and applying them with git apply --check first, and it refused two corrupt patches rather than forcing them.
  • The deploy failed six times. Three attempts crashed inside Python, one failed in a nested build, one on a test, and one because the build environment couldn't launch a tool. More on that in Part 5.

The review loop also had no brakes. The custom audio player reached its fourteenth review round, and it still didn't look like the design. Late on September 7 I asked why the website still wasn't moving and what had been accomplished in the past 24 hours. Soon after, I told Codex to hand the deployment problem to Claude Code, with everything it had found.

🔁 Claude Code as the orchestrator

In the second phase, Claude Code chose each fix, wrote a precise brief, had Codex implement it, and then checked the result itself, making small corrections directly when that was quicker. The loop for each of the twelve fix rounds:

  1. Claude Code wrote a brief: the design reference, the exact problem, the files in scope, and the tests to run.
  2. A fresh Codex run implemented it and ran focused tests. When review found more, a follow-up Codex run fixed it, or Claude Code made a small fix itself.
  3. Claude Code reviewed the diff, took screenshots at desktop and phone widths, compared them with the design, and ran the full test gate itself.
  4. One commit per round, then a staging deploy when it made sense.
One fix round: Claude Code writes the brief, a fresh Codex run implements and tests it (with follow-up runs or small direct fixes when review finds more), and Claude Code reviews the diff, screenshots it against the design and runs the full test gate before committing.
fig 03 One fix round: Claude Code writes the brief, a fresh Codex run implements and tests it (with follow-up runs or small direct fixes when review finds more), and Claude Code reviews the diff, screenshots it against the design and runs the full test gate before committing.

Across the phase that meant 30 Codex runs launched by Claude Code (18 implementations, 12 reviews and audits), plus 59 agents those runs spawned, 46 of them reviewers. The whole alignment phase, from putting the redesign on staging to the last fix, used 0.66 billion Codex tokens and 0.36 billion Claude tokens, $632 at API prices against $1,671 for the Codex build. Staged deploys took about a minute and a half each once the deploy path was fixed.

The pattern continued after the draft. When the real seven-post batch exposed blog layouts that the small fixtures had hidden, I asked for parallel Codex fixers. Three agents worked in separate worktrees on the index, the drop cap, and article metadata and figure labels, followed by one narrow-width pass. Claude Code rendered the real seven posts with a local screenshot harness, verified each worktree, and merged the fixes. The useful split was not just parallelism: Codex owned bounded changes, while the orchestrator kept the visual judgment and integration.

🪞 Each caught the other's mistakes

In this project, the useful property was that the two stacks failed differently. Some examples from the logs:

  • Codex's focused tests missed what the full gate caught. A fix renamed a function argument, and three older tests still used the old name. Codex's own run checked only the tests near its change. Claude Code's full gate failed, and the fix preserved what the old tests meant.
  • Claude Code's review caught a layout bug the tests passed. Expanded episode notes rendered inside the narrow 72-pixel cover column. The browser doesn't place the children of a <details> element as grid items when the element itself is set to display: contents. Claude Code found it in screenshots; geometry tests now pin it.
  • Unapproved copy crept in, and review caught it. Codex added four Chinese strings nobody had approved, and Claude Code removed them in review. An unapproved introduction line on the Projects page also came out in the first fix round.
  • Codex caught Claude Code's plan bugs. Reviewing Claude Code's deploy plan, Codex found a missing -- in an npm command that would have silently dropped a flag the deploy depended on.
  • I could change who reviewed what. A read-only Codex audit fed an AEO/SEO plan, but I asked Claude, not Codex, to review that plan. Claude found three must-fix issues: a protocol conflict over an "Updated" date, under-specified route plumbing for tag hubs, and a trailing-slash change that was too broad.
  • Claude Code misdiagnosed too. It first called a set of killed review jobs a false alarm from the tool harness. The system journal showed that at least one was a real kernel out-of-memory kill, and it corrected itself in the log.

⚖️ What each mode is for

Use goal mode for a well-specified build you can leave running; use an interactive orchestrator when the work is judging results against a picture.

Where each setup was strong in this project: goal mode for long parallel builds, rigorous review on frozen diffs, and days of work with only intermittent input from me; the interactive orchestrator for visual verification, operations and deployment, and changing course within minutes.
fig 04 Where each setup was strong in this project: goal mode for long parallel builds, rigorous review on frozen diffs, and days of work with only intermittent input from me; the interactive orchestrator for visual verification, operations and deployment, and changing course within minutes.

Goal mode shone where the spec was clear. It kept going for days with only intermittent input from me, and it kept its review standard. Its weaknesses were visibility, brakes and coordination cost: I asked it for status 20 times, medium findings kept reopening review rounds until I set a rule, and its root agent alone used 1.39 billion tokens re-reading a conversation that never stopped growing.

The orchestrator shone where the task was looking and deciding: comparing a screenshot with a design, reading a deploy log, choosing the next fix. It turned each of my four visual complaints into a commit within about 90 minutes. Its weakness was cost per fix. Every fix started a fresh Codex run that had to learn the repository cold; those runs averaged 21.7 million tokens, about twice a goal-mode builder's 10.3 million. And at API prices, Claude Code's own sessions cost more on the deploy path than on the redesign. Part 6 has the breakdown.

If I did it again, I'd start with the orchestrator setting up a working deploy path, then hand the build to goal mode with a review stopping rule, then bring the orchestrator back to audit against the design.

Which part of your agent workflow needs a person watching, and which could run for five days alone?

How this post was made: drafted by Claude Code from my own session logs, checked against those logs by a separate Codex agent, and edited by me. The cover illustration is AI-generated; the charts are drawn from the numbers above.