← Blog

What only the human caught: 20 corrections across 162 prompts

Part 4: I sent the agents 162 messages over fifteen days. Twenty were corrections, and they fall into two groups: five visual or product judgments, fifteen about how the agents were working. Neither group was caught by the checks in place.

Part 4 of 6 in "Rebuilding williamliu.ai with agents." Start with the overview.

TL;DR

  • 💬 I sent 162 messages in fifteen days: 67 new requests, 38 questions, 34 approvals, 20 corrections and 3 others.
  • 👀 Twenty of my 43 messages to Codex asked what was going on.
  • 🎨 Five were visual or product judgments, and none was something the checks in place could decide for me.
  • 🧭 Fifteen were about how the agents worked: pace, priority, scope and who should review a plan.
  • 🗑️ I liked the design's audio player, but the one built after 14 review rounds was far from it; building and removing it took at least 114 million tokens.
  • 🧑‍⚖️ The agents stopped to ask me 23 times, and everything that could touch production waited for a yes.

Agents did almost all of the building and reviewing; what they needed from me was taste, pace and priority, and no amount of review replaced any of the three. I counted every message I typed into Codex and Claude Code through the launch, 162 of them, and classified each one. This part is about the 20 that corrected something, and what they have in common. Five were judgments about the site or product; fifteen changed how the agents worked. Neither group was caught by the checks in place.

My 162 messages by type: 67 new requests, 38 questions, 34 approvals or decisions, 20 corrections and 3 other. The corrections split into 5 visual or product judgments and 15 about how the agents were working.
fig 01 My 162 messages by type: 67 new requests, 38 questions, 34 approvals or decisions, 20 corrections and 3 other. The corrections split into 5 visual or product judgments and 15 about how the agents were working.

👀 Nearly half my Codex messages asked for status

In the Codex phase, 20 of my 43 messages were status questions, by my classification. "what is the progress" appears three times, "show me the progress" once. "have you started parallel implementation of the tasks?", "did you start parallel tasks subagents", "are you using multiple subagents to work at the same time?" and two more like them show that I asked five times whether the work was actually running in parallel. Five more messages were just "continue" or "resume".

On the first day of the build I asked it to "regularly show me the progress (which milestone is being implemented and what are the subagents beingexecuted in parallel)". I still kept asking. The status the agents wrote down was in their own terms: candidate hashes, review rounds, repair passes. I wanted to know whether the website was closer to done.

What I sent Codex, 43 messages: 20 status questions, 5 "continue" or "resume", 6 corrections, 6 new requests, 5 approvals and 1 other question.
fig 02 What I sent Codex, 43 messages: 20 status questions, 5 "continue" or "resume", 6 corrections, 6 new requests, 5 approvals and 1 other question.

The lesson I took: an autonomous run needs a progress view written for the person paying for it. Milestones done, milestones left, what is blocked and on whom. When the only view is the agent's own log, the human becomes a polling loop.

🎨 Five visual or product corrections no test was looking for

Every correction in this group was a judgment about how a reader experiences the product; four changed the site, and the fifth corrected the account of why I abandoned the player. The first four came from using the staging site:

  1. "There is also another small bug that the title has a '.' at the end". Page titles read "Technical podcasts." and "About." The design's own prototype had those periods, and production had the same style ("Technical podcasts.", "Blog."). The agents copied the design faithfully. I decided I didn't want them, and the brand guide changed with the fix.
  2. "The media player's style is still wrong. Revert the media player style back to the one before the redesign". The player that had been built looked far from the design's, so this one overruled the approved spec, which had called for the custom player. More on this below.
  3. “with the 3 episode displayed, it is hard for user to know that there are more episodes. add a clear indication like "View all xxxx episodes" there.” The podcast page showed three episodes per series, exactly as designed. For the 60-episode series, nothing hinted at the other 57. A copy check flagged my wording as longer than the site's button guideline; the link shipped with my wording anyway.
  4. "The + sign shall fill the cirle in the top right corner. now it is too small". The expand button met its 44-pixel size requirement and passed its tests. The plus inside it was tiny.
  5. "I uploaded two screenshots of what the media play looks like in the design. I like that design. But the implemented player is very far from the design. So I have to give up that. make sure when you talk about the media player used many tokens and I decided to revert it." This did not change the site again. It corrected the story: I liked the design and gave up the implementation, not the idea.
Five visual or product corrections: four site fixes—trailing periods removed, the custom player replaced, a "View all 60 episodes" link added and the plus sign enlarged—and one correction to the account of why I gave up the player.
fig 03 Five visual or product corrections: four site fixes—trailing periods removed, the custom player replaced, a "View all 60 episodes" link added and the plus sign enlarged—and one correction to the account of why I gave up the player.

The screenshots I sent with the first correction showed four more defects. The home page's episode rows lost their episode codes after a live update. The footer was missing a link and stacked in the wrong typeface. A rule under the page title was narrower than the rules below it. The episode player stretched its buttons across the page. The missing episode codes only happened on the live staging site, after its scripts loaded fresh data, which none of the existing local checks caught.

🗑️ The player I had to give up

I liked the design's audio player, but the one the agents built was far from it, so after 14 review rounds I gave it up; building and removing it took at least 114 million tokens. The design had a full player with a fine waveform you could click to seek, and a slim mini player in every list: a round play button, a thin progress line, the time. The agents built a player with keyboard support, a fallback to the native player and synchronized state across the page. It went through 14 review rounds. It still didn't look like the design.

The design's players against what was built, at the same scale: the design's full player has a fine two-tone waveform and a single speed chip; the built one has chunky bars and three boxed buttons. The design's mini player is a round play button and a thin progress line; the built one is a text "Play" button, three boxed buttons and a slider.
fig 04 The design's players against what was built, at the same scale: the design's full player has a fine two-tone waveform and a single speed chip; the built one has chunky bars and three boxed buttons. The design's mini player is a round play button and a thin progress line; the built one is a text "Play" button, three boxed buttons and a slider.

The agents that built and reviewed it by name used 86 million tokens, and the earlier rounds aren't in that count. When I saw it on staging, I asked for the player the site had before. That revert took another 28 million tokens, touched 43 files and deleted 1,850 lines, 1,263 of them tests. The site now uses the browser's native audio control everywhere, keeps the sticky player bar the old site already had on episode pages, and still lets only one episode play at a time.

The lesson is about what review checks. The review rounds tested whether the player worked, and it did. Only the fourteenth round compared it with the design, and even after that round's fix it looked far from the design's player. A side-by-side look at the first working version would have raised that before most of those rounds ran.

🧭 Fifteen corrections about how the agents worked

Most of my corrections told the agents to stop doing something well that no longer mattered. Grouped by what they were trying to fix:

  • Pause. "pause the process, I will resume it soon"; "pause the execution".
  • Stop polishing and move. "non critical or high priority feedback do not block the milestone from moving forward. just document the rest of the feedback into a backlog list and give them the proper priority." Two days later: "why the webstie is still not moving on fast. what did you accomplished in the past 24 hours", then "do not hang on the non-critical issues. you don't need to get 100% bug free or a perfect code."
  • Don't give up either. "never stop before you try again", after the build had stopped and was waiting.
  • Change priority. "given there is a big redesign pending deplyment. my first priority is to fix the deployment issue to be able to verify the redesign." Claude Code had spent about a day researching and planning a larger overhaul of how the site stores and deploys media. It was a good plan for the wrong week. I deferred it, and the deploy fix landed the next morning.
  • Plan corrections. Three messages corrected details of that deferred plan, and one stopped work so I could review it first.
  • Report failures. "the previous staging deploy failed because the process run out of memory, debug the issue and fix it."
  • Set scope and choose the reviewer. "we dont need specific code for my website redesign article drafts. also the website code can commite if clean" removed the website-specific code package from the series. "Using Fable to Review the plan one more time after you are done with the current edits" redirected the final plan review from Codex to Claude.
Fifteen corrections to how the agents worked by theme: 2 pause, 3 stop polishing and move, 1 keep trying, 2 change priority, 3 plan details, 1 stop to review, 1 failure report, and 2 about scope and reviewer choice.
fig 05 Fifteen corrections to how the agents worked by theme: 2 pause, 3 stop polishing and move, 1 keep trying, 2 change priority, 3 plan details, 1 stop to review, 1 failure report, and 2 about scope and reviewer choice.

The pattern is consistent. The agents' behavior favored thoroughness: another review round, another hardening pass, a better plan. My messages kept pointing at one thing: a working site on a staging URL that I could look at. The severity rule changed the review policy for everything after it. Until then, medium findings also had to be fixed or explicitly waived before a milestone could move; after it, only critical and high findings could stop one.

📐 After September 12, real content broke the baseline

The launch extension produced more fixes without turning every prompt or objection into a human correction. On September 14, I uploaded a screenshot of the staged Blog page and wrote, "uploaded how the blog page looks like, check if it aligns with the design". The classifier kept that prompt as a new request, not a correction, even though it led to five fix commits across the index and article layouts.

Every gate had been green. The screenshot baseline was the problem: its fixture posts had short titles, one-line summaries, no covers and few tags. That content made the wrong composition look fine and then froze it into the baseline. The replacement fixture has a long title, a multi-sentence summary, a cover and several tags. A visual baseline is only as good as the case it freezes.

The same distinction matters for copy. During the Writing-to-Blog rename, an agent changed the positioning sentence to "podcasts and blog posts". The copy checker rejected that wording. My later "change the blogs." was classified as an approval or decision, not a correction; it selected the final line: "Technical podcasts and a blog on language models, coding agents, and engineering practice." The count is of my messages by function, not every defect fixed or every objection an agent raised.

🧑‍⚖️ The questions the agents asked me

The agents stopped for me 23 times, in two different ways, and every stop that involved production, spending or deletion waited for an explicit yes. Claude Code made 18 structured AskUserQuestion calls containing 33 questions; all 18 calls got an answer. Codex blocked itself five times, for a total of almost six hours of waiting on me: the design, the corrected spec, sandbox permissions, uploading a specific commit to a specific preview project, and a diagnostic step.

Some of those questions were the most consequential decisions in the project:

  • Whether to delete about 2,200 stale staging records from a database production shares, while production writes were failing (Part 5).
  • Which existing URLs had to keep working exactly after the redesign.
  • Whether to keep the home page on its wider layout at large screen widths (I kept it).
  • Every push that could reach production. The redesign went live on September 16, after my explicit command.

🧭 What to keep for yourself

Keep taste, pace and priority; give the agents the rest.

  • Look at the running thing early, next to the design. Review rounds can prove it works; here only the last of fourteen compared it with the design.
  • Put real content in visual baselines. A short fixture can make the wrong layout look correct and preserve the mistake.
  • Write the stopping rule on day one. Decide which severities block and where the rest goes.
  • Ask for progress in your own terms. Milestones done and left, blockers and who owns them.
  • Say what matters this week. A good plan for the wrong week is still a detour.
  • Own every irreversible yes. Production, deletion, money.

In Encoded Judgment Becomes the New Moat, I argued evaluation and escalation patterns—not all judgment—can become infrastructure. The Real AI Coding Race: Which Tool Best Encodes Engineering Judgment named task contracts, boundaries, approvals, verification and review. This project confirms it: severity, retry, plan-review, visual-check and representative-fixture rules, plus “View all” and plus-size criteria, can be encoded.

Some calls were not derivable from inputs: removing periods reversed design and production precedent; abandoning the player reversed the spec; pauses and that week's deploy-first priority depended on context. Irreversible approvals still need a human. That sharpens Hiring AI: The Assembly Line Worker Era Is Ending, where I wrote “your taste becomes infrastructure”: taste only becomes infrastructure after a person makes it explicit.

Part 5 is about the approvals that mattered most: keeping production safe while staging a redesign, and the day staging broke production writes.

Which of your corrections to an agent could have been a rule you wrote down on day one?

How this post was made: drafted by Claude Code from my own session logs, checked against those logs by a separate Codex agent, and edited by me. The cover illustration is AI-generated; the charts are drawn from the numbers above.