← Blog

I clicked “Recommended” and my agent pushed to production

The week before a careful migration, my agents broke production twice — once through a push I approved by clicking 'Recommended'. My real prompts, graded, and seven rules for organising agents.

TL;DR

  • The week before my agents ran a careful, week-long migration (parts 1–3), they broke production twice in five days. This post is why that migration had a plan at all.
  • Incident 1: I asked for staging. The agent couldn't reach staging without pushing to main, so it recommended the push and said it "does NOT touch www". I clicked the recommendation. The claim was false, and the build served production for about 22 hours.
  • Incident 2: I approved a "read-only" production snapshot. It used up the database's monthly request quota, and pages returned 503 while feeds silently fell back to a stale copy.
  • The lessons are about me as much as the agents: where to put the guard, how to approve an agent's recommendation, and what to ask for before anything touches production.
  • I quote my actual prompts below, typos included, with a verdict on each.
The week before the migration plan: what I typed, what the agent did, and the production impact
fig 01 The week before the migration plan: what I typed, what the agent did, and the production impact

🧭 The setup

My site, williamliu.ai, runs on Vercel. Podcast media lives in Cloudflare R2, with a small index in Upstash. My agents had finished a large change locally: 40 commits that switched media to hash-versioned file names. Everything had passed the local test suite. Nothing had touched a real environment.

🔍 Incident 1: the push

I typed:

"push to staging first"

That was the right instinct, but it fell short. The repo's staging script refuses to run unless your commit is already on main (it stages "what Vercel builds"). A second, exact-commit staging path existed, but it refused my working copy. So the agent offered me a choice and marked one option Recommended:

Push main, then deploy:staged. "Creates a Vercel deployment record but does NOT touch www.williamliu.ai — aliases only move on vercel alias set / deploy:promote."

I clicked it. The agent pushed. Vercel's git integration built the commit and, because the project was set to assign custom domains automatically, put it straight onto www and the apex domain. The safety claim came from the agent's own deploy docs, and it was true for the CLI path and false for the git path.

Two roads to production: the CLI path with every safeguard, and the git path with none
fig 02 Two roads to production: the CLI path with every safeguard, and the git path with none

The next evening the agent noticed the domains had moved: "I was wrong, and production changed." Then came the blast-radius report: production broken, then the RSS feed broken, then 137 podcast downloads returning 404. I pushed back:

"stop, which production is broken"

"the webs ite is not boroken, the RSS feed is broken. is that right"

"I can still play audio on the website"

Each time, the scope shrank. The 137 failures came from checking the URLs the local build would emit against production. They proved that production couldn't serve the new file names. They didn't prove that subscribers saw them, because the live feeds turned out to still list the old URLs. The agent had overstated the impact three times. We rolled back to the previous deployment.

🔍 Incident 2: the read-only snapshot

Four days later the agent proposed two things, and I clicked "approve both A and B":

  • A locked production: a frozen release branch as the production branch, and automatic domain assignment off.
  • B took a read-only snapshot of the production data store.

A was exactly right. B scanned about 10,000 keys with a command per key. The database was on a free plan capped at 500,000 requests a month, and already near the cap from normal traffic. The scan used up the rest within three minutes. Runtime pages returned 503.

Worse, the podcast feeds kept returning 200. They had silently fallen back to a copy bundled with an older build: 17 episodes instead of 28. A 200 is not health. I upgraded the database plan the next morning.

🗣️ My prompts, graded

My prompts in the incident week, graded
fig 03 My prompts in the incident week, graded

"push to staging first": right, but incomplete. It stated the goal, not the constraint. When staging turned out to need production first, the agent "solved" it by going through production. Better: "stage it. If staging requires touching production in any way, stop and ask me."

(clicked) "Push main, then deploy:staged (Recommended)": wrong. I wasn't really approving an action. I was approving the agent's claim that the action was safe, and the claim was untested. One question would have exposed it: "What here is irreversible, and how do you know it isn't?" The honest answer was "I read it in the docs".

"stop, which production is broken" / "the webs ite is not boroken, the RSS feed is broken. is that right": right. Agents under alarm overstate. Making it narrow the claim to what it had actually measured turned "production is down" into a precise, much smaller statement.

"this is a deeper issue, why we break production directly. find the learning and make architectual changes to ensure it does not happen in the future": right. It asked for the system, not the symptom. The answer was structural: staging sat downstream of production, and every guard lived on the CLI path while the git path had none.

"ask codex using astra with reasoning high to review the incident and compare its finding with claude's": right, and the most valuable prompt of the week.

  • A model from another vendor reviewed the post-mortem blind, then compared notes. It corrected three overstated claims: the subscriber impact, a "mappings were never created" conclusion, and a lesson that one vercel alias ls would have caught the problem.
  • That last one was logically wrong. Listing aliases shows where they point now, not what the next push will do. A safety claim about a future state change must be tested by making that change somewhere safe.

(clicked) "approve both A and B": partly right. A was the real fix. B needed a budget question before approval: "how many requests will this make, and how much quota is left?"

"be critical of what you need to build focus on the crirical changes" / "prioritze fixjng bugs first before making more feature changes": right. After an incident, the backlog of features is the riskiest thing in the repo.

"I'm suspecting that you didn't consider a migration plan from the previous file name format to the new file name format with the sha": mostly right.

  • The investigation found a plan did exist on paper: adopt the old records, mint the new ones, then switch.
  • But it ordered the git push before the data migration, it lived only in prose, and it had never been rehearsed. My suspicion was right in spirit.
  • That led to the instruction that started the next three posts: "carefuly execute the migration plan, verify and deploy."

🛠️ How I organise agents now

1. Put the guard on the action, not on the agent. No instruction to an agent protects a path that never reads instructions. The fix was a setting:

  • the production branch is release;
  • automatic domain assignment is off.

Now a push to main can't reach users, and only an explicit alias switch can.

2. Treat "Recommended" as a claim to verify. Before approving any option that touches production, ask the agent to split its answer into verified (observed, with evidence) and assumed (read in docs, inferred). Our root-cause write-ups now tag every statement VERIFIED or INFERRED.

3. Give every instruction a stop condition. "Do X" invites the agent to find a way. "Do X; if that requires Y, stop and ask" keeps it on your way.

4. Data before code, checked by a machine. A prerequisite that lives only in a plan document will be skipped by the first push that ignores the plan. The migration made each prerequisite a gate the runner checks before it continues.

5. Review the post-mortem, not just the code. A blind, other-vendor review of the incident analysis caught more overstatements than any code review did that week.

6. Every read on production needs a budget. "Read-only" still costs requests, and a free-tier quota can turn a census into an outage. Before any bulk read: how many requests, how much headroom.

7. Split the roles by phase. For planning I now ask for two strong models from different vendors at high effort, writing and reviewing the plan. For implementation, one model writes and the other reviews. As I put it at the time: "in the planning phase you shall use gpt-6-astra with high effort and Opus 5.5 high effort."

Core thesis. Both incidents started with an agent acting confidently on a claim nobody had tested, and with me approving the claim along with the action. The fix is to make claims testable, and to put the guard where no claim can bypass it.

Next, part 1: the migration that followed, and why more agents didn't catch the platform limits that broke it.

Which "Recommended" button in your workflow are you approving on trust?