← Blog

What I delegated to AI agents in a production release — and the three things I kept

By the end of the release my agents answered consent prompts and released checkpoints. The three decisions I kept, what cross-vendor review was worth, and my real prompts graded.

TL;DR

  • At the start of the week I was the y key on every production prompt. By the end, the agents were releasing checkpoints and answering consent prompts, and I made three kinds of decisions.
  • Delegate comparisons. Keep intent: the release window, the irreversible switch, and how much load you put on production.
  • Cross-family review was worth it: it blocked about 40% of about 60 verdicts and caught failures that would have hit production. But it couldn't see the platform, and it needs a stop rule.
  • The reviewer prompt and the stop rule are below.

🔄 From "press y" to three decisions

Early on, every production write ended at an interactive Proceed? [y/N]. The orchestrator would write a shell script, hand it to me, and ask me to type y only if a 64-character digest matched.

I asked why it needed me to verify an exact match. The honest answer was that it could compare strings better than I could. What it lacked was authority, not ability. After that:

  • I stopped running scripts the agents wrote. If a step is delegated, the agent runs it. If not, I approve one sentence.
  • Consent prompts went to an exact-match runner (part 2).
  • Checkpoints went to the orchestrator, with approvals tied to evidence hashes.

That removed hours of handoffs a day. It also exposed what I actually needed to own.

🗣️ My actual prompts, graded

Here's what I typed to the orchestrator during the release, verbatim with the typos, and what each one did. Most were right, but not all.

Six real prompts graded: four right, one partly right, one wrong-ish
fig 01 Six real prompts graded: four right, one partly right, one wrong-ish

“why you do not have the confidence to check whtehre the prompt exactly matches, why you need a human to verify that” — right. It separated authority from ability. The agent wasn't unsure about the comparison; it hadn't been given permission. Once that was named, consent moved to an exact-match runner and hours of handoffs disappeared.

“why you keep delaying it to tomorrow. do it today” — right. Agents optimise for not breaking things, and someone has to own the cost of waiting. The run started that evening. It hit a new platform limit overnight, but the delay wouldn't have prevented that. It would only have found it a day later.

“auto-approve the fresh plan if counts match” — right, and the best of the bunch. Delegation with a condition a machine can check: object count, reference count, zero blockers, zero deletions. The agent knew exactly when it was allowed to say yes, and when it had to stop.

“why do you need so many passes” and “It causes me upstash memory if you pull very oftent” — right, with one wrong word. It's the request quota, not memory. But putting a budget on verification was the right instinct: the plan had five hour-long passes against the live site, each one reading the datastore thousands of times.

“you shall just do it yourself, drive it from there” — partly right. Delegating the next step was right. But that step was deleting 379 keys from a production store, and I didn't state the scope. It was safe only because the agent bound its own consent to the exact digest and key count from the dry run. Better: “run the reset yourself; approve only deletions=379 and this digest; stop on anything else.” Delegate the action and state the bounds in the same sentence.

“pause any codex run” — wrong-ish. “Any” included another agent's background work, which had nothing to do with the release. The agent froze those processes with SIGSTOP, and a watchdog in the other session killed two of them as hung. Better: “stop starting new Codex runs; let the in-flight ones finish.” Say which runs and how to stop them. “Pause” means different things to different processes.

The pattern: my best prompts were either a sharp why? that exposed a hidden assumption, or a delegation with a checkable condition. My weakest gave an agent an action with no scope.

🧱 The three things I kept

1. The release window. The agents' caution kept pushing the schedule out a day at a time: re-pin, re-plan, wait for tomorrow morning. I asked them to do it today. They found that one phase could start that evening, and it did. Agents optimise for not breaking things. Someone has to own the cost of waiting.

2. The irreversible step. Pointing the domain at the new deployment needed my explicit go. I gave it once, in advance, scoped to that session's switch. That way it didn't block a checkpoint at noon on a Saturday, and the authority was still mine.

3. Load on production. After the switch, the plan called for five full verification passes against the live site. Each pass sent thousands of requests through the site's functions, and each request read Upstash, where I have a monthly request quota. I asked why so many. The honest answer was that the cadence assumed passes of a few minutes, not an hour. I ended the heavy watch after two clean passes and switched to ten URLs every 30 minutes. An agent will verify until something stops it. Verification has a cost, and the owner of that budget has to set the stop.

⚖️ What the reviewer was worth

60 review verdicts in 7 days: 25 blocked on a CRITICAL or HIGH finding, 33 passed
fig 02 60 review verdicts in 7 days: 25 blocked on a CRITICAL or HIGH finding, 33 passed

The setup: OpenAI Codex implemented, a Claude Opus reviewer from a different vendor reviewed, and only CRITICAL or HIGH findings blocked.

What it caught that would have hurt on release day:

  • npx waiting for an install prompt inside a pseudo-terminal. That would have hung the first upload and left a reservation needing manual repair.
  • A rollback handler that died with its logging pipe on SIGTERM, which disabled the safety net exactly when it was needed.
  • A rollback that reported success without reading the aliases back.
  • A recovery command that couldn't resume a crash in the middle of a seal.

What it couldn't catch: anything about the real platform: Upstash's Lua limits (part 1) and Vercel's query-preserving redirects (part 2). It reads code and runs local tests, so ground truth has to come from rehearsal.

What it cost: 10–20 minutes per review and two or three rounds on some changes. Once it was wrong: it suggested a post-rollback check that could never pass, then flagged its own suggestion in the next round with evidence. Findings that carry evidence can correct themselves. That only works if you require evidence.

Steal this: the reviewer prompt skeleton

Review <commit> (parent <sha>) in <repo>. It runs against PRODUCTION <when>.
Do it yourself, synchronously. READ-ONLY w.r.t. production: never run <list of write commands>.
Context (verified facts): <what's already proven, with numbers>.
Focus, highest stakes first:
  1. <the failure that would hurt most, phrased as a question to answer>
  2. <rollback / fail-closed paths: signals, pipes, partial failure, readback>
  3. <is every guard actually tested? mutate 3-4 guards, report killed/survived>
Output: findings tagged CRITICAL/HIGH/MEDIUM/LOW with file:line and a concrete fix;
at most ~8 real findings; MEDIUM/LOW get a one-line "why not blocking".
End with exactly: VERDICT: READY TO COMMIT — <C>C/<H>H/<M>M/<L>L  or  VERDICT: NOT READY — …

Three things in it matter most:

  • Name the failure you fear, so the reviewer doesn't spend its budget on style.
  • Make the reviewer mutate the guards. "Tests pass" is not "tests would fail".
  • End with a verdict line a script can parse, so the gate is mechanical.

Steal this: the stop rule

  • Only CRITICAL and HIGH block. MEDIUM and LOW go into a residuals list, with the one-line reason each is acceptable.
  • If a round's blocking findings are all test-completeness items with no production defect, stop iterating. List the invariants already covered by mutation-verified tests, and ask the reviewer to either name a genuinely new class of invariant or accept the suite. One more possible test can always be found.
  • Match review depth to the stakes. A one-line wiring fix gets a 10-minute focused review, and a new consent mechanism gets the full treatment.

💡 The division of labour I'd use again

Agents: implementer, cross-vendor reviewer, orchestrator, rehearsal harness. Me: release window, irreversible switch, production load budget
fig 03 Agents: implementer, cross-vendor reviewer, orchestrator, rehearsal harness. Me: release window, irreversible switch, production load budget
Work Who
Implementation, tests first implementer model
Adversarial review a model from a different vendor, bounded, with verdicts a script can parse
Running phases, releasing checkpoints, answering consent prompts orchestrator, with exact matching and evidence hashes
Ground truth about the platform rehearsal on the real backend (agents run it, a human scopes it)
Release window, irreversible switch, production load budget me

Core thesis. Agents can carry almost all of a production release. They can't own its cost of delay, its irreversibility, or its budget. Make those three explicit, give them to a person, and delegate everything that's a comparison.

That's the series:

  • Part 1: more agents, same blind spot. Rehearse on the real backend.
  • Part 2: a runner that fails closed makes the agents' friction safe.
  • Part 3: delegate comparisons; keep intent.

Which decision in your release process is really a comparison wearing a human's name?