← Blog

13 launches, 0 bugs in the release, 1 switch: a release runner you can hand to an agent

Thirteen launches, zero bugs in the release itself, one clean switch: the five mechanisms that made an agent-driven release fail safe, with code you can copy.

TL;DR

  • The second half of the migration (upload about 2,000 files, map them, deploy, switch the domain) ran as one long script under systemd, driven by an AI orchestrator. It took 13 launches before the switch went through.
  • Not one of the 12 stops was caused by the release itself. Five were caused by the orchestrator agent, four by our own verification tooling, two by the network, and one by a time gate working as designed.
  • None of them harmed production, because of five mechanisms. Each one is below, with code you can copy.

📊 Every stop, by cause

Runner stops by cause: orchestrator agent 5, verification tooling 4, network 2, time gate 1, the release itself 0
fig 01 Runner stops by cause: orchestrator agent 5, verification tooling 4, network 2, time gate 1, the release itself 0
Cause Stops Examples
Orchestrator agent (state or edits) 5 left a stale build receipt from the previous attempt; a sed edit failed and it relaunched without checking; copied a count from a rehearsal that didn't match production; a Python path missing for one script
Our verification tooling 4 refused a known-inert page; ledger rebuild differed in order (Python hash randomisation); allowlist missing in one step; rejected CDN redirects that keep the query string
Network 2 TLS handshake timeout to R2 at file 767 of the upload; Vercel CLI's final fetch failed after a successful deploy
Time budget gate 1 not enough window left for the remaining phases, so it stopped as designed
The release being deployed 0
Timeline of 13 launches across Friday and Saturday, colored by cause; launch 13 switched the domain
fig 02 Timeline of 13 launches across Friday and Saturday, colored by cause; launch 13 switched the domain

If you're deciding whether to let an agent drive a release, this table is the honest answer. The agent and its tools cause most of the friction. A runner that fails closed is what makes that friction safe.

Runner flow: resume checks, HOLD 1, writes with exact consent, HOLD 2, deploy without domain, verify every URL, HOLD 3 and 4, switch with rollback armed
fig 03 Runner flow: resume checks, HOLD 1, writes with exact consent, HOLD 2, deploy without domain, verify every URL, HOLD 3 and 4, switch with rollback armed

🛠️ Mechanism 1: approvals tied to evidence hashes

The runner pauses at four checkpoints. At each one it writes the machine-checked evidence to a hold file and waits for a "go" file. The go file must be newer than the hold file and must contain the hold file's SHA-256, so a stale or pre-written approval can't release it.

hold() {  # hold N title evidence-file
  local n=$1 sha
  cp "$3" "$S/HOLD-$n.txt"; sha=$(sha256sum "$S/HOLD-$n.txt" | cut -c1-64)
  until [[ -f $S/GO-$n && $S/GO-$n -nt $S/HOLD-$n.txt ]] && grep -q "$sha" "$S/GO-$n"; do
    [[ -f $S/STOP ]] && fail "STOP at HOLD-$n"; sleep 10
  done
}
# reviewer releases it:  echo "reviewed $(sha256sum HOLD-3.txt | cut -c1-64) <note>" > GO-3

Every approval in the log names exactly which evidence it approved. My agent released 14 checkpoints this way and I released none, and the audit trail is the same either way.

🛠️ Mechanism 2: consent by exact match, not by eye

The tools ask Proceed? [y/N] after printing a JSON summary. A small pseudo-terminal wrapper answers instead of a human. It compares the whole summary with strict types, and only on the first prompt:

def strict_eq(a, b):
    if type(a) is not type(b): return False          # 0 != False, 1 != True
    if isinstance(a, dict): return a.keys() == b.keys() and all(strict_eq(a[k], b[k]) for k in a)
    if isinstance(a, list): return len(a) == len(b) and all(map(strict_eq, a, b))
    return a == b

# on "Proceed? [y/N]": answer "y" only if this is prompt #1 and strict_eq(last_plan_json, EXPECTED); else "n"

EXPECTED comes from sealed evidence: the plan digest, counts, target and object key. Test it the way you'd test production code. Feed it the tool's real confirmation output, then mutate one field and expect "n".

🛠️ Mechanism 3: every phase is resumable and proves where it stands

A rerun never trusts a local "done" flag. It reads the store and classifies:

  • Uploads use reservation IDs derived from key and hash. A rerun skips everything already acknowledged and re-uploads the one file whose reply was unknown, under the same reservation. The resumed run did exactly that: 1,180 uploaded, 767 skipped, 0 failures.
  • Mappings are classified as applied-exactly, absent or mismatched. The consent runner only sees the absent ones, and any mismatch stops the run.
  • The resume entry point (RESUME_AT=mappings) re-runs every start-up check before skipping ahead.

🛠️ Mechanism 4: retry only the failure you understand

for try in 1 2 3 4 5 6; do
  run_upload > "$log"; rc=$?
  [[ $rc -eq 0 ]] && break
  transient=$(grep -cE '^unready [^ ]+: (direct PUT reply unknown|PUT outcome unknown)' "$log")
  other=$(grep -cE '^(refused|Traceback|Error)|requires audited reconciliation' "$log")
  [[ $rc -eq 2 && $transient -ge 1 && $other -eq 0 ]] || break     # anything else: stop
  sleep 20
done
[[ $rc -eq 0 ]] || fail "upload failed after $try tries"

A blanket retry would have hidden two real problems we hit earlier in the week. This one retries only the single failure class we had proven to be safe to retry.

🛠️ Mechanism 5: test the verifier like production code

The gate before the domain switch checks every URL the new build emits, about 80,000 of them, plus every media byte. It is also code, and it caused four of the twelve stops. The most expensive one:

Vercel 308 redirects keep the request's query string. /assets/x.m4a?v=123 → https://media.example/x.m4a?v=123. The gate required the Location to equal the media URL with no query, so one pass reported 4,758 failures on a site that was working perfectly.

Checker false failures before and after fixes: simulation 1,835 to 0, baseline capture 17 to 0, candidate pass 4,759 to 0
fig 04 Checker false failures before and after fixes: simulation 1,835 to 0, baseline capture 17 to 0, candidate pass 4,759 to 0

Two rules came out of it:

  1. If the check fails on the system you're replacing too, suspect the checker first. One direct probe of the old deployment showed the same redirect.
  2. Your offline simulation must reproduce the live failure before you trust its green. Our simulation's stub answered redirects the way the gate expected, not the way Vercel does. The fix required red on the old checker (4,759 failures) and then green on the new one.

🧩 Platform facts worth knowing (Vercel, as of Oct 2026)

  • vercel deploy --prod --skip-domain creates a production deployment, but doesn't move your custom domains or the project's "current production" pointer. Crons follow "current production", not your domain aliases. We verified this before and after every deploy.
  • The project's default *.vercel.app alias does move on deploy.
  • npx <cli> inside a pseudo-terminal asks "Ok to proceed?" before installing a missing package. Pin the CLI locally, or it hangs an unattended run. The reviewer caught this one before it stranded an upload.

⚖️ The trade-off

Strict start-up checks made some launches refuse for reasons that weren't real problems, like the count copied from a rehearsal. Each refusal cost a fix and a relaunch, typically 5–15 minutes. I'll take a runner that refuses too often over one that proceeds on a guess. The real fix is to derive every expectation from sealed evidence, never from the last run.

Next: what I delegated to the agents, what I kept, and what the reviewer was worth.

Which stop in your last release came from the tooling rather than the thing you shipped?