← Blog

More agents, same blind spot: what our AI review loop couldn't see

An implementer model and a cross-vendor reviewer still let three real Upstash limits reach production. Why more agents add opinions but not ground truth, and the rehearsal that fixed it.

Part 0 (prequel) covers why this migration had a plan at all: the push that broke production.

TL;DR

  • I ran a week-long production data migration with a team of AI agents: one model implemented, a model from a different vendor reviewed, and an orchestrator ran the sessions. Over the week the reviewer sent back about 60 verdicts.
  • It still failed in production three times, each time on a real platform limit (Upstash's Lua runtime) that no agent had ever touched.
  • The reason isn't model quality. The implementer wrote the code and the test fake; the reviewer read both. Everyone agreed, and nobody had touched reality. Adding agents adds opinions, not ground truth.
  • The fix: measure the platform first, encode the measurements in the fake, and rehearse on the real backend in a fenced scratch namespace. A checklist and the guard code are below.

🧭 The setup

My site, williamliu.ai, serves podcast audio from Cloudflare R2 and keeps a small media index in Upstash (serverless Redis). I was moving about 3,900 media records to a new content-addressed format: every file name carries its hash, and Lua scripts update the index atomically.

The division of labour:

  • Orchestrator: a Claude session that planned and ran each production session.
  • Implementer: OpenAI Codex, writing each change tests-first.
  • Reviewer: a Claude Opus reviewer. It gave every finding a severity with file and line evidence, and any CRITICAL or HIGH finding blocked the change.

By the numbers, the process worked: roughly 60 review verdicts in a week, about 40% of them blocking. Then production said no, three days running.

🔍 Three failures nobody in the loop could have seen

Everything a model wrote or read stayed green; first contact with real Upstash was a production day
fig 01 Everything a model wrote or read stayed green; first contact with real Upstash was a production day

Day 1: registry overflow. Batches of 100 records failed. Upstash's Lua table.concat has a hard ceiling counted in elements, not bytes. On our database, 2,561 entries worked and 2,562 failed. A 100-record batch needed about 3,500.

Batch of 100 needs 3,530 concat elements against a 2,561 ceiling; batch of 20 needs 730
fig 02 Batch of 100 needs 3,530 concat elements against a 2,561 ceiling; batch of 20 needs 730

Day 2: a field silently dropped. On Upstash, cjson.null is nil. Decoding {"a":null,"b":1} and encoding it again gives {"b":1}. Real Redis keeps a null sentinel, and so did our fake. All ~3,900 records were written without one nullable field, and the seal correctly refused them.

Day 3: HTTP 400. The body said ERR Error running script: execution timed out. Upstash caps the CPU each script call may use. On our data, sealing 250 records took 0.15 s and 500 timed out at 0.27 s, so the practical budget is about a quarter of a second per call.

Script CPU by records per call: 64 ≈0.04 s, 100 0.07 s, 250 0.15 s, 500 timed out at 0.27 s
fig 03 Script CPU by records per call: 64 ≈0.04 s, 100 0.07 s, 250 0.15 s, 500 timed out at 0.27 s

🗣️ The prompt that changed the plan

On the third bad day I typed this to the orchestrator, typos and all:

“you do it everyday but nothing is accomplished. wasting time”

What it got right: it named the pattern instead of the latest bug. Each day ended with a correct, reviewed fix for yesterday's failure, and a new production surprise the next day. That message made the agent stop patching and ask why the loop kept missing things, which led straight to the rehearsal rule below.

What it lacked: a direction. Frustration tells an agent something is wrong, not what to change. A better version: “before the next production session, prove the whole write path on the real backend.” That's what I should have asked for on day one.

💡 Why more agents didn't help

Look at where each artifact came from:

Artifact Written by Checked against
Migration code implementer model tests
Test fake of the store implementer model the same tests
Tests implementer model the fake
Review reviewer model code + tests + fake

Every check ran against something a model had written. The reviewer was excellent at finding contradictions inside that loop. It caught races, missing guards and broken rollback logic (more in part 3). But it had no way to know that Upstash's Lua differs from Redis, because nothing it could see said so.

Core thesis. Agent review raises internal consistency. Only contact with the real system raises correctness about the real system. Budget for both.

Agents make this worse in one way. A human engineer who has been burned by a platform carries that scar tissue into the next review. A model reviewer starts every review fresh, unless the scar is written into the fake or the checklist.

🛠️ What fixed it

  1. Measure before you build. A few pure-compute probe scripts against the real Upstash database took about 20 minutes and produced numbers that hold: the concat ceiling, the script time budget, null behaviour.
  2. Put the measurements into the fake. Our fake now drops nulls, refuses more than 2,561 concat entries, and has an operation budget for script work. Each of the three bugs now fails a unit test on the old code. The fake encodes the platform's limits, not the model's assumptions about them.
  3. Rehearse on the real backend, fenced. A harness runs the actual write path (begin → batches → seal → strict reads → mappings) against production Upstash, in a scratch key prefix, at real scale: 2,000 objects and 200 identities. The real ~2,000-record seal then ran in 2.5 s, paged 64 records per call.

Steal this: the scratch-namespace guard

The harness needs write credentials, so the guard is the whole safety story. Every command goes through it:

PREFIX = "m4ctest:"

def guarded(cmd):
    verb = cmd[0]
    if verb == "EVAL":
        keys = cmd[3:3 + cmd[2]]            # only declared KEYS
    elif verb == "SCAN":
        assert cmd[2] == "MATCH" and cmd[3].startswith(PREFIX)
        keys = []
    elif verb in {"GET", "SET", "DEL", "HGET", "HSCAN", "SMEMBERS", "INCR", "TYPE"}:
        keys = cmd[1:] if verb == "DEL" else cmd[1:2]
    else:
        raise AssertionError(f"unreviewed command {verb}")
    assert all(str(k).startswith(PREFIX) for k in keys), (verb, keys[:2])
    return ORIGINAL_POST(cmd)

It refuses to start if the prefix already has keys, and it deletes everything it created in a finally block.

Steal this: the "measure the platform" checklist

Before the first production session, probe and record:

  • [ ] Script/function CPU or wall-time limit, found by sweeping input size (e.g. 100 / 250 / 500 records).
  • [ ] Stack, concat or argument-count ceilings, swept to the exact failing size.
  • [ ] Null / empty / missing-field round-trip through every serializer in the path.
  • [ ] Request size limits, both body and reply.
  • [ ] Monthly request quota versus the read volume of your verification and rehearsal plan. (Ours mattered: see part 3.)
  • [ ] Then: encode each one in the fake, and add a regression test that fails on naive code.

⚖️ The cost I accepted

Rehearsing against production Upstash means running with write credentials near production data. I accepted that with the prefix guard, the empty-prefix check, and cleanup. The alternative was what we did first: discover each limit on a production day I had blocked out.

Next: the release runner I'd hand to an agent again — and the 13 launches it took, broken down by who caused each stop.

What's in your agents' test fakes that nobody has measured against the real system?