โ† Blog

Keeping production safe while staging a redesign, and where it failed

Part 5: the redesign stayed off production until it was ready, by design. Getting it onto a staging URL took six failed deploys, a suspect CPU, an upload step that ate more than 50 GB of memory and a hidden dependency on private files. Then staging publishes filled a database production shares, and production writes stopped for about two hours.

Part 5 of 6 in "Rebuilding williamliu.ai with agents." Start with the overview.

TL;DR

  • ๐Ÿ”’ Production changes only on my explicit command; the agents could break the staging preview freely, and did.
  • ๐Ÿ’ฅ The first six staging deploys failed for four different reasons; one made the CPU's firmware a suspect.
  • ๐Ÿงน A clean checkout exposed tests that depended on private files that exist only on my machine.
  • ๐Ÿงฎ The deploy's upload step grew past 50 GB of memory until my machine ran out; building locally and uploading the result took 290 MB.
  • ๐Ÿ—„๏ธ Staging and production shared one 256 MB database; about seven staging publishes filled it.
  • โฑ๏ธ Production writes failed for about two hours, until a cleanup I approved brought the database down to 158.5 MB.
  • ๐Ÿš€ On September 16, the preview could not be promoted. Pushing main deployed production; rebuilding and publishing its content artifacts raised Upstash from 156 to 177 MB.

The redesign stayed off production until it was ready, and the safety rules that ensured it worked; the failure came through the one resource staging and production still shared. This part covers what it took to get a large redesign onto a staging URL without risking the live site, the night that went wrong anyway, and how the redesign ultimately reached production. I'm writing it in full because the lessons are in the details.

The safety layers around production: an explicit command from me for any production deploy; previews on isolated URLs behind access protection; prebuilt previews recorded as not promotable; staging content in its own database namespace. The gap: that namespace lived in the same 256 MB database as production.
fig 01 The safety layers around production: an explicit command from me for any production deploy; previews on isolated URLs behind access protection; prebuilt previews recorded as not promotable; staging content in its own database namespace. The gap: that namespace lived in the same 256 MB database as production.

๐Ÿ”’ The rule: production waits for my yes

The agents could break the redesign's staging preview freely, but no push to the production branch, no production deploy and no promotion happened without my explicit instruction in the conversation. This is a standing rule in the repository's agent instructions, and the redesign spec repeated it. Staging could be broken while milestones were in progress. Production stayed on its known-good version until I said otherwise.

Two mechanisms backed the rule. First, previews live on their own URLs behind the host's access protection, and the test suite reaches them with a bypass token bound to that exact host. Second, the deploy path the agents built records previews as not promotable, so one can't be pushed to production by accident. All of the redesign's building, previewing and deploy testing happened on staging and on my machine. When the redesign did go to production, it went on my explicit command, after everything described here.

๐Ÿ’ฅ Six failed deploys, four different causes

When Codex first tried to deploy the finished redesign to staging, six attempts failed, and treating them as one problem would have been the mistake. Three attempts crashed inside Python's YAML parser. One failed inside a nested build without enough output to diagnose. One failed a test about access grants. The sixth reached the hosting provider's build step and couldn't launch a tool the build needed.

The six failed staging deploys on September 7: attempts 1, 2 and 4 crashed inside Python's YAML parser; attempt 3 failed in a nested build; attempt 5 failed a grant test; attempt 6 could not launch a build tool.
fig 02 The six failed staging deploys on September 7: attempts 1, 2 and 4 crashed inside Python's YAML parser; attempt 3 failed in a nested build; attempt 5 failed a grant test; attempt 6 could not launch a build tool.

Codex kept every failed log, and that paid off. When I pushed it to find root causes instead of polishing, the logs showed four problems, not one. Each needed a different fix and a different owner.

๐Ÿ–ฅ๏ธ When the CPU is a suspect

The Python crashes looked like corrupted memory inside the interpreter, and old CPU firmware became a plausible suspect. The signature was strange. A value that should always be text came back as an integer, specifically the integer stored right next to it in memory. It happened at random times in whichever test file happened to be running hot. The same test passed when rerun alone.

Codex's root-cause investigation found the first clue on the machine: kernel logs held eight Python segmentation faults since September 2, all on the same logical CPU. Claude Code then dug further. My desktop has a 13th-generation Intel Core i9, and its microcode, the CPU's own firmware, was version 0x10f. Intel recommends 0x12F or later for a known instability in these chips; mine was older because the Linux microcode package had been installed and later removed. I installed the update and rebooted; the CPU now reports 0x133.

The update didn't make the crashes go away. The same signature appeared again during the heavy load right after the reboot, and the next day a staging deploy died with a segmentation fault. So the CPU is still a suspect, not a proven cause, and I don't have a proven cause. What changed is how the agents read a green test run. If the host can miscompute, one passing run is weaker evidence than it looks. A result now counts only after several signals agree, including the real exit code, a zero traceback count and the health markers.

๐Ÿงน A clean checkout finds what your machine hides

Deploying from a clean copy of the code exposed tests that only passed because private files existed on my workstation. The preview build deliberately leaves out my private podcast drafts, which are ignored by git. Without them, a date validator rejected the empty preview list. A test expected a folder that only existed locally. Two more test groups assumed private media and an installed command were present. The repair touched six files and changed no publishing logic. A fresh checkout then passed the full local gate, including 401 browser tests, with 3 deliberately skipped.

The upload step had its own surprise. The deploy tool was re-uploading its own 5.1 GB build cache, about 15,000 files, on every deploy. Its archive step kept growing past 50 GB of memory until my machine ran out and I had to kill it by hand. A later retry reached about 32 GB before Claude Code stopped it. The durable fix was to stop uploading the source tree at all: build a clean checkout locally, audit the output, and upload only the finished site. That path peaks at about 290 MB, and a staged deploy now takes about a minute and a half.

Peak memory of one staging deploy: more than 50 GB, until the machine ran out of memory, when the tool archived and uploaded the source tree; about 290 MB when uploading a locally built site.
fig 03 Peak memory of one staging deploy: more than 50 GB, until the machine ran out of memory, when the tool archived and uploaded the source tree; about 290 MB when uploading a locally built site.

๐Ÿ”‘ A protected preview is its own environment

A staging URL has its own access rules, its own data and its own cache, and each one bit us.

  • Access. The test suite's bypass token had to be bound to the exact preview host. The site's media now redirects to a separate media domain, and the redesign's test transport refused to follow those redirects, so the token couldn't leak to another host. A small runner that drops secret headers on redirect fixed the tests without weakening that rule.
  • Data. Previews read their content from a staging copy that nothing kept current. The first full staging run after the deploy fix failed 95 of 889 checks, all traced to stale staging data. The fix was to publish current content to staging before verifying, which led to the incident below.
  • Cache. One diagnostic request that Claude Code made before the staging data was published got cached at the edge for up to a day, so the preview kept serving the stale answer. The fix we used was a fresh deploy. The rule now: never probe a preview route before publishing its data.

The last full staging run in that fix cycle, on the build with the "View all" fix, passed 2,475 checks with none failing. The final fix, the enlarged plus sign, passed the 343-test local browser gate and was redeployed without another full staging run.

๐Ÿ—„๏ธ The night staging broke production writes

About seven staging publishes in one day filled the database that production shares, and for about two hours every write to it failed, production included. Here is what happened.

The site stores published pages and feed variants in an Upstash Redis database. Production and staging are separate namespaces, prod: and stg:, inside one database capped at 256 MB. Each full staging publish writes a complete new generation of pages, blog and feed data. The busiest series alone has about 43 scheduled feed variants, so each publish adds an estimated 20 to 30 MB. Old staging generations are not removed automatically. On September 11, Claude Code published staging about seven times, once per batch of fixes, so I could test each one.

At about 05:00 UTC on September 12, writes began failing with ERR DB capacity quota exceeded. Threshold: 268435456 bytes, Usage: 269174008 bytes: the database had gone just past its 256 MB quota. A full database refuses every write, and the production namespace was in the same database. Production's page-view counters, publishes and scheduled-release activations could not write. The publishing tool reported only HTTP Error 400: Bad Request. Claude Code found the cause by printing the database's own error message.

Database size during the incident: about seven staging publishes push it just past the 256 MB cap at 05:00 UTC; the approved staging cleanup brings it to 158.5 MB; one more publish raises it to 180.5 MB; a second cleanup ends at 155.3 MB.
fig 04 Database size during the incident: about seven staging publishes push it just past the 256 MB cap at 05:00 UTC; the approved staging cleanup brings it to 158.5 MB; one more publish raises it to 180.5 MB; a second cleanup ends at 155.3 MB.

Claude Code stopped and asked me how to proceed. A cleanup tool existed, but it deletes data from a shared database, and that needs my approval. The dry run found 2,229 staging records no longer referenced by anything, against 437 still in use. I approved the staging-only cleanup. The cleanup itself kept failing on connection timeouts, because the client opens a new encrypted connection for every command, so it ran in repeated passes. Deletions from interrupted passes still stood. By about 07:05 UTC writes were working again, and the database was at 158.5 MB. After the next publish it was at 180.5 MB; a second cleanup finished at 155.3 MB. Claude Code has run a cleanup after each staging publish since then. No production data and no media were touched.

The impact, as precisely as I can state it: production writes failed for about two hours. No scheduled episode release fell in that window. Page-view counts from those two hours were probably lost, and I haven't measured how many.

๐Ÿงญ What I changed, and what I'd tell you

Separate namespaces isolate names, not capacity; a shared database needs a capacity check, a cleanup and error messages that say what went wrong.

  • Clean up after every staging publish. This is now a written rule in the agents' notes. Making the publish tool do it automatically is still a follow-up.
  • Check headroom before publishing. Refuse a publish when the free space is smaller than one publish needs. This is still a follow-up.
  • Surface the real error. The publishing tool should print the database's error body, not a bare 400. Also a follow-up.
  • Give staging its own database if it will ever publish at this rate. That is the real fix, and I haven't made it yet.
  • Deploy from a clean checkout, early. Your machine hides dependencies your deploy will find.
  • Keep every failed log. Six failures turned out to be four problems.

๐Ÿงช What staging found before launch

The final staging failures were explained without finding a redesign regression: 12 of the 15 came from staging data mismatches, and the other three were TLS handshake timeouts from the build box. The artifact namespace still carried the old "Writing" label and included staging-only posts. The three timeouts cleared on rerun. That left the final suite at 2,380 passes and 15 explained failures before launch.

Staging had already exposed two other boundaries. The blog artifact builder's media base URL defaults to production even when its target is staging, so the first preview pointed at the production media path and had to be rebuilt with the staging base. The same builder rejects links to /projects/ and /about/ because its boundary render has no site pages; for the staged copy, I made those links absolute. That gap was still open at launch, but the series posts were not published, so it did not block the site.

๐Ÿš€ September 16: launch was a push, not a promotion

The safety rule held, but the launch mechanism was different from the staged-to-promote path I had built: I gave the command, and pushing main triggered the production deploy. The prebuilt preview could not be promoted because it was recorded as non-promotable: its preview token is a production-only secret. On this project, the hosting provider's git build moved the canonical aliases on its own.

That deployed the site's chrome, not its current content. Podcasts, series, episodes, the Blog index, feeds and the sitemap are served from published artifacts, so after the push they had to be rebuilt in the new design and published to the production namespace. It took three attempts. The first stopped because the production page builder fails closed on its update-record inventory and needed credentials the staging builders had not needed. The second stopped while constructing the cache purger because VERCEL_TOKEN had been stripped; that happens before phase 0, so nothing was written. The third succeeded, with Upstash TLS handshake timeouts retried under the same request ID.

Deployment verification then found an incomplete production page-route set: the redesign had added /about/, /projects/ and the /zh/ routes. I seeded 143 routes, bringing the set to 503, and verification passed with live content matching the local public/ output. The gated production-preview surface stayed unpublished because its media shadow had never been bootstrapped in R2; the public site was unaffected. The production content publish moved Upstash from 156 to 177 MB. Superseded production generations are not removed automatically, so production cleanup still requires my approval.

The last part of the series adds it all up: where 6.26 billion tokens went.

Does your staging environment share anything with production that could fill up?

How this post was made: drafted by Claude Code from my own session logs, checked against those logs by a separate Codex agent, and edited by me. The cover illustration is AI-generated; the charts are drawn from the numbers above.