← Blog

GPT-6 Astra asked for a second computer. Opus 5 asked for a reboot

Two coding-agent sessions responded differently to the same crash evidence, showing how diagnostic judgment and a clear human handoff can outweigh elaborate orchestration.

🔴 Seven crashed deploys. Twelve segfaults in the kernel log. One CPU core with its name on every one of them. Codex on GPT-6 Astra read that log, wrote "hypothesis, not proof," and asked me for a GitHub Actions runner. Fifty-eight minutes later Claude Code on Opus 5 read the same machine, ran two commands nobody had bothered to run in four days, and told me to type sudo apt install intel-microcode. The rules just changed on what "blocked" is allowed to mean. 🧭

⚙️ The bug: one flipped bit, seven costumes

A CPU on stale microcode was flipping bits in CPython's inline caches, and every "different" crash in four days was that single fault reading the wrong memory slot. ⚙️ Once you see it, the whole log collapses into one line.

Here is what it looked like from the outside. My site ships staged-then-promote. Nothing hits the host until a preflight gate is green: build, validation, a few hundred Python tests, a few hundred browser checks. Six preview attempts between T+3m and T+1h30m. Exit 134 from a corrupted malloc free list under an integer subtraction. A TypeError saying a string was an int, then a segfault inside the garbage collector. A nested build that died with its stderr swallowed. Conda Python 3.12.2 segfaulting in the same YAML reader as Homebrew Python 3.14.4. A real argparse bug. And a spawn uv ENOENT that the orchestrator inflicted on itself by building a "controlled" PATH without the directory uv lives in.

The same session had been hitting exit 139 since T−4 days. Five times. It had already tried a different Python. ⚠️

Seven failed pre-upload deployment attempts and recorded Python crashes
fig 01 Seven failed pre-upload deployment attempts and recorded Python crashes

Here is what the crash was. PyYAML's Reader keeps buffer, a string, next to pointer, an int. Every failure said buffer was an int. No Python code can do that. Since 3.11, CPython's specializing interpreter (PEP 659) keeps a specialized LOAD_ATTR's inline cache in the entries immediately following the instruction; one flipped bit and the read lands on the neighbour. CPython's own name for it opened Claude's session: Fatal Python error: Executing a cache. Both 3.12 and 3.14 have inline caches. That is why the interpreter swap did nothing.

Evidence of an inline-cache read returning the adjacent attribute
fig 02 Evidence of an inline-cache read returning the adjacent attribute

🚨 Codex had the answer on screen and filed it under "unproven"

Codex found the CPU 8 pattern and the 0x10f microcode, then its own process ate the finding and spat out a request for a different computer. 🔍 This is the part that should sting if you run agents in production.

The orchestrator, from T (the start of the first preview attempt, which every time below is measured from), was Codex CLI 0.153.4 on gpt-6-astra, reasoning effort high, with sub-agents on gpt-5.6-sol, gpt-5.6-terra, and gpt-5.6-luna. Every model name here is lifted from the transcripts, not remembered. T+1h41m, verbatim: "eight recorded Python segfaults over several days all identify logical CPU 8. This machine is an i9-13900KF with an old BIOS and microcode 0x10f; Intel currently recommends BIOS microcode 0x12F or later." It even linked Intel's support notice, which recommends a BIOS with microcode 0x12F or later.

Then the process ate the finding:

  • A sub-agent was sent to compare CPU 8 against CPU 16. A provider safety check blocked it twice. No result, honestly reported, honestly abandoned.

  • The diagnosis written an hour earlier, "not evidence for hardware failure," was never retracted. The new lead got appended as a "serious hypothesis."

  • The filed report imagined exactly one remedy, a BIOS flash, and declined it in writing: "no flash instruction is given here."

  • So the ask became infrastructure. Which second machine may I use? May I create a GitHub Actions workflow? Goal marked blocked at T+2h26m.

It never ran dpkg -l intel-microcode. Ubuntu ships Intel's fix as a package that loads at boot. No flash. One reboot. Codex had the CPU model, the current microcode, the recommended microcode, and the one-core pattern, and it turned all four into a request for a different computer. Then it spawned Claude Code as a sub-process with "No BIOS/microcode/voltage/clock changes, reboot" written into the brief, and got back the same conclusion, because it had asked for it. At T+5h22m, attempt seven. Exit 139. Same question.

Codex records CPU 8 and stale microcode, then requests another machine
fig 03 Codex records CPU 8 and stale microcode, then requests another machine

🔍 Opus 5, 58 minutes, two commands

Claude took Codex's finding, ruled out software in four commands, ran dpkg, and found the one fact four days of notes had missed: the microcode package had been installed and then removed. 🎯

At T+5h25m I opened a bare Claude Code session and typed "deploy website to staging." The first three bookkeeping steps, including issuing that first deploy, ran on claude-sonnet-5; every step from the first failure on ran on claude-opus-5. Attempt one hit a stray zip. Attempt two hit the cache abort. A rerun produced the 'int' error; a third run passed. The model's line: "This is a bug, not a flake I should paper over."

Fourteen minutes in, it found Codex's session logs on disk, read them, and said so: "Prior sessions have been fighting this exact crash for two days. Let me check the CPU/microcode evidence they flagged." Let me be exact about credit, because the point of this article depends on it: Codex found the CPU 8 pattern and the 0x10f microcode. Claude inherited that finding and asked the next question. 📌

Then it did what a good on-call engineer does at 2 a.m. Eliminate software first. No C loader. Standard GIL build, no JIT. 50 GB free, no OOM. No monkeypatching. Then the two commands:

$ grep -m1 microcode /proc/cpuinfo
microcode  : 0x10f
$ dpkg -l intel-microcode
rc  intel-microcode  ...   ← removed, config-files only

rc. Installed once, removed later. The CPU was on bare, years-old firmware, before the entire fix chain Intel shipped for what it called elevated core voltages causing Vmin shift during idle and light activity, ending in 0x12B in 2024 and 0x12F in 2025 for machines that sit lightly loaded for days. That is my workstation with two idle agents parked on it. 🎯

Attempt three crashed in a different test file. The model named the signature, Reader.buffer reading back as Reader.pointer, and then it stopped: "I don't want to keep spending 25-minute coin flips on your behalf."

Claude checks the same crash evidence and the installed microcode package
fig 04 Claude checks the same crash evidence and the installed microcode package

🛠️ Neither agent could fix it. Only one knew what to do about that

The fix needed root, a reboot, and my password, which no agent should have, and the whole difference between the two harnesses is what each did at that wall. 🤝

Root needs my password, which does not belong in any transcript. The reboot kills the agent's own session and every other process on the box, including two Codex instances that had been up for a day. That is a human decision by construction, and I am glad both harnesses treat it that way.

Codex looked for a way around it. Claude made the handoff cheap: three options with costs, sudo -n tried and reported, the session log and a memory note written before the reboot, then two commands, a verification line, the rollback, and the caveat I would have written myself: "The microcode update is the experiment, not a confirmed fix." When I tried to run sudo from inside the session and it choked on the missing TTY, it explained why and refused to suggest piping my password through stdin.

Human handoff for the root command and reboot
fig 05 Human handoff for the root command and reboot

📊 After the reboot

One package install bought the first clean full gate since the fight began, and one recurrence during the post-boot storm keeps me from calling it proof. ✅

Microcode went 0x10f to 0x133, per the kernel, loaded from version 3.20260210 of Ubuntu's intel-microcode package. Gate run one died because the reboot had stopped the test database. Run two threw the 'int' signature once, under a load-15 post-boot storm, in a file that then passed five of five. Run three, calm system: exit 0, zero tracebacks, zero segfaults. The staging deploy then refused because my local commit was not pushed yet, which is the guard doing its job.

The session said so itself: one green run against an intermittent fault is evidence, and a miscomputing host can flip a byte silently as easily as it crashes loudly. But it is the only experiment in four days that changed the result, and it cost one package.

Post-reboot microcode version and validation results
fig 06 Post-reboot microcode version and validation results

📊 The bill for not solving the problem

About 30x more tokens on the side that did not find the bug, and most of them went to Codex reviewing Codex. 📉

Subscriptions on both sides, so nothing here is cash. These are API-equivalent numbers at list price, computed from the raw usage records. Codex, from its first preview attempt at T to the moment I gave up on it at T+5h25m, priced models only:

Model API calls Uncached input Cached input Output Cost
gpt-6-astra 786 2.66M 90.6M 341k $134.24
gpt-5.6-sol 114 0.60M 9.9M 51k $7.37
gpt-5.6-terra 125 0.39M 11.6M 77k $4.03
gpt-5.6-luna 77 0.15M 4.6M 32k $0.16
Total 1,102 3.80M 116.7M 501k $145.80
Model-by-model API-equivalent cost and token accounting
fig 07 Model-by-model API-equivalent cost and token accounting

Plus 588 calls and 66 million input tokens to an approvals reviewer logged as codex-auto-review, which OpenAI does not price, so I am not pricing it either. Rates from OpenAI's pricing page.

Claude Code on Opus 5, first attempt to diagnosis, 20 minutes: 47 messages, 5.34M cache reads, 121k cache writes, 27k output. $4.57. Through the finished handoff at 58 minutes: $5.31. Rates from Anthropic's pricing page, and the arithmetic reproduces the $5.46 that Claude Code's own telemetry reported for the review Codex spawned.

In fairness: through its own root-cause report at T+2h26m, Codex was at $71.13, and it shipped three reviewed commits that day. None of them touched the problem I had.

🧭 Process bought filings. Judgment bought the fix

Both models knew what Raptor Lake instability is; the gap was three decisions about what to do with evidence, and no amount of review process substitutes for them. 🧭

GPT-6 Astra is a top-tier coding model and Opus 5 is Anthropic's current Opus-tier model, not a magic one. When the symptom is impossible in software, stop looking in software. Ask the cheapest decisive question before asking for more hardware. And when you hit the one thing only a human can do, make the handoff two commands, not a workflow proposal.

Timeline of the Codex and Claude investigations and microcode update
fig 08 Timeline of the Codex and Claude investigations and microcode update
Cartoon comparing five hours and $146 with twenty minutes and $4.57
fig 09 Cartoon comparing five hours and $146 with twenty minutes and $4.57

I have argued that the real AI coding race is about encoded judgment and that encoded judgment is the moat. Codex did the finding. Claude did the deciding. This week I watched a harness with more process than any I have used, finders, adversaries, referees, an approvals reviewer, evidence filing, produce 66 million reviewer tokens and a blocked ticket. Judgment produced dpkg -l.