What it took to turn a Claude Design prototype into a buildable spec
Part 2: a design prototype carries taste, sample data and its own opinions about who wins a conflict. Here is how I turned mine into a spec agents could build against, and why every gate passing still left 20 places that didn't match.
Part 2 of 6 in "Rebuilding williamliu.ai with agents." Start with the overview.
TL;DR
- π¨ A Claude Design handoff is a working mock-up with sample data; eight of its mechanisms must never ship.
- βοΈ Decide who wins a conflict before anyone builds; mine needed a 13-row ledger.
- π Record the design file by its hash: my first research pass reviewed a stale copy.
- π« The prototype invented facts about me; the spec allowed only content I could verify.
- π§ͺ Every build gate passed, and an audit still found 20 places that didn't match the design; closing them took nine fix rounds.
- π A choice can follow the spec and still look wrong at a width nobody checked, or with content shorter than readers will see.
A design prototype tells you what the site should feel like; a spec tells agents what to build when the prototype, its own documentation and your real content disagree. I designed the new williamliu.ai in Claude Design. The handoff was excellent at the first job and silent or wrong on the second. This part covers the work between "the design looks right" and "an agent can implement this without guessing".

The blog did not yet support inline video, so the player that renders this walkthrough in place was built for it. Older renderers still show the poster as a link.

π¨ What the handoff is, and what can never ship
The handoff is a runnable mock-up of the whole site, and most of its machinery exists to make the mock-up work, not to be deployed. Claude Design exported a ZIP with a written README, a tokens.json of colors, type and spacing, ten reference screenshots, and a prototype you open in a browser. The prototype is React compiled in the page by Babel, with its content hardcoded in a data file.
My site is a static site built by Python scripts, with feeds, scheduled releases and a publishing pipeline that other things depend on. So before any code, the spec listed what must never ship:
- React compiled in the browser, and client-side hash routing.
- The hardcoded sample data file.
- Language switching by rewriting data in place.
- Simulated audio playback.
- The design tweak panel and the image-drop placeholders.
- The phone-frame preview harness.
Everything else, the look, the layout, the interactions, had to be rebuilt on the existing architecture. The spec's first rule was that the redesign changes presentation only: no change to how posts and episodes are stored, dated, scheduled or fed.
βοΈ Decide who wins before anyone builds
The handoff disagreed with itself in 13 places, and its own instructions said the prototype should win; I overruled that before a single line was written. The handoff's agent notes say: "When the README and the prototype disagree, the prototype wins." But the README carried the explicit decisions: breakpoints, accessibility, container widths, the Chinese-language requirements. The prototype encoded mock-up choices that conflicted with it.
So I approved a precedence order on day one. My explicit decisions come first, then the README, then tokens, screenshots and prototype styles where the README is silent. Real production content beats mock data every time. The spec then settled each known conflict in a ledger, one row per disagreement:

The ledger did two jobs. It stopped each agent from re-litigating the same question, and it gave reviewers something to check against. A reviewer who sees a 52-pixel title can look up that 52 was decided and move on.
This precedence order gives Hard Question of Who Owns the Agent a concrete answer here: whoever sets the order decides whose defaults the agent executes.
π Fingerprint the design file first
The first research pass analyzed the wrong design file, and its "sound" verdict had to be thrown away. There were two copies of the handoff on my machine: a 96 KB archive in my home folder and a 394 KB archive in the website repository. The research started on the stale one. An independent review of the draft spec found three serious gaps, led by missing Chinese-language behavior. The stale archive had no Chinese-language files at all; the current one did.
The fix was procedural. The spec now names the one authoritative archive by size and SHA-256 hash and declares the other stale. The research was amended to record that the earlier verdict was invalid. The spec grew from 394 lines to more than 500 before it passed review. Recording a hash is trivial; not recording it cost a full research-and-review cycle.
π« The prototype invented facts about me
The mock-up's sample data looked like my site but contained claims that were false, so the spec allowed only content that already existed and that I had approved. The prototype's projects page listed fidx as a Rust project with "1.2k" stars. fidx is written in Python, and the star counts were invented. Its About page described my career: "Earlier work spans applied ML research, large-scale training, and inference systems." I never wrote that. It also listed a fourth project, kvprobe, that wasn't among my verified projects.

The spec's rule was blunt: no mock counts, no dates, roles or stacks without a source, no biography I hadn't written. The About page shipped with a two-sentence lead that passed a brand copy review, after the first proposal was rejected for repeating the navigation. The Chinese side got the same rule: only translations supplied with the design or approved by me. It still needed enforcing later. During the fix rounds, unapproved copy had to come out twice: an introduction line on the Projects page, removed in the first fix round, and four Chinese strings that Codex added and Claude Code removed in review. Visual alignment is not a license to write copy.
π§ͺ Every gate passed; the page still didn't match
After every milestone had passed its tests and adversarial reviews, a design audit found 20 places where the site didn't match the design, 6 of them serious. The milestones were tested thoroughly: unit tests, browser tests on four screen sizes, accessibility scans, visual snapshots. But those tests checked the behavior the agents had specified, and the snapshots compared the site with its own earlier self. None of them asked whether the site looked like the design.
So once the redesign was on a staging URL, I had Claude Code audit it against the design itself. A Codex audit read the prototype and the code side by side, and a visual pass rendered both at the reference screenshots' size and at phone width. The audit's verdict: not conforming, with 20 deviations the plan didn't account for, 6 high, 12 medium, 2 low. The visual pass found more. The navigation had underlines the design didn't. Subscribe links were plain text where the design had bordered buttons. The episode sidebar was unstyled, with a long source link overflowing it. My own screenshots from the staging site added four more.

Three re-audits took the count from 20 to 6, then 2, then a single low-severity item, which the ninth fix round closed. One deviation was accepted on purpose: the design's footer links to a combined podcast feed that doesn't exist. The audit also found a bug that was already live in production: episode transcripts rendered the transcript file's header, "Generated: β¦", as if it were part of the conversation.
Coda, September 14, 2026: After all nine fix rounds and every gate were green, my screenshots of the real series posts on staging found five more places that did not match the handoff: the index row grid, the one-line dek, the article drop cap, the meta separators and the figure labels. The screenshot baselines had been captured with short fixture posts, so they had locked in a composition that looked fine only with short content. We fixed those five and added a long-content fixture. The lesson in the next section now has a second half: check at the widths readers use, with the content readers read.
π Right by the spec, wrong to the eye
A choice can follow the spec exactly and still look off at a screen width the review never used. While writing this series, I took fresh screenshots at 1280 pixels wide. The home page's content starts 124 pixels left of the logo above it. The reason is in the ledger. The README says wide pages use a 1180-pixel container, the prototype's own styles used 980 pixels everywhere, and the ledger picked the README. The navigation bar kept 980. At the reference screenshots' width of 909 pixels, both containers are wider than the screen, so both shrink to fit and line up. Every visual comparison with the design ran at 909 pixels or at phone width.

I looked at it and kept it: the wide home page is what the README asked for. The lesson is about the method. Check the design at the widths your readers use, with the content readers readβnot only the width the screenshots happen to be or the short content in a fixture.
π§ A checklist for your own design handoff
Treat the prototype as the brief and the spec as the contract, and audit the result against the design, not against your own tests.
- Hash the input. Name the one authoritative design file by size and hash before anyone reviews anything.
- Write the precedence order. Decide whether your decisions, the written notes, the tokens or the prototype wins, and write the known conflicts down.
- List what never ships. Prototype machinery, sample data and demo controls.
- Allow only real content. Every number, date, role and translation needs a source or your approval.
- Audit after the gates pass. Compare the running site with the design, at several widths, including the ones your readers use.
- Test with real-sized content. Build screenshot baselines from real posts or long-content fixtures, not only short samples.
Part 3 covers who did the building: a Codex goal session with 333 agents, then Claude Code directing Codex one fix at a time.
What does your design handoff say should win when the mock-up and the written notes disagree?
How this post was made: drafted by Claude Code from my own session logs, checked against those logs by a separate Codex agent, and edited by me. The cover illustration is AI-generated; the charts are drawn from the numbers above.