← Blog

Spotify's 90% token win is a 0.6% rounding error on my logs

Measurements of file reads and answer completeness show why token-saving hooks should follow workload evidence rather than a headline benchmark.

🚨 Two viral posts said Spotify cut Claude Code tokens by 90% by blocking big file reads; on my own logs that hook would save 0.6%. The villain is a workflow: installing a token-saving hook because a screenshot said so, without measuring what it touches. It burns the one thing a small team cannot buy back, engineering attention, and it can make the agent worse. Measure before you adopt; this article shows what to measure. That discipline protects you from importing somebody else's benchmark, workload, and trade-off into your team.

All the code and data behind these numbers are public: shared-code on GitHub.

Viral read-hook claims compared with measurements from the author’s logs
fig 01 Viral read-hook claims compared with measurements from the author’s logs

🚨 86% of steps, 2% of fresh input

Agents spend most of their steps reading, and that reading is a rounding error on what I pay for.

Different denominators: read actions are 86% of steps, but file content is 2.1% of fresh input in the measured Claude Code sessions
fig 02 Different denominators: read actions are 86% of steps, but file content is 2.1% of fresh input in the measured Claude Code sessions

The podcast claim, from Alexander Whedon of Subquadratic, is that 86% of a frontier agent's steps are read steps. Fine. Steps. I counted tokens instead across my own logs: 2.1% of fresh input on Claude Code and 6.1% on Codex. Background review is a third, but a 350-line block reaches only about 3% there because Codex slices; codex exec has no hook to install anyway. That is where a summary is most dangerous, because a reviewer is paid to notice what nobody asked about.

Content already read also comes back every turn, filling the window and bringing compaction forward. Cheap re-sent tokens are not a free context window.

Here is the line that will annoy people: if you installed a read-blocking hook this week because of a screenshot, you optimised a rounding error and added a 10 to 30 second stall to every large file. Count your own read share, per stream, before you touch anything. A benchmark can tell you where to look. It cannot tell you which stream deserves a new rule.

🔍 Codex read faster and got 2 of 6 fully right. Claude Code got 6 of 6

On identical questions, fast cheap reading dropped lines and lost answers; slow complete reading cost up to $1.12 a question and got everything right.

Six trials comparing Codex and Claude Code response times and answer completeness
fig 03 Six trials comparing Codex and Claude Code response times and answer completeness

Three read-heavy questions, two runs each, Claude Code 2.1.266 versus Codex CLI 0.153.4. Codex ran gpt-5.3-codex-spark; Claude ran claude-fable-5-1 in five runs and claude-opus-5 in one. This is about how the harnesses read, not which model is smarter. Claude Code caps output too: about 1 in 150 Read results and 1 in 200 shell outputs in my logs.

Codex's output cap chopped a 797-line file to 491 lines and it returned 6 of 12 methods. Twice. Claude Code's 1,118-line whole read came back intact; that run cost $1.12 in API-equivalent terms. Fast reading that drops lines is a bug with a good latency number. Time it and check the answer on your own repo. A faster trace is only useful if the answer still covers the work the reader asked for.

⚙️ What Mazmanov built is better than what went viral

The 90% is real and narrow: it is how much file text the main model stopped seeing on three reading tests, with no cost, task, or correctness number, and the one quality result is a missed bug.

Spotify’s 90% measures fewer tokens seen by the main model, not task success or total spend
fig 04 Spotify’s 90% measures fewer tokens seen by the main model, not task success or total spend

Dimitri Mazmanov's post describes a hook that blocks whole reads over 350 lines and sends them to Gemini 2.5 Flash for bullets. The README benchmark reports 82% to 94% fewer tokens in three read scenarios. The worker missed a subtle thread-safety bug; the 277-point Hacker News thread said, “No mention of correctness or task success rate.”

What is right in it: enforcement beats instructions, because CLAUDE.md rules decay over a long session and a hook can say no; a smaller window helps even when re-sent tokens are cheap, because compaction arrives later; and writing code straight to disk attacks output tokens, the expensive category.

The rules did change. Prefix caching made the first read expensive, and a ten-times-cheaper reader made delegation possible. That turned read cost into a per-stream question; viral posts turned it back into one global constant. Ask what the headline number measures before you copy the setup. A headline can describe a real result and still be the wrong decision rule for your machine.

🛠️ Install the hook when your logs say so

Adopt the hook on an unattended stream where big reads are more than a tenth of what you pay for; skip it everywhere else.

Conditions for testing a read hook in interactive work and unattended background review
fig 05 Conditions for testing a read hook in interactive work and unattended background review

Reads must rarely become edits, the worker must be much cheaper or local, and you check sampled answers against the full file. Skip interactive sessions, dense files, review, and debugging, where the whole point is noticing the unrequested line. Treat each check as a number to pull from your logs first. If you cannot pull it, you are guessing about the place where the hook will hurt.

🧭 Your threshold is in your logs, not in a screenshot

One constant cannot fit two harnesses whose typical read differs eight-fold, and it is certainly not tuned to yours.

Median reads of 28 lines for Claude Code and 230 for Codex, versus Spotify’s 350-line threshold
fig 06 Median reads of 28 lines for Claude Code and 230 for Codex, versus Spotify’s 350-line threshold

This week, use the shared-code on GitHub to count file-content bytes against your harness counters, split interactive from background, and look at whole reads over a threshold as a share of fresh input. I have argued before that input-token cost lives elsewhere: Opus 4.7 uses 58% more input tokens than GPT-5.5 and the claude-mem token-burn post. That is the trap: a clean chart turns an implementation choice into an identity test, and nobody asks whether the saving is large enough to matter before anyone writes the hook.