A high score is not a good score. What a year of AI evaluation taught me
Useful AI evaluation compares versions under consistent conditions, investigates failure patterns, and continually strengthens tasks, harnesses, and acceptance criteria.
TL;DR
-
🚨 A 90%+ score is often a warning sign: your eval may be getting easier.
-
📈 A +5 point gain matters more than a 95% score in isolation.
-
🧪 Compare model versions in the same harness, not against fixed score targets.
-
🔍 Failure patterns matter more than the final headline score.
-
⚙️ Evaluation is a quality system: model quality and harness quality must co-evolve.
-
🏁 No single model wins every benchmark.
Over the last year, I've spent a lot of time evaluating production GenAI products and running side-by-side comparisons between new frontier models and existing baselines. That first-hand work changed how I read benchmark numbers.
I still believe a five-point improvement can be meaningful. But the meaning only exists in context. A +5 from model version A to B, under the same tasks, same harness, and same scoring policy, can indicate real progress. A standalone 93% or 97% score without that context tells me much less.
That's the core thesis:
Absolute score alone doesn't matter.
Comparisons matter. Version-to-version movement matters. Failure pattern movement matters. Chasing a threshold in isolation is usually how teams confuse comfort with quality.
📈 Keep relative progress, drop threshold worship
In many teams, evaluation gets reduced to a single objective: "get above X%." That sounds practical, but it often creates the wrong behavior. Once people optimize to hit a static threshold, they tune around the measurement instead of improving the system.
When I compare models in practice, I care less about whether one model is "high" and more about what changed versus the prior version under identical conditions. Same prompts, same tools, same retry policy, same judge criteria. If a newer model improves materially in that setup, that's useful evidence. If it only looks good on a separate benchmark with different assumptions, confidence should stay low.
This is consistent with broader evaluation guidance from HELM, which argued for multidimensional reporting over one-number simplification because model behavior differs meaningfully across settings. The technical performance chapter of the 2026 Stanford AI Index made the point even more sharply: "Evaluations intended to be challenging for years are saturated in months, compressing the window in which benchmarks remain useful for tracking progress."
🚨 Very high scores are often warnings, not celebrations
A 90%+ score can look like victory.
Operationally, it's often a diagnostic alarm. 🚨
When a test is that easy, one of two things is usually happening: either the model is truly overqualified for that task, or the evaluation is no longer discriminative enough for the decisions you're making. In production, the second case is common. Data gets stale. Criteria go soft. Edge cases disappear. The metric climbs while real-world risk remains.
MMLU, once the standard general-knowledge benchmark, is now saturated above 88% for frontier models and no longer differentiates leading systems cleanly. Contamination and benchmark saturation are now central concerns. If benchmark content or close variants leak into training data, top-line scores can inflate while true generalization is unchanged, as shown in Investigating Data Contamination in Modern Benchmarks for Large Language Models.
The response has been a new generation of refreshed, contamination-aware benchmarks. LiveBench adds new questions monthly. LiveCodeBench sources fresh coding problems from LeetCode, Codeforces, and AtCoder. LiveCodeBench Pro pushes harder with Olympiad-medalist judgment.
Humanity's Last Exam was designed as a harder replacement for saturated academic benchmarks. Yet even HLE went from 8% top score in January 2025 to over 50% by April 2026 — a reminder that today's hard benchmark can become tomorrow's leaderboard theater.
So yes, celebrate progress.
But when scores approach ceiling: celebrate less, investigate more.
🔍 The main output of eval is failure discovery, not score reporting
A lot of teams treat evaluation as a reporting pipeline: produce a number, ship a dashboard, move on. In production GenAI, that misses the highest-value output.
The highest-value output is the set of failure patterns you identify early enough to fix.
Examples from real systems are rarely just "wrong answer."
They look more like:
-
brittle tool sequences
-
overconfident false positives
-
weak recovery after partial tool failure
-
unstable behavior under slight prompt perturbations
-
silent regressions after seemingly positive model upgrades
A single scalar score won't tell you where those failure classes live.
Behavioral testing methods have made this point for years. CheckList showed that held-out accuracy can mask critical behavioral bugs, including in strong commercial systems. Dynabench pushed further by making benchmarking dynamic and adversarial, forcing evaluation to evolve as models improve.
The 2026 coding benchmark landscape makes this concrete. On SWE-bench Verified, the top six models are separated by just 1.3 percentage points — so compressed that ranking differences reflect scaffold tuning and prompt engineering more than raw model capability.
Yet the same models show dramatically different profiles on newer benchmarks. A March 2026 comparison reported Claude Opus 4.6 near the top of SWE-bench Verified while trailing GPT-5.3 Codex by 12 points on Terminal-Bench, which measures command-line agent performance. No single model wins every evaluation. If you pick a model based on one leaderboard, you may be optimizing for the benchmark's biases, not your use case.
That matches my experience directly: the eval process is valuable because it reveals what to improve next, not because it prints a bigger number.
⚙️ Evaluation is a system quality function, not only a model test
In agentic GenAI products, the harness is often as important as the model.
By harness, I mean the full execution layer:
-
prompting strategy
-
tool routing
-
fallback policy
-
retry logic
-
memory handling
-
output constraints
-
judge/scoring design
Two teams can run the same base model and get very different quality because their harnesses differ.
This is why "which model won?" is often the wrong first question. The better question is: "Which model-harness combination produced the best quality profile for our real tasks?"
You can see this dynamic in coding evaluations too. SWE-bench Verified improved reliability by human-validating instances and tightening task quality, a move further explained in OpenAI's Introducing SWE-bench Verified.
The harder SWE-bench Pro makes the harness-vs-model point even more stark. Under the SEAL standardized scaffold, Claude Opus 4.5 scores 45.9%. When the same model runs in Claude Code's custom agent framework, it jumps to 55.4%.
That's a ~10-point swing from scaffolding alone.
The lesson generalizes: measurement design choices can move outcomes as much as model swaps.
From a governance perspective, this aligns with NIST AI RMF thinking: measure is one function inside a broader loop of mapping context, governing risk, and managing ongoing change — not a one-time model score ritual. The NIST AI RMF Playbook makes the same point in more operational terms.
🧗 Make evaluation continuously harder
If your system is improving, your evaluation should get harder on purpose.
That means:
-
adding more adversarial slices
-
refreshing task sets
-
increasing ambiguity
-
tightening acceptance criteria
-
introducing new failure traps based on recent incidents
Otherwise, you get the worst pattern: score increases that reflect benchmark familiarity rather than product robustness.
The evolution of SWE-bench is the clearest example. The original 2023 benchmark asked whether models could resolve real GitHub issues. Then came SWE-bench Lite for faster iteration, SWE-bench Verified for cleaner human-validated tasks, and now SWE-bench Pro for harder, multi-language, more contamination-resistant software engineering work. SWE-bench Live pushes the same idea further by adding fresh tasks from recent GitHub activity.
That progression matters because each new variant exists to restore signal as older variants become easier, more familiar, or more contaminated. Verified was once the gold standard; now frontier models cluster near the top, and the next useful question has shifted to harder variants and cleaner harness comparisons.
The broader industry shift toward continuously updated and contamination-aware benchmarks is the same signal, as seen in LiveBench, LiveCodeBench, and SWE-bench Live. LiveBench is still actively maintained, with new model updates landing in April 2026, which is exactly what a living benchmark should look like. But internal evaluation needs the same discipline. Teams should treat eval set difficulty as a managed variable, not a fixed artifact.
When we fail to do that, we accidentally reward stagnation. When we do it well, metrics may become harder to improve — but those gains become far more trustworthy.
🛠️ A practical operating model I now use
For production decisions, I now use a five-part eval posture:
-
📊 Compare deltas, not trophies. Compare candidate models against current baselines under the same harness and policies.
-
🚨 Watch the ceiling. Once tasks saturate, refresh or raise difficulty immediately.
-
🔍 Track failure movement. Report failure classes alongside scores, and make failure movement a top-level KPI.
-
⚙️ Evaluate the harness. Test tool reliability, retry strategy, judge robustness, and orchestration — not just model output text.
-
🧗 Keep hardening. Run continuous hardening cycles so evaluation difficulty tracks system capability.
This approach is slower than posting leaderboard snapshots, but it produces decisions you can defend.
🧭 Closing view
I'm still pro-score.
I'm just anti-score-without-context.
A five-point gain can be very meaningful. A very high score can be a warning. The true value of evaluation is the quality-improvement loop it creates: identify failures, fix system behavior, re-test under harder conditions, and repeat.
In 2026, strong GenAI teams won't win by chasing isolated thresholds. They'll win by treating evaluation as an evolving quality system across model, harness, data, and criteria.
That is the difference between looking good on paper and shipping reliable products.