Evals and accuracy
Benchmarks stopped predicting which model works. Here is what we use instead.
- Among frontier models, public benchmark scores no longer tell you which one will hold up in your work.
- A score describes a setup, not a model. Effort settings, safety filters, graders and test versions all move the number.
- Gut feel is no fix either. Developers in one trial believed they had been 20% faster, and took 19% longer.
- Choose with 20 to 50 real tasks from your own work, run several times, and count cost per completed task.
Among frontier models, public benchmark scores no longer tell you which one will hold up in your work. Choose by leaderboard rank and you pay for gains that may never reach your tasks, while missing the failures that only your tasks reveal. We choose models with a test built from our own work, and we think every team running AI in production should too.
Vendors have started to say so. Anthropic's launch page for Claude Opus 5.5, on 22 September 2026, says the model leads on Anthropic's benchmarks, then adds that "at these levels of capability we've found that benchmark margins have become a less reliable guide to real-world differences."
Seven months of rising scores
After Claude Sonnet 4.6 on 17 February 2026, Anthropic released seven generally available models in seven months, ending with Opus 5.5. Each launch page reported better results than the model before it. These are Anthropic's claims about its own models.
Seven launches in seven months, in Anthropic's words
16 AprilOpus 4.7
"a notable improvement on Opus 4.6 in advanced software engineering"
28 MayOpus 4.8
With "improvements across benchmarks"
9 JuneFable 5
"state-of-the-art on nearly all tested benchmarks"
30 JuneSonnet 5
"close to that of Opus 4.8, but at lower prices"
24 JulyOpus 5
"close to the frontier intelligence of Claude Fable 5 at half the price"
1 SeptemberFable 5.1
Among "the world's most advanced models for coding and knowledge work"
22 SeptemberOpus 5.5
"performs at the level of Claude Fable 5.1 on most work and costs 40% less to run than Opus 5"
Each launch meant deciding whether to switch. The scores always pointed to the newer model, and could not say whether it would finish your tasks.
A score describes a setup, not a model
Most readers take a score as a property of the model. We disagree. A score belongs to a setup: the model, its thinking time, the machine, the safety filters, the grader and the test version. Change any of them and the number moves.
- Opus 5.5's published results use "max effort", the setting that lets it think longest, unless noted otherwise. Its API default is medium.
- Anthropic found that sandbox resources alone moved Terminal-Bench 2.0 scores by 6 percentage points, and says gaps "below 3 percentage points deserve skepticism" until setups are matched.
- Fable 5.1 and Mythos 5.1 are "the same underlying model" but scored differently on Terminal-Bench 4.0, because safety filters stopped Fable 5.1 on some tasks.
- Terminal-Bench went from version 2.0 on the Sonnet 4.6 page to 4.0 on the Opus 5.5 page, so scores from different launches cannot be compared directly.
None of this is hidden. It sits in footnotes and engineering posts that leaderboards leave out.
Why public scores drift from real work
The anthropologist Marilyn Strathern summed up Goodhart's law, named after the economist Charles Goodhart, in 1997. "When a measure becomes a target, it ceases to be a good measure." Public leaderboards became targets.
Leaderboards get tuned for. The Leaderboard Illusion, a 2025 study of the crowd-voted Chatbot Arena, now LMArena, found that a few providers test private variants and can publish only the best score. It counts 27 from Meta before Llama 4. LMArena disputes several of its numbers.
Meta's launch post for Llama 4 Maverick cited an LMArena score from "an experimental chat version", which LMArena said was "a customized model to optimize for human preference". The released model ranked below months-old rivals.
Tests leak into training data, which is called contamination. OpenAI stopped reporting SWE-bench Verified in February 2026, after finding that every frontier model it tested could reproduce original fixes or problem text for some tasks. In an Anthropic BrowseComp run, Opus 4.6 worked out that it was being tested, then "located and decrypted the answer key."
Tests saturate. Top models all score near the ceiling, and the test stops separating them. Stanford's AI Index 2025 lists MMLU, a widely used knowledge test, among the saturated benchmarks, and a re-check of MMLU estimated that 6.49% of its questions contain errors. When one question in 15 is wrong, a two-point gap tells you very little.
Tests check something narrower than the job. In a March 2026 METR study, maintainers reviewed 296 AI-written pull requests for their own projects. They would not have merged roughly half of those that passed SWE-bench Verified's tests, although the agents got no chance to revise.
We do not read any of this as one company cheating. Any public number that drives sales gets optimized for, and nothing makes that carry over to your tasks.
Gut feel is not the fix
The tempting alternative is to try each new model for a week and trust your impression. We disagree.
Expected and felt faster, measured slower
| Developers' estimate | Measured | |
|---|---|---|
| Expected before the trial | -24% | +19% |
| Believed afterwards | -20% | +19% |
METR now marks those results as out of date for current tools. Our lesson is about measurement. The developers' own estimate was off by nearly 40 percentage points.
An impression tells you where to test. It is not the result.
What independent testers found on Opus 5.5
Opus 5.5 gets the same test, and the independent results are mixed.
Four independent looks at Opus 5.5
| METR | Artificial Analysis | Vals AI | BullshitBench | |
|---|---|---|---|---|
| What it found | "an incremental improvement above Fable 5.1 on our quantitative evaluations, rather than a discontinuous jump" | First on its intelligence index | Effectively tied with Sonnet 5.5 | Pushed back clearly on 62% to 63% of nonsense prompts; the test checks whether a model calls them out |
| The catch | METR tested it before release, and Anthropic could review and edit that summary | "level with Opus 5 on cost per task" at max effort, so the 40% saving depends on the effort setting | Sonnet 5.5 costs about two-thirds as much per test | Opus 4.8 pushed back on 95% |
| Source | METR summary | Artificial Analysis | Vals Index | BullshitBench |
METR
- What it found
- "an incremental improvement above Fable 5.1 on our quantitative evaluations, rather than a discontinuous jump"
- The catch
- METR tested it before release, and Anthropic could review and edit that summary
- Source
- METR summary
Artificial Analysis
- What it found
- First on its intelligence index
- The catch
- "level with Opus 5 on cost per task" at max effort, so the 40% saving depends on the effort setting
- Source
- Artificial Analysis
Vals AI
- What it found
- Effectively tied with Sonnet 5.5
- The catch
- Sonnet 5.5 costs about two-thirds as much per test
- Source
- Vals Index
BullshitBench
- What it found
- Pushed back clearly on 62% to 63% of nonsense prompts; the test checks whether a model calls them out
- The catch
- Opus 4.8 pushed back on 95%
- Source
- BullshitBench
Anthropic's own system card lists regressions too, such as being "more likely to follow malicious instructions planted in text a user pastes into their own prompt." On public numbers, Opus 5.5 is one of several models near the top, and the order depends on the test, the effort setting and how fallbacks are counted.
What predicts real performance
We look at five things instead.
Five things to measure instead
Your own eval set
A fixed list of real tasks from your own work, failures included, each with a written definition of done and scored the same way every time. Kapoor and colleagues, in AI Agents That Matter, found that benchmarks built for model developers get mistaken for what application teams need.
Cost per completed task
Divide everything you spent, including retries and the time a person spent fixing output, by the tasks that passed review. Kapoor and colleagues also found that "for substantially similar accuracy, the cost can differ by almost two orders of magnitude."
Consistency across repeated runs
Run every task several times and count it as passed only if it passes every time, which tau-bench calls pass^k. A model that succeeds 75% of the time passes three runs in a row only about 42% of the time. Princeton's HAL reliability study found that "recent capability gains have yielded only small improvements in reliability."
Long tasks
METR's time horizon is the length of task, measured by how long it takes a skilled person, that a model completes with a 50% success rate. METR found 80% horizons four to six times shorter. Run your longest real tasks and note where each model stops, loops or drifts.
Behavior when wrong
Count confident wrong answers separately from "I don't know". In OpenAI's SimpleQA example, o4-mini scored 24% accuracy with a 75% error rate. gpt-5-thinking-mini scored 22% with a 26% error rate, because it declined to answer 52% of the time. An accuracy leaderboard ranks the first one higher.
Price per token misleads. Opus 4.7 and Sonnet 5 brought a new tokenizer, the step that splits text into the billable units called tokens, and Anthropic says "the same input can map to more tokens", up to 1.35 times as many. Claude 4.7 and later models also reject non-default temperatures, so measure the spread across runs.
Agents can also oversell their work. In an Epoch AI study from 7 October 2026, Claude Fable 5 and GPT-5.6 Sol claimed better results than their methods achieved, after picking the best of several runs. Anthropic also says Opus 5.5 "often suspects it is being evaluated", one more reason to test on real work.
How we choose a model now
Any team can copy this method.
Choosing a model on your own work
Collect 20 to 50 real tasks from recent work, including failures
Anthropic's eval guide calls that a great start. With 30 tasks, one task moves the score about three points, so ignore smaller gaps.
Write down what done means for each task
Name the person who checks it.
Run each candidate at the effort setting you would pay for in production
Run it at least three times per task.
Record four numbers per model
Tasks passed in every run, cost per completed task, confident wrong answers, and where long tasks broke.
Choose the cheapest model that clears the bar, and keep a second that also clears it
Fable 5 was unavailable to every user for 19 days after US export controls, and Opus 5.5 hands most cybersecurity tasks to Opus 4.8.
Re-run the set on every model release and every prompt change
Pin the model version in production.
Start with last month's work
If you are choosing a model this quarter, pull 30 tasks from last month's work and write down what done means for each. At extendfuture, this is the evidence we ask for before any AI system goes live: accuracy as a number, cost per outcome, and a person on the calls that matter. If you want help building the set, talk to us.
Sources
Anthropic, vendor claims about its own models:
- Launch pages: Sonnet 4.6, Opus 4.7, Opus 4.8, Fable 5 and Mythos 5, Redeploying Fable 5, Sonnet 5, Opus 5, Fable 5.1 and Mythos 5.1, Opus 5.5
- Claude docs: models overview, model deprecations and the Fable 5.1 overview, which gives its 1 September release date
- Claude Opus 5.5 system card, 22 September 2026
- Anthropic engineering: Quantifying infrastructure noise in agentic coding evals, 5 February 2026; Demystifying evals for AI agents, 9 January 2026; Eval awareness in Claude Opus 4.6's BrowseComp performance, 6 March 2026
Independent evaluations of Opus 5.5:
- METR, Claude Opus 5.5 summary, 22 September 2026
- Artificial Analysis, Claude Opus 5.5 takes the top spot on the Artificial Analysis Intelligence Index, 22 September 2026
- Vals AI, Vals Index and Opus 5.5 model page
- Peter Gostev, BullshitBench, results updated 29 September 2026
How public scores mislead:
- Marilyn Strathern, "Improving ratings": audit in the British University system, European Review, 1997
- Singh and colleagues, The Leaderboard Illusion, April 2025, and LMArena's response, May 2025
- Llama 4 Maverick: Meta's launch post, 5 April 2025; The Verge, 8 April 2025; TechCrunch, 11 April 2025
- OpenAI, Why we no longer evaluate SWE-bench Verified, 23 February 2026
- Liang, Garg and Moghaddam, The SWE-Bench Illusion, 2025, on memorization in SWE-bench Verified
- METR, Many SWE-bench-passing PRs would not be merged into main, 10 March 2026
- Stanford HAI, AI Index 2025, technical performance
- Gema and colleagues, Are We Done with MMLU?, 2024, revised January 2025
What to measure instead:
- Becker, Rush, Barnes and Rein, METR, Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity, July 2025, and METR's blog post, which now marks the results out of date
- Kapoor and colleagues, AI Agents That Matter, July 2024
- Yao and colleagues, tau-bench, June 2024
- Princeton HAL, agent reliability study
- Kwa and colleagues, METR, Measuring AI Ability to Complete Long Software Tasks, March 2025
- OpenAI, Why language models hallucinate, 5 September 2025
- Epoch AI, Can AI automate AI R&D yet? Early evidence from InnovationEval, 7 October 2026
Related posts: why errors compound in agent loops, cost, not intelligence, is the bottleneck and the agent is a compiler.