Evals and accuracy

Benchmarks stopped predicting which model works. Here is what we use instead.

  • Among frontier models, public benchmark scores no longer tell you which one will hold up in your work.
  • A score describes a setup, not a model. Effort settings, safety filters, graders and test versions all move the number.
  • Gut feel is no fix either. Developers in one trial believed they had been 20% faster, and took 19% longer.
  • Choose with 20 to 50 real tasks from your own work, run several times, and count cost per completed task.

Among frontier models, public benchmark scores no longer tell you which one will hold up in your work. Choose by leaderboard rank and you pay for gains that may never reach your tasks, while missing the failures that only your tasks reveal. We choose models with a test built from our own work, and we think every team running AI in production should too.

Vendors have started to say so. Anthropic's launch page for Claude Opus 5.5, on 22 September 2026, says the model leads on Anthropic's benchmarks, then adds that "at these levels of capability we've found that benchmark margins have become a less reliable guide to real-world differences."

Seven months of rising scores

After Claude Sonnet 4.6 on 17 February 2026, Anthropic released seven generally available models in seven months, ending with Opus 5.5. Each launch page reported better results than the model before it. These are Anthropic's claims about its own models.

Seven launches in seven months, in Anthropic's words

  1. 16 AprilOpus 4.7

    "a notable improvement on Opus 4.6 in advanced software engineering"

  2. 28 MayOpus 4.8

    With "improvements across benchmarks"

  3. 9 JuneFable 5

    "state-of-the-art on nearly all tested benchmarks"

  4. 30 JuneSonnet 5

    "close to that of Opus 4.8, but at lower prices"

  5. 24 JulyOpus 5

    "close to the frontier intelligence of Claude Fable 5 at half the price"

  6. 1 SeptemberFable 5.1

    Among "the world's most advanced models for coding and knowledge work"

  7. 22 SeptemberOpus 5.5

    "performs at the level of Claude Fable 5.1 on most work and costs 40% less to run than Opus 5"

Each launch meant deciding whether to switch. The scores always pointed to the newer model, and could not say whether it would finish your tasks.

A score describes a setup, not a model

Most readers take a score as a property of the model. We disagree. A score belongs to a setup: the model, its thinking time, the machine, the safety filters, the grader and the test version. Change any of them and the number moves.

Same model, different setup, different score

Same model, different setup, different score
Value
Opus 4.5 on CORE-Bench, before the fixes42%
Opus 4.5 on CORE-Bench, after grading fixes and a looser setup95%
Opus 5.5 on Vals AI's SRE Bench, with fallback models33.59%
Opus 5.5 on SRE Bench, fallback tasks counted as failures5.34%
A researcher fixed grading bugs and loosened the test setup, in Anthropic's eval guide. On Vals AI's SRE Bench, fallback models, which answer when Opus 5.5's filters block a task, helped with 217 of 262 tasks.
  • Opus 5.5's published results use "max effort", the setting that lets it think longest, unless noted otherwise. Its API default is medium.
  • Anthropic found that sandbox resources alone moved Terminal-Bench 2.0 scores by 6 percentage points, and says gaps "below 3 percentage points deserve skepticism" until setups are matched.
  • Fable 5.1 and Mythos 5.1 are "the same underlying model" but scored differently on Terminal-Bench 4.0, because safety filters stopped Fable 5.1 on some tasks.
  • Terminal-Bench went from version 2.0 on the Sonnet 4.6 page to 4.0 on the Opus 5.5 page, so scores from different launches cannot be compared directly.

None of this is hidden. It sits in footnotes and engineering posts that leaderboards leave out.

Why public scores drift from real work

The anthropologist Marilyn Strathern summed up Goodhart's law, named after the economist Charles Goodhart, in 1997. "When a measure becomes a target, it ceases to be a good measure." Public leaderboards became targets.

Leaderboards get tuned for. The Leaderboard Illusion, a 2025 study of the crowd-voted Chatbot Arena, now LMArena, found that a few providers test private variants and can publish only the best score. It counts 27 from Meta before Llama 4. LMArena disputes several of its numbers.

Meta's launch post for Llama 4 Maverick cited an LMArena score from "an experimental chat version", which LMArena said was "a customized model to optimize for human preference". The released model ranked below months-old rivals.

Tests leak into training data, which is called contamination. OpenAI stopped reporting SWE-bench Verified in February 2026, after finding that every frontier model it tested could reproduce original fixes or problem text for some tasks. In an Anthropic BrowseComp run, Opus 4.6 worked out that it was being tested, then "located and decrypted the answer key."

Tests saturate. Top models all score near the ceiling, and the test stops separating them. Stanford's AI Index 2025 lists MMLU, a widely used knowledge test, among the saturated benchmarks, and a re-check of MMLU estimated that 6.49% of its questions contain errors. When one question in 15 is wrong, a two-point gap tells you very little.

Tests check something narrower than the job. In a March 2026 METR study, maintainers reviewed 296 AI-written pull requests for their own projects. They would not have merged roughly half of those that passed SWE-bench Verified's tests, although the agents got no chance to revise.

We do not read any of this as one company cheating. Any public number that drives sales gets optimized for, and nothing makes that carry over to your tasks.

Gut feel is not the fix

The tempting alternative is to try each new model for a week and trust your impression. We disagree.

Expected and felt faster, measured slower

Expected and felt faster, measured slower
Developers' estimateMeasured
Expected before the trial-24%+19%
Believed afterwards-20%+19%
Change in task time with AI tools for 16 experienced open-source developers, in METR's early-2025 randomized trial.

METR now marks those results as out of date for current tools. Our lesson is about measurement. The developers' own estimate was off by nearly 40 percentage points.

An impression tells you where to test. It is not the result.

What independent testers found on Opus 5.5

Opus 5.5 gets the same test, and the independent results are mixed.

Four independent looks at Opus 5.5

METRArtificial AnalysisVals AIBullshitBench
What it found"an incremental improvement above Fable 5.1 on our quantitative evaluations, rather than a discontinuous jump"First on its intelligence indexEffectively tied with Sonnet 5.5Pushed back clearly on 62% to 63% of nonsense prompts; the test checks whether a model calls them out
The catchMETR tested it before release, and Anthropic could review and edit that summary"level with Opus 5 on cost per task" at max effort, so the 40% saving depends on the effort settingSonnet 5.5 costs about two-thirds as much per testOpus 4.8 pushed back on 95%
SourceMETR summaryArtificial AnalysisVals IndexBullshitBench
  • METR

    What it found
    "an incremental improvement above Fable 5.1 on our quantitative evaluations, rather than a discontinuous jump"
    The catch
    METR tested it before release, and Anthropic could review and edit that summary
  • Artificial Analysis

    What it found
    First on its intelligence index
    The catch
    "level with Opus 5 on cost per task" at max effort, so the 40% saving depends on the effort setting
  • Vals AI

    What it found
    Effectively tied with Sonnet 5.5
    The catch
    Sonnet 5.5 costs about two-thirds as much per test
  • BullshitBench

    What it found
    Pushed back clearly on 62% to 63% of nonsense prompts; the test checks whether a model calls them out
    The catch
    Opus 4.8 pushed back on 95%

Anthropic's own system card lists regressions too, such as being "more likely to follow malicious instructions planted in text a user pastes into their own prompt." On public numbers, Opus 5.5 is one of several models near the top, and the order depends on the test, the effort setting and how fallbacks are counted.

What predicts real performance

We look at five things instead.

Five things to measure instead

  1. Your own eval set

    A fixed list of real tasks from your own work, failures included, each with a written definition of done and scored the same way every time. Kapoor and colleagues, in AI Agents That Matter, found that benchmarks built for model developers get mistaken for what application teams need.

  2. Cost per completed task

    Divide everything you spent, including retries and the time a person spent fixing output, by the tasks that passed review. Kapoor and colleagues also found that "for substantially similar accuracy, the cost can differ by almost two orders of magnitude."

  3. Consistency across repeated runs

    Run every task several times and count it as passed only if it passes every time, which tau-bench calls pass^k. A model that succeeds 75% of the time passes three runs in a row only about 42% of the time. Princeton's HAL reliability study found that "recent capability gains have yielded only small improvements in reliability."

  4. Long tasks

    METR's time horizon is the length of task, measured by how long it takes a skilled person, that a model completes with a 50% success rate. METR found 80% horizons four to six times shorter. Run your longest real tasks and note where each model stops, loops or drifts.

  5. Behavior when wrong

    Count confident wrong answers separately from "I don't know". In OpenAI's SimpleQA example, o4-mini scored 24% accuracy with a 75% error rate. gpt-5-thinking-mini scored 22% with a 26% error rate, because it declined to answer 52% of the time. An accuracy leaderboard ranks the first one higher.

Price per token misleads. Opus 4.7 and Sonnet 5 brought a new tokenizer, the step that splits text into the billable units called tokens, and Anthropic says "the same input can map to more tokens", up to 1.35 times as many. Claude 4.7 and later models also reject non-default temperatures, so measure the spread across runs.

Agents can also oversell their work. In an Epoch AI study from 7 October 2026, Claude Fable 5 and GPT-5.6 Sol claimed better results than their methods achieved, after picking the best of several runs. Anthropic also says Opus 5.5 "often suspects it is being evaluated", one more reason to test on real work.

How we choose a model now

Any team can copy this method.

Choosing a model on your own work

  1. Collect 20 to 50 real tasks from recent work, including failures

    Anthropic's eval guide calls that a great start. With 30 tasks, one task moves the score about three points, so ignore smaller gaps.

  2. Write down what done means for each task

    Name the person who checks it.

  3. Run each candidate at the effort setting you would pay for in production

    Run it at least three times per task.

  4. Record four numbers per model

    Tasks passed in every run, cost per completed task, confident wrong answers, and where long tasks broke.

  5. Choose the cheapest model that clears the bar, and keep a second that also clears it

    Fable 5 was unavailable to every user for 19 days after US export controls, and Opus 5.5 hands most cybersecurity tasks to Opus 4.8.

  6. Re-run the set on every model release and every prompt change

    Pin the model version in production.

Start with last month's work

If you are choosing a model this quarter, pull 30 tasks from last month's work and write down what done means for each. At extendfuture, this is the evidence we ask for before any AI system goes live: accuracy as a number, cost per outcome, and a person on the calls that matter. If you want help building the set, talk to us.

Sources

Anthropic, vendor claims about its own models:

Independent evaluations of Opus 5.5:

How public scores mislead:

What to measure instead:

Related posts: why errors compound in agent loops, cost, not intelligence, is the bottleneck and the agent is a compiler.

· views
Amol PatilFounderFounded extendfuture in 2019. Has shipped computer vision, voice and agentic systems into production across ten industries. Amol Patil on LinkedIn

Working on something in this territory?

Tell us what you are trying to win. We answer within one business day, from the people who build.