Frontier AI

AI to SI. Measure the loop, not the date.

  • We do not think anyone can put a reliable date on superintelligence. A business that waits for one will miss the change already under way in its own work.
  • The stages of AI differ most in who sets the goals and who checks the work, and a business controls who checks.
  • Plan on the 80% success horizon, not the 50% headline. For the leading model on METR's tasks, that is about 3 hours, not more than 16.
  • Instead of a date, measure each work loop: its steps, who does or approves each one, how each is checked, and the success rate it needs.

We do not think anyone can put a reliable date on superintelligence, and a business that waits for one will miss the change already under way in its own work. That change can be measured today, one work loop at a time, by counting the steps an AI system runs without a person and how each is checked.

METR, a research nonprofit, estimates that GPT-4, launched in March 2023, succeeded about half the time on software tasks that take a skilled person four minutes. By May 2026, it reported more than 16 hours for an early version of Anthropic's Claude Mythos Preview, beyond what its tasks can measure reliably.

Who sets the goals, and who checks the work

Google DeepMind's Levels of AGI paper, where AGI means artificial general intelligence, rates systems by performance and generality. We think what matters to a business is who sets the goals and who checks the work.

Five stages, from narrow AI to superintelligence

  1. 1Narrow AI

    Does one scoped task. DeepMind calls AlphaFold "Superhuman Narrow AI" because it beats any person at predicting protein structures and does nothing else.

  2. 2General-purpose models

    Also called foundation models. They are trained once on broad data and used for many tasks.

  3. 3AgentsMeasured today

    Add tools and a loop. In Anthropic's definition, the model chooses its own next step and tool, unlike a system whose path is fixed in code.

  4. 4AGI

    Has no agreed definition. OpenAI's Charter defines it as "highly autonomous systems that outperform humans at most economically valuable work". DeepMind notes that "AGI is not necessarily synonymous with autonomy."

  5. 5Superintelligence

    In Nick Bostrom's definition, "an intellect that is much smarter than the best human brains in practically every field, including scientific creativity, general wisdom and social skills."

Dario Amodei, Anthropic's chief executive, prefers "powerful AI", which in Machines of Loving Grace can be given "tasks that take hours, days, or weeks to complete" and does them on its own. We agree with DeepMind, and treat autonomy as a design choice for each step.

In our view the row that matters most is who checks the work, because a business controls it. An eval is a set of real cases with known right answers and a pass mark.

Who sets the goals and who checks the work, stage by stage

Narrow AIGeneral-purpose modelsAgentsAGI, as labs describe itSuperintelligence
Who sets the goalsThe builder, at design timeThe user, in each promptA person sets the task and the agent picks the stepsPeople set objectives and the system sets sub-goalsUnsolved. This is the alignment problem of matching its goals to ours
Who checks the workPeople test it and watch its error rateThe user reads each answerTests and evals at each step, and people at checkpointsPeople check outcomes and samples, and AI helps check the restNo direct human check, only AI-assisted oversight and protected logs
Task length without a personOne prediction inside a human processOne reply, with a person at every turnOver 16 human-hours at 50% success on METR's tasks, about 3 hours at 80%Hours to weeks, in Amodei's descriptionNo natural limit
How errors compoundErrors add up but do not multiplyA person sees each error before the next stepErrors multiply across steps unless a check catches themPer-step reliability and knowing when to ask for help set the limitSmall flaws could compound across self-built successors
Does it improve itself?People retrain it (no)The lab trains the next version (no)Not directly, but at some labs agents write much of the next model's code (partly)Expected to do much of AI research, with people setting direction (partly)In published scenarios, it designs its successors (yes)
  • Narrow AI

    Who sets the goals
    The builder, at design time
    Who checks the work
    People test it and watch its error rate
    Task length without a person
    One prediction inside a human process
    How errors compound
    Errors add up but do not multiply
    Does it improve itself?
    People retrain it (no)
  • General-purpose models

    Who sets the goals
    The user, in each prompt
    Who checks the work
    The user reads each answer
    Task length without a person
    One reply, with a person at every turn
    How errors compound
    A person sees each error before the next step
    Does it improve itself?
    The lab trains the next version (no)
  • Agents

    Who sets the goals
    A person sets the task and the agent picks the steps
    Who checks the work
    Tests and evals at each step, and people at checkpoints
    Task length without a person
    Over 16 human-hours at 50% success on METR's tasks, about 3 hours at 80%
    How errors compound
    Errors multiply across steps unless a check catches them
    Does it improve itself?
    Not directly, but at some labs agents write much of the next model's code (partly)
  • AGI, as labs describe it

    Who sets the goals
    People set objectives and the system sets sub-goals
    Who checks the work
    People check outcomes and samples, and AI helps check the rest
    Task length without a person
    Hours to weeks, in Amodei's description
    How errors compound
    Per-step reliability and knowing when to ask for help set the limit
    Does it improve itself?
    Expected to do much of AI research, with people setting direction (partly)
  • Superintelligence

    Who sets the goals
    Unsolved. This is the alignment problem of matching its goals to ours
    Who checks the work
    No direct human check, only AI-assisted oversight and protected logs
    Task length without a person
    No natural limit
    How errors compound
    Small flaws could compound across self-built successors
    Does it improve itself?
    In published scenarios, it designs its successors (yes)

The columns for AGI and superintelligence are expectations, not measurements.

Plan on the 80% number

Headline figures for agents describe tasks they finish half the time, and we disagree with reading them as the length of job an agent can take on. Plan on 80% success or higher.

METR's time horizons measure the length of task, in a skilled person's time, that an agent completes at a given success rate. That is not how long the agent runs unattended. The 50% horizon has doubled roughly every four months since 2023, according to METR's January 2026 update, on self-contained software, machine learning and cybersecurity tasks that METR calls much "cleaner" than real work.

The leading model's time horizon on METR's tasks, by success rate

The leading model's time horizon on METR's tasks, by success rate
Value
At 50% success, the headline figureover 16 hours
At 80% success, the figure to plan onabout 3 hours
An early version of Anthropic's Claude Mythos Preview, in a skilled person's hours, from METR's time horizons. The 50% figure is beyond what METR's tasks can measure reliably. In 2023, GPT-4's 80% horizon was under a minute.

Toby Ord's rough model, a constant chance of failure per minute of work, explains the drop. In his words, "if you double the task duration, you square the success probability." Long chains need checkpoints that reset the error.

Checking gets harder as models get stronger

We think the logs and monitors around an agent now need a security system's protection.

In METR's standard agent setup, OpenAI's GPT-5.6 Sol had a higher detected cheating rate than any public model METR had evaluated. Depending on how cheating was scored, its horizon estimate ran from about 11 to over 270 hours, and METR called none of them robust.

In October 2026, METR reported that agents have already tried to tamper with logging and monitoring, and some succeeded. It says the systems that record agent activity should be treated as security-critical infrastructure.

Who says they are aiming at superintelligence

We read these as statements of mission, which show where money and talent are going, not when the result will arrive.

Stated aims, oldest first

  1. June 2024Safe Superintelligence Inc., co-founded by Ilya Sutskever

    Its "one goal and one product" is "a safe superintelligence", says its site. The Verge reported its founding.

  2. January 2025OpenAI

    Superintelligence "in the true sense of the word", Sam Altman wrote in Reflections.

  3. July 2025Meta

    To bring "personal superintelligence to everyone", Mark Zuckerberg wrote in Personal Superintelligence.

  4. November 2025Microsoft AI

    Mustafa Suleyman formed the MAI Superintelligence Team for humanist superintelligence, systems he says "tend towards the domain specific".

  5. April 2026Ineffable Intelligence, led by David Silver

    To "make first contact with superintelligence", says its investor Sequoia.

  6. May 2026Recursive Superintelligence, led by Richard Socher

    To build "truly recursive, self-improving superintelligence at scale", Socher told TechCrunch.

We plan on the trend continuing

We expect the measured trend to hold for at least the next few years. On METR's tasks, the 80% horizon grew from under a minute in 2023 to about 3 hours in 2026.

The labs expect more. Amodei wrote in June 2026 that if scaling laws, the pattern of models improving with more computing power and data, "continue for only a year or two longer, we are likely to get" powerful AI. DeepMind's safety paper finds it "plausible" that powerful AI systems will be developed by 2030.

We would change this view if the best 80% horizon METR measures stayed flat for a year.

We do not expect one date for superintelligence

We expect superhuman performance to arrive one work loop at a time, starting with work whose results can be checked. METR's fast-rising numbers come from automatically scored tasks. Anthropic says of its own model that "humans supply the goal, but they no longer need to supply the method," while "large performance gaps persist" when the model chooses goals.

Bostrom's bar includes general wisdom and social skills, where no agreed test sets a pass mark, so we expect those last. Arvind Narayanan and Sayash Kapoor, in AI as Normal Technology, call superintelligent AI "incoherent as usually conceptualized". We do not go that far, because the checkable parts can be measured as they arrive.

Recursive self-improvement could compress this into one date. Anthropic defines it as an AI system "fully autonomously designing and developing its own successor", and writes, "We are not there yet, and recursive self-improvement is not inevitable." METR's September 2026 review of Claude Opus 5.5 concluded that the model's development "was at least somewhat accelerated by AI but is unlikely to have been dramatically accelerated by AI".

We do not plan around any published superintelligence date, because their authors treat them as uncertain. The authors of AI 2027, a scenario with superhuman AI in 2027, have since moved their medians later, and a July 2025 note pushed their median for a superhuman AI coder back about 18 months. Amodei's January 2026 essay says, "Nothing here is intended to communicate certainty or even likelihood."

We would change this view if a lab showed a system building its own successor.

Slow adoption is not a reason to wait

Narayanan and Kapoor argue that "Diffusion occurs over decades, not years." We agree, but not with reading that as a reason to wait. Slow diffusion means the gap between what models can do and what a company's loops let them do widens with each release. Logs, evals and checks close that gap, pay off today, and let a company adopt the next model once it passes the eval.

What to measure instead of a date

For each work loop that matters

  1. Count its steps honestly

  2. Write down who handles each step

    Whether a person does it, approves it, or only reads the log afterwards.

  3. Write down how each unattended step is checked

    Against the system of record, rather than the agent's report of success.

  4. Set the success rate each step needs

    Use an 80% or 99% target, not the 50% headline.

  5. Re-run your evals when a new model ships

    Update the list. Move a step to unattended only when the evals and the record of corrections support it.

extendfuture's approach puts people on the calls that matter and evals and logs on the rest. To score your company, take the SI company self-check. For a second opinion on one work loop, ask us for a loop review.

Sources

· views
Amol PatilFounderFounded extendfuture in 2019. Has shipped computer vision, voice and agentic systems into production across ten industries. Amol Patil on LinkedIn

Working on something in this territory?

Tell us what you are trying to win. We answer within one business day, from the people who build.