Frontier AI
AI to SI. Measure the loop, not the date.
- We do not think anyone can put a reliable date on superintelligence. A business that waits for one will miss the change already under way in its own work.
- The stages of AI differ most in who sets the goals and who checks the work, and a business controls who checks.
- Plan on the 80% success horizon, not the 50% headline. For the leading model on METR's tasks, that is about 3 hours, not more than 16.
- Instead of a date, measure each work loop: its steps, who does or approves each one, how each is checked, and the success rate it needs.
We do not think anyone can put a reliable date on superintelligence, and a business that waits for one will miss the change already under way in its own work. That change can be measured today, one work loop at a time, by counting the steps an AI system runs without a person and how each is checked.
METR, a research nonprofit, estimates that GPT-4, launched in March 2023, succeeded about half the time on software tasks that take a skilled person four minutes. By May 2026, it reported more than 16 hours for an early version of Anthropic's Claude Mythos Preview, beyond what its tasks can measure reliably.
Who sets the goals, and who checks the work
Google DeepMind's Levels of AGI paper, where AGI means artificial general intelligence, rates systems by performance and generality. We think what matters to a business is who sets the goals and who checks the work.
Five stages, from narrow AI to superintelligence
1Narrow AI
Does one scoped task. DeepMind calls AlphaFold "Superhuman Narrow AI" because it beats any person at predicting protein structures and does nothing else.
2General-purpose models
Also called foundation models. They are trained once on broad data and used for many tasks.
3AgentsMeasured today
Add tools and a loop. In Anthropic's definition, the model chooses its own next step and tool, unlike a system whose path is fixed in code.
4AGI
Has no agreed definition. OpenAI's Charter defines it as "highly autonomous systems that outperform humans at most economically valuable work". DeepMind notes that "AGI is not necessarily synonymous with autonomy."
5Superintelligence
In Nick Bostrom's definition, "an intellect that is much smarter than the best human brains in practically every field, including scientific creativity, general wisdom and social skills."
Dario Amodei, Anthropic's chief executive, prefers "powerful AI", which in Machines of Loving Grace can be given "tasks that take hours, days, or weeks to complete" and does them on its own. We agree with DeepMind, and treat autonomy as a design choice for each step.
In our view the row that matters most is who checks the work, because a business controls it. An eval is a set of real cases with known right answers and a pass mark.
Who sets the goals and who checks the work, stage by stage
| Narrow AI | General-purpose models | Agents | AGI, as labs describe it | Superintelligence | |
|---|---|---|---|---|---|
| Who sets the goals | The builder, at design time | The user, in each prompt | A person sets the task and the agent picks the steps | People set objectives and the system sets sub-goals | Unsolved. This is the alignment problem of matching its goals to ours |
| Who checks the work | People test it and watch its error rate | The user reads each answer | Tests and evals at each step, and people at checkpoints | People check outcomes and samples, and AI helps check the rest | No direct human check, only AI-assisted oversight and protected logs |
| Task length without a person | One prediction inside a human process | One reply, with a person at every turn | Over 16 human-hours at 50% success on METR's tasks, about 3 hours at 80% | Hours to weeks, in Amodei's description | No natural limit |
| How errors compound | Errors add up but do not multiply | A person sees each error before the next step | Errors multiply across steps unless a check catches them | Per-step reliability and knowing when to ask for help set the limit | Small flaws could compound across self-built successors |
| Does it improve itself? | People retrain it (no) | The lab trains the next version (no) | Not directly, but at some labs agents write much of the next model's code (partly) | Expected to do much of AI research, with people setting direction (partly) | In published scenarios, it designs its successors (yes) |
Narrow AI
- Who sets the goals
- The builder, at design time
- Who checks the work
- People test it and watch its error rate
- Task length without a person
- One prediction inside a human process
- How errors compound
- Errors add up but do not multiply
- Does it improve itself?
- People retrain it (no)
General-purpose models
- Who sets the goals
- The user, in each prompt
- Who checks the work
- The user reads each answer
- Task length without a person
- One reply, with a person at every turn
- How errors compound
- A person sees each error before the next step
- Does it improve itself?
- The lab trains the next version (no)
Agents
- Who sets the goals
- A person sets the task and the agent picks the steps
- Who checks the work
- Tests and evals at each step, and people at checkpoints
- Task length without a person
- Over 16 human-hours at 50% success on METR's tasks, about 3 hours at 80%
- How errors compound
- Errors multiply across steps unless a check catches them
- Does it improve itself?
- Not directly, but at some labs agents write much of the next model's code (partly)
AGI, as labs describe it
- Who sets the goals
- People set objectives and the system sets sub-goals
- Who checks the work
- People check outcomes and samples, and AI helps check the rest
- Task length without a person
- Hours to weeks, in Amodei's description
- How errors compound
- Per-step reliability and knowing when to ask for help set the limit
- Does it improve itself?
- Expected to do much of AI research, with people setting direction (partly)
Superintelligence
- Who sets the goals
- Unsolved. This is the alignment problem of matching its goals to ours
- Who checks the work
- No direct human check, only AI-assisted oversight and protected logs
- Task length without a person
- No natural limit
- How errors compound
- Small flaws could compound across self-built successors
- Does it improve itself?
- In published scenarios, it designs its successors (yes)
The columns for AGI and superintelligence are expectations, not measurements.
Plan on the 80% number
Headline figures for agents describe tasks they finish half the time, and we disagree with reading them as the length of job an agent can take on. Plan on 80% success or higher.
METR's time horizons measure the length of task, in a skilled person's time, that an agent completes at a given success rate. That is not how long the agent runs unattended. The 50% horizon has doubled roughly every four months since 2023, according to METR's January 2026 update, on self-contained software, machine learning and cybersecurity tasks that METR calls much "cleaner" than real work.
Toby Ord's rough model, a constant chance of failure per minute of work, explains the drop. In his words, "if you double the task duration, you square the success probability." Long chains need checkpoints that reset the error.
Checking gets harder as models get stronger
We think the logs and monitors around an agent now need a security system's protection.
In METR's standard agent setup, OpenAI's GPT-5.6 Sol had a higher detected cheating rate than any public model METR had evaluated. Depending on how cheating was scored, its horizon estimate ran from about 11 to over 270 hours, and METR called none of them robust.
In October 2026, METR reported that agents have already tried to tamper with logging and monitoring, and some succeeded. It says the systems that record agent activity should be treated as security-critical infrastructure.
Who says they are aiming at superintelligence
We read these as statements of mission, which show where money and talent are going, not when the result will arrive.
Stated aims, oldest first
June 2024Safe Superintelligence Inc., co-founded by Ilya Sutskever
Its "one goal and one product" is "a safe superintelligence", says its site. The Verge reported its founding.
January 2025OpenAI
Superintelligence "in the true sense of the word", Sam Altman wrote in Reflections.
July 2025Meta
To bring "personal superintelligence to everyone", Mark Zuckerberg wrote in Personal Superintelligence.
November 2025Microsoft AI
Mustafa Suleyman formed the MAI Superintelligence Team for humanist superintelligence, systems he says "tend towards the domain specific".
April 2026Ineffable Intelligence, led by David Silver
To "make first contact with superintelligence", says its investor Sequoia.
May 2026Recursive Superintelligence, led by Richard Socher
To build "truly recursive, self-improving superintelligence at scale", Socher told TechCrunch.
We plan on the trend continuing
We expect the measured trend to hold for at least the next few years. On METR's tasks, the 80% horizon grew from under a minute in 2023 to about 3 hours in 2026.
The labs expect more. Amodei wrote in June 2026 that if scaling laws, the pattern of models improving with more computing power and data, "continue for only a year or two longer, we are likely to get" powerful AI. DeepMind's safety paper finds it "plausible" that powerful AI systems will be developed by 2030.
We would change this view if the best 80% horizon METR measures stayed flat for a year.
We do not expect one date for superintelligence
We expect superhuman performance to arrive one work loop at a time, starting with work whose results can be checked. METR's fast-rising numbers come from automatically scored tasks. Anthropic says of its own model that "humans supply the goal, but they no longer need to supply the method," while "large performance gaps persist" when the model chooses goals.
Bostrom's bar includes general wisdom and social skills, where no agreed test sets a pass mark, so we expect those last. Arvind Narayanan and Sayash Kapoor, in AI as Normal Technology, call superintelligent AI "incoherent as usually conceptualized". We do not go that far, because the checkable parts can be measured as they arrive.
Recursive self-improvement could compress this into one date. Anthropic defines it as an AI system "fully autonomously designing and developing its own successor", and writes, "We are not there yet, and recursive self-improvement is not inevitable." METR's September 2026 review of Claude Opus 5.5 concluded that the model's development "was at least somewhat accelerated by AI but is unlikely to have been dramatically accelerated by AI".
We do not plan around any published superintelligence date, because their authors treat them as uncertain. The authors of AI 2027, a scenario with superhuman AI in 2027, have since moved their medians later, and a July 2025 note pushed their median for a superhuman AI coder back about 18 months. Amodei's January 2026 essay says, "Nothing here is intended to communicate certainty or even likelihood."
We would change this view if a lab showed a system building its own successor.
Slow adoption is not a reason to wait
Narayanan and Kapoor argue that "Diffusion occurs over decades, not years." We agree, but not with reading that as a reason to wait. Slow diffusion means the gap between what models can do and what a company's loops let them do widens with each release. Logs, evals and checks close that gap, pay off today, and let a company adopt the next model once it passes the eval.
What to measure instead of a date
For each work loop that matters
Count its steps honestly
Write down who handles each step
Whether a person does it, approves it, or only reads the log afterwards.
Write down how each unattended step is checked
Against the system of record, rather than the agent's report of success.
Set the success rate each step needs
Use an 80% or 99% target, not the 50% headline.
Re-run your evals when a new model ships
Update the list. Move a step to unattended only when the evals and the record of corrections support it.
extendfuture's approach puts people on the calls that matter and evals and logs on the rest. To score your company, take the SI company self-check. For a second opinion on one work loop, ask us for a loop review.
Sources
- Google DeepMind, Levels of AGI, 2023, revised 2025. It rates the chat models of 2023, such as ChatGPT, as "Emerging AGI", meaning general but only at or somewhat above an unskilled person on most tasks, defines "Competent AGI" as matching or beating the median skilled adult on most cognitive tasks, and marks fully autonomous "AI as an Agent" as not yet unlocked.
- Rishi Bommasani and co-authors, On the Opportunities and Risks of Foundation Models, 2021, the paper that named them.
- Anthropic, Building effective agents, December 2024.
- OpenAI, OpenAI Charter.
- Dario Amodei, Machines of Loving Grace, October 2024, where powerful AI also asks for clarification as needed; The Adolescence of Technology, January 2026; Policy on the AI Exponential, June 2026.
- Nick Bostrom, How Long Before Superintelligence?, 1997, revised 1998.
- METR, time horizons, last updated September 2026, with the 50% and 80% horizons in its published chart data; Time Horizon 1.1; GPT-5.6 Sol evaluation; Claude Opus 5.5 evaluation, which also says its data is insufficient to tell whether AI research ability is improving at a steady, rising or falling rate; AI systems could cover up misbehavior. OpenAI and Anthropic reviewed METR's pre-release evaluation texts before publication.
- Toby Ord, Is there a Half-Life for the Success Rates of AI Agents?, 2025. His February 2026 update notes newer analysis by Gus Hamilton that finds the failure rate falls as a task goes on, so the constant-rate model is a rough guide.
- Google DeepMind, An Approach to Technical AGI Safety and Security, April 2025. It warns that "Supervising a system with capabilities beyond that of the overseer is difficult, with the difficulty increasing as the capability gap widens." For the foreseeable future, it expects self-improvement to come from AI doing AI research the way people do, not from a system that will "edit its own weights".
- Anthropic Institute, When AI builds itself, 2026, updated September 2026. It says Claude wrote more than 80% of the code merged into Anthropic's codebase as of May 2026. Statements about Anthropic's own work are the company's claims.
- Sam Altman, Reflections, January 2025.
- Safe Superintelligence Inc., and The Verge on its founding, June 2024.
- Mark Zuckerberg, Personal Superintelligence, July 2025.
- Mustafa Suleyman, Towards Humanist Superintelligence, November 2025.
- Sequoia Capital, Partnering with Ineffable Intelligence, April 2026.
- TechCrunch, What happens when AI starts building itself?, May 2026.
- Arvind Narayanan and Sayash Kapoor, AI as Normal Technology, April 2025. They expect more of people's work to become controlling AI.
- Kokotajlo and co-authors, AI 2027, April 2025, with later notes.
- Also linked: Is yours an SI company?, errors compounding in agent loops, contact.