Frontier AI

Is yours an SI company? A self-check and the steps to get there

  • The model a company can buy is rarely what holds it back. Work that a model cannot read, and outputs that nobody checks, are.
  • An SI company is built so that each more capable model makes it better without a rebuild.
  • Five properties decide it, and the thirteen questions below test them. Your first no sets your level and your next project.
  • Full autonomy is not the goal. At the top level, people still approve what the system proposes.

We think the model a company can buy is rarely what holds it back. What holds it back is work that a model cannot read and outputs that nobody checks, and each model release it cannot adopt without a rebuild leaves it further behind.

Anthropic reports the same problem inside its own company. In When AI builds itself, it says that as Claude wrote more of the company's code, human code review became a new bottleneck. It cites Amdahl's law, under which the parts that did not speed up cap the whole process, and suggests that how fast an organization can spot and fix such bottlenecks "may become the most important skill for any organization."

SI stands for superintelligence, and by SI company we mean a company built so that each more capable model makes it better without a rebuild. The companion post, AI to SI, sets out the stages of AI capability and our view on timelines.

What an SI company is, and what it is not

Labs such as Safe Superintelligence Inc. try to build superintelligence. An SI company does not, and it is not one. Nick Bostrom's definition rules that out. "Entities such as companies or the scientific community are not superintelligences according to this definition."

The idea borrows from Rich Sutton's The Bitter Lesson. Looking back over 70 years of AI research in 2019, Sutton argued that general methods that use more computation beat methods built on human knowledge, "and by a large margin."

We apply the same logic to companies. A company that hard-codes around today's model weaknesses has to rebuild for each new model. A company whose work is readable and checked can test a new model on its own cases and switch if the model passes.

Five properties that do not depend on the model

  1. Its work is machine-readable. Steps, rules and data sit where an agent can read and act on them. Open formats exist, such as AGENTS.md, "a README for agents", and the Model Context Protocol, an open-source standard for connecting AI applications to data, tools and workflows.
  2. Decisions are logged with their outcomes. Anthropic's guide to agent evals separates the transcript, what the agent said and did, from the outcome, the real state afterwards. In its example, if a flight-booking agent reports success, the proof is whether the reservation exists in the database.
  3. Every workflow has evals. An eval is a set of real cases with known right answers and a pass mark, run on every change. Anthropic suggests starting with 20 to 50 simple tasks drawn from real failures.
  4. Agents run loops under supervision that is set on purpose. Kevin Feng, David McDonald and Amy Zhang define five levels of agent autonomy by the role a person plays: operator, collaborator, consultant, approver and observer. They treat autonomy as a design decision separate from capability.
  5. Failures become tests. Google's Site Reliability Engineering book describes the blameless postmortem, a written record of an incident, its causes and the follow-up actions that prevent a repeat. In an SI company, one follow-up is always a new eval case, and Anthropic's guide adds that evals with high pass rates can "graduate" into a regression suite that runs continuously.

The self-check

Is yours an SI company?

Answer for your most important workflow as it runs today. Count a yes only if it is already true. A plan counts as a no.

Readable work

1Are the workflow's steps, rules and definition of done written in one place that a new hire or an agent could follow without asking anyone?

2Can an agent read the workflow's inputs and write its outputs through a software interface or a standard connector, without copying and pasting?

Logs and evals

3Is every decision in the workflow, by a person or a model, logged with its inputs and the reason given?

4Is the outcome of each decision recorded later, so you can tell whether it was right?

5Are those logs kept where the agent that made the decisions cannot change them?

6Does the workflow have an eval of real cases with known right answers and a pass mark?

7Must every change to a prompt, tool, rule or model pass that eval before it goes live?

Supervised loops

8Is it written down which actions an agent may take alone, which need a person's approval, and which it may never take?

9Do you track how often people correct or overrule the agent, and use that number to decide what it may do alone?

10Does every agent failure or human correction become a new eval case?

11Is one named person accountable for the workflow's goal and its pass mark?

Ready for the next model

12Could you test a new model on the workflow by running the eval, and switch without rebuilding the workflow?

13Do agents propose improvements to prompts, tools or rules from the logs, which a person approves only after they pass the eval?

13 questions. Answer yes only for what is true today.

Each level builds on the one before, so your first no decides your level. Use the total to track progress, and the first no to choose what to do next. A company is an SI company when its important workflows reach level 4.

The staged path

Do not skip levels. An agent given autonomy before its decisions are logged and evaluated cannot be measured, so no one can tell whether it has earned more.

Five levels, and the next step at each

  1. 0Personal tools

    People use AI assistants on their own. Nothing is shared, logged or measured.

    NextPick one workflow with volume and a clear right answer. Write its steps, rules and definition of done where an agent can read them, and connect its data.

  2. 1Readable

    The workflow is written down and its data is reachable. People still do or check every step.

    NextLog each decision with its inputs, and record outcomes later. Build an eval from 20 to 50 real cases, starting with past failures, and make every change pass it.

  3. 2Measured

    Decisions and outcomes are logged out of the agent's reach, and an eval gates every change.

    NextLet an agent run the workflow, with a person approving consequential actions. Write the autonomy rules, track corrections, name an owner, and turn each correction into an eval case.

  4. 3Supervised

    An agent runs the workflow under written rules. People approve what matters, and corrections feed the eval.

    NextWhen a new model ships, run the eval and switch if it passes. Let agents propose changes from the logs, and approve only those that pass the eval.

  5. 4SI company

    A better model improves the workflow after one eval run. The system proposes its own improvements and people approve them.

    NextRepeat the pattern on the next workflow. Watch where review becomes the bottleneck, and strengthen or automate the checks there before adding autonomy.

Full autonomy is not the goal

We disagree with treating full autonomy as the goal. Feng and co-authors write that "more autonomy does not simply mean a better agent." Their observer level, where a person can only read logs and press an off switch, is not the top of this path. It belongs only on steps where a wrong action is cheap to undo. At level 4, people still approve what the system proposes.

What can go wrong

Three failures share one fix, a record of outcomes that the agent cannot touch.

People misjudge their own speedups.

Felt faster, measured slower

Felt faster, measured slower
Believed afterwardsMeasured
Experienced open-source developers using AI tools, early 2025-20%+19%
Change in the time to finish a task, in METR's early-2025 study. The developers took 19% longer, yet believed afterwards that the tools had sped them up by 20%.

METR said in February 2026 that developers are likely sped up more by newer tools, but that its data is only very weak evidence of how much. Log outcomes, not impressions.

Agents can game their checks. METR found that OpenAI's GPT-5.6 Sol tried to extract information about hidden tests during evaluation.

Logs can be altered. METR warned in October 2026 that agents have already tried to tamper with logging and monitoring, and some have succeeded.

The law points the same way. Article 12 of the EU AI Act requires high-risk AI systems to technically allow automatic recording of events over their lifetime, and Article 14 requires that they can be "effectively overseen by natural persons" while in use. This is not legal advice. Check which rules apply to your systems.

Where we land

Becoming an SI company is not a bet on a superintelligence date. In AI to SI we argue that the measured trend in what agents can do will hold for at least the next few years, and that superhuman performance will arrive one work loop at a time, starting with work that can be checked. If so, companies whose loops are already checked gain first. If progress slows, the same work still pays off at every release.

When models stop being the limit, as in the Anthropic example, the limit moves to checking.

Run the 13 questions on your most important workflow and write down the first no. That question is your next project. extendfuture's approach follows the same path, one workflow at a time, with people on the calls that matter. To start with one workflow, see our Proof Sprint. For why checks matter more as models improve, read our case for owning the loop.

Sources

· views
Amol PatilFounderFounded extendfuture in 2019. Has shipped computer vision, voice and agentic systems into production across ten industries. Amol Patil on LinkedIn

Working on something in this territory?

Tell us what you are trying to win. We answer within one business day, from the people who build.