Creuto is now an OpenAI Select Partner Read More

AI & Machine Learning

What is superintelligence? The definition, minus the hype

What is superintelligence? The definition, what benchmarks actually measure as of October 2026, what nobody can measure, and why we give no date.

What is superintelligence? The definition, minus the hype

What is superintelligence? It is a system that outperforms the best human minds in practically every field, not a system that beats most people at many tasks. The gap between those two sentences is the whole subject, and the most rigorous scientific assessment of AI we have does not use the word at all.

That assessment is the International AI Safety Report 2026, chaired by Yoshua Bengio, submitted on 24 February 2026, and guided by over 100 experts from more than 30 countries and international organisations. It runs to 220 pages across three risk domains: malicious use, malfunctions and loss of control, and systemic risks. Search it for "superintelligence" and you find nothing in the analysis. The words "superintelligent" and "superhuman" appear only in the bibliography, inside the titles of papers it cites.

That is not an oversight. It is what a body of evidence looks like when it restricts itself to things that can be measured. This post separates the definition from the marketing: what the word means, what today's systems are actually scored on, what no one can currently score at all, and who is forecasting what. Everything below is labelled either measured or forecast. We do not give a date.

What is superintelligence? The definition that has stood since 1998

The working definition in the literature comes from Nick Bostrom's paper How Long Before Superintelligence?, first published in the International Journal of Futures Studies in 1998 and still on his own site. It reads:

By a "superintelligence" we mean an intellect that is much smarter than the best human brains in practically every field, including scientific creativity, general wisdom and social skills.

Three load-bearing words there. Best, not average. Practically every field, not many fields. And general wisdom and social skills, which are named alongside scientific creativity rather than treated as a rounding error.

Bostrom's definition also rules things out on purpose. He writes that entities such as companies or the scientific community are not superintelligences under it, even though they outperform any individual. The claim is about a single intellect, not an aggregate.

Superintelligence vs artificial intelligence, and where AGI sits between them

Artificial intelligence is the broad category: systems performing tasks that usually need human intelligence. Artificial general intelligence is the middle rung, and the safety report defines it once, in its glossary, with one word doing the heavy lifting: "a hypothetical AI model or system that equals or surpasses human performance on all or almost all cognitive tasks". Superintelligence, or ASI, is the rung above that: not parity with humans across the board, but decisive advantage over the best of them across the board.

So the ladder is narrow to general to super, and the report's own vocabulary tells you where the evidence stops. It has a chapter on measured capabilities. It has a glossary entry, flagged hypothetical, for AGI. It has nothing for the rung above.

Why "better than most people at many tasks" is a different claim

Here is the strongest version of the opposite argument, because it deserves to be stated properly. Frontier systems now beat most humans at a very wide range of cognitive work. Measured: the safety report records leading models scoring over 90% on undergraduate-level examinations from chemistry to law and over 80% on graduate-level science questions in GPQA, and notes that in July 2025 models from Google DeepMind and OpenAI reached gold-medal-level scores at the International Mathematical Olympiad, solving five of six problems under competition-like conditions. If a system out-argues you on law and out-solves you on olympiad geometry, calling it merely narrow feels like denial.

The answer is in the same report, in the word it uses over and over: jagged. Its own summary of current capability is that "leading systems may excel at some difficult tasks while failing at other, simpler ones". It lists the failures concretely: systems get derailed by simple errors during multi-step projects, still generate false statements, produce inconsistent outputs given identical inputs, and cannot yet integrate with robotic components to perform basic physical tasks such as housework.

Bostrom's definition has no room for jagged. "Practically every field" is a claim about the floor of a capability profile, not the ceiling. A system that wins an olympiad medal and then miscounts the words in a paragraph has a very high ceiling and a floor that is nowhere near the best humans. Breadth of excellence is the thing being claimed, and breadth is exactly what the evidence says is uneven.

What is actually measured today, as of 1 October 2026

The cleanest public evidence on general reasoning is the ARC-AGI family, because its tasks are designed so that memorising the training set does not help. ARC Prize states plainly what it is scoring: "skill-acquisition efficiency on unknown tasks", and defines AGI as "a system that can match the learning efficiency of humans". Their own framing of the benchmark is a measurement of efficiency, not of raw score.

From the official ARC Prize leaderboard, whose underlying data file was generated on 30 September 2026:

BenchmarkBest system (30 Sep 2026)ScoreHuman baseline
ARC-AGI-1 (semi-private)Several, incl. GPT-6.1 Sol (High)98.5%Human Panel 98.0%
ARC-AGI-2 (semi-private)GPT-6 Astra (Max)95.0%Human Panel 100%

Measured: on ARC-AGI-1 the leading systems are now above the human panel, and the cheapest system at 98.5% did it at about six cents per task against $17 per task for the human panel. On ARC-AGI-2, which adds symbolic interpretation, compositional reasoning and context-dependent rule application, the best score is 95.0% against a human panel that solved 100%. ARC Prize has since published ARC-AGI-3, an interactive benchmark that scores how efficiently an agent explores novel environments and acquires goals over time rather than whether it returns the right final answer.

Two things follow, and they point in opposite directions. Systems have crossed a human panel on a test built specifically to resist memorisation, which is a real result. And the organisation that built that test immediately built a harder one and states the criterion it still applies: "As long as there is a gap between AI and human learning, we do not have AGI." Not superintelligence. AGI, the rung below.

The measurement that travels furthest: how long a task can be

The most useful single number for anyone planning AI engineering work is not an exam score. It is task duration. The safety report cites a study finding that AI systems complete well-specified software engineering tasks that take human experts 30 minutes around 80% of the time, and that this task length has been doubling roughly every seven months. Forecast, explicitly labelled as one in the report: if that trend continues, systems could complete tasks lasting several hours by 2027 and several days by 2030. The trend is measured; the extrapolation is not a finding.

What nobody can currently measure

The report names the limit on all of this in one phrase: the evaluation gap. Its executive summary puts it as "performance on pre-deployment tests does not reliably predict real-world utility or risk". Elsewhere it is blunter: many common capability evaluations are outdated, affected by data contamination when models are trained on the same questions used to test them, or focused on a narrow set of tasks.

It then names the gap in the literature itself: "There is no single, comprehensive, and continuously updated synthesis of AI capabilities, leading to a fragmented and often outdated understanding of the field." With no widely accepted taxonomy for capabilities, policymakers navigate a patchwork of benchmarks.

So the honest list of what cannot currently be measured is short and consequential. There is no accepted test for general wisdom or social skill, the two things Bostrom's definition names alongside scientific creativity. There is no test that establishes the floor of a capability profile, only tests that probe selected points on it. The report says performance depends heavily on the specific test examples and prompt used, making it hard to prove with high confidence that a system cannot do something. And there is no measurement of superintelligence at all, because no one has proposed a test that a system passes only if it exceeds the best humans in practically every field.

That last point is the one worth carrying away. The absence is not a gap waiting for a better benchmark next quarter. You cannot score a threshold whose definition ranges over every field, including ones nobody has enumerated.

Who is forecasting what, and why we will not give you a date

Dated, attributed forecasts exist and are worth reading as forecasts. Forecast: Mark Zuckerberg wrote on 30 July 2025 that "Developing superintelligence is now in sight", in a letter setting out Meta's goal of personal superintelligence for everyone. Forecast: Demis Hassabis told a Stanford Graduate School of Business audience in a conversation published on 18 June 2026 that "Ten years from now, I think we'll realize that we were standing in the foothills of the singularity now."

In the same conversation Hassabis said something more useful about the genre than any of the dates in it. Speaking of his peers at other labs: "I think they're being way too certain, I would say, with some of their pronouncements. Where I think actually there's just huge uncertainty."

The expert assessment agrees. Rather than a date, the safety report presents four scenarios for 2030 developed by the OECD, concluding that "by 2030, AI progress could plausibly range from stagnation to rapid improvement to levels that exceed human cognitive performance". Its own conclusion states that contributors differ in their views on how quickly capabilities will improve. A range that wide from 100-plus experts is not a forecast you can plan a product around, and presenting it as one would be dishonest.

We do not forecast a date for superintelligence, and you should treat anyone selling you software on the strength of one as selling you the forecast rather than the software.

What the definition changes about what you build

The practical consequence is narrower than the debate suggests. Nothing you can license today is superintelligent under Bostrom's definition, and the systems you can license are jagged in ways the benchmarks understate. That makes the design question about containment of failure, not about capability ceilings.

In the systems we build, that shows up in unglamorous places: bounding each model call to a task short enough that the measured reliability holds, keeping a deterministic path for anything with money or safety attached, and writing acceptance criteria that test the output rather than the model. The same argument applies to how you stage the work, which is the subject of building AI-ready software architecture, and to where generative AI belongs in a product at all.

If a vendor's roadmap depends on a capability threshold nobody can currently measure, that is not a roadmap. Ask instead what task length they have measured, on what benchmark, on what date, and what happens in your product on the runs that fail.

Frequently asked questions

Superintelligence means an intellect much smarter than the best human brains in practically every field, including scientific creativity, general wisdom and social skills. That is Nick Bostrom's 1998 definition, and it is a claim about breadth of excellence rather than about beating average people at a long list of tasks.

Superintelligence is not the same as AGI. The International AI Safety Report 2026 defines artificial general intelligence as a hypothetical system that equals or surpasses human performance on all or almost all cognitive tasks. Superintelligence is the rung above: decisive advantage over the best humans, not parity with people in general.

No system today meets the definition of superintelligence, and no accepted test for it exists. The International AI Safety Report 2026 describes current capabilities as jagged, with leading systems excelling at hard tasks while failing at simpler ones, and never uses the word superintelligence in its own analysis.

Artificial superintelligence, or ASI, is software that would outperform the best human expert in nearly every field at once, including judgement and social skill, not just in one domain. A chess engine is not superintelligent because its advantage does not extend past chess.

Nobody can settle that from evidence today. The International AI Safety Report 2026 presents four OECD scenarios for 2030 ranging from stagnation to progress exceeding human cognitive performance, and states that its contributors differ on how quickly capabilities will improve. Creuto does not forecast a date.

The definition most often cited comes from Nick Bostrom's paper How Long Before Superintelligence?, published in the International Journal of Futures Studies in 1998. Bostrom later added four postscripts to that paper, the last in March 2008, revising his own confidence downwards rather than upwards.

Benchmarks measure selected points on a capability profile, not its floor. On the official ARC Prize leaderboard data generated on 30 September 2026, the best ARC-AGI-2 score was 95.0% against a human panel at 100%, while ARC-AGI-1 leaders reached 98.5% against a 98.0% panel.

Written by

Akash Mohapatra

Akash Mohapatra

Co Founder & Director

30 Sep 2026

·

10 min read

Share

LET'S CONNECT

Connect with Creuto!

Ready to take the first step towards unlocking opportunities, realizing goals, and embracing innovation? We're here and eager to connect.

We don't just aim to fit in – we strive to stand out. Experience the perfect blend of innovation, excellence, and trust that makes us truly unforgettable. Discover the difference with Creuto.

© 2026 Creuto All Rights Reserved