Creuto is now an OpenAI Select Partner Read More
A superintelligence business strategy that survives any forecast: reversible decisions, abstraction at the model boundary, humans on irreversible actions.

The useful part of a superintelligence business strategy is the part that does not depend on superintelligence arriving. Strip the forecasting out and what remains is a planning question you can answer this quarter: which decisions would you make differently if model capability keeps improving at the current rate, and which would you make the same way either way? Almost everything in the second group is just good engineering.
That is not a comforting answer and it is not a neutral one. It is a claim with a consequence: if it is right, the correct response to the superintelligence debate is to spend your budget on architecture, data and controls you can justify on this year's P&L, and to spend almost nothing on readiness programmes whose only payoff is a capability jump nobody can date.
A roadmap is a sequence of commitments, and commitments differ in one way that matters more than any forecast: how expensive they are to undo. A model provider you can swap in a sprint is a different kind of decision from a data model you have replicated into four downstream systems, even when both are described in the same slide as "our AI strategy".
So sort every AI-adjacent decision on your roadmap along two axes. First, is it reversible or irreversible — can you walk it back in weeks, or have you taught your organisation to depend on it? Second, does the payoff depend on capability, or not — does this only pay off if models get substantially better, or does it pay off at today's capability?
| Decision | Reversible? | Depends on capability? | Do it now? |
|---|---|---|---|
| Abstract the model boundary | Yes | No | Yes |
| Instrument and keep your own task data | Yes | No | Yes |
| Humans on irreversible actions | Yes | No | Yes |
| Model bets — picking a winner and building to it | Yes | Yes | Only if cheap to unwind |
| Agents acting unsupervised on production | No | No | No |
| ASI readiness programmes | No | Yes | No |
The bottom-left quadrant — reversible, and valuable at today's capability — is where a sane roadmap spends its money. It is also, inconveniently for anyone selling a transformation, indistinguishable from competent software engineering.
The International AI Safety Report 2026, chaired by Yoshua Bengio and backed by experts nominated by over 30 countries, is the most useful input here precisely because it declines to give you a date. Its summary of current capability is that systems are "jagged": strong on well-scoped tasks, still derailed by simple errors during multi-step projects, still producing false statements.
Its finding on agents is narrower than the marketing and broader than the scepticism. Agents have demonstrated the ability to complete a variety of software engineering tasks with limited human oversight, but they cannot yet complete the range of complex tasks and long-term planning required to fully automate many jobs. Both halves of that sentence are load-bearing.
The report also names the problem that should change how you read every benchmark chart you are shown this year. It calls it an "evaluation gap": existing evaluation methods do not reliably reflect how systems perform in real-world settings, because many evaluations are outdated, contaminated by training data, or narrow. Systems perform impressively in pre-deployment evaluation and more poorly in real conditions.
You can watch that happen on a public leaderboard. On the ARC Prize leaderboard, the semi-private ARC-AGI-2 set has gone from a research challenge to near-saturation: as of the official export dated 30 September 2026, the top entry scores 95% against a human panel at 100%. A benchmark that discriminated between frontier models a year earlier now barely separates them. Planning a five-year roadmap off a number with that half-life is not strategy.
On the risk that the whole debate circles — systems operating outside anyone's control — the report is deliberate. Expert opinion on the likelihood varies greatly, some experts consider such scenarios implausible, and current systems show early signs of relevant capabilities but not at levels that would enable loss of control. Its conclusion for planners is the honest one: managing the risk could require substantial advance preparation despite the uncertainty, because the likelihood, nature and timing remain unusually ambiguous.
Put provider calls behind one interface you own, with your own prompt templates, your own retry and fallback policy, and your own evaluation harness. The test of whether the abstraction is real is not that it exists — it is whether you have actually run a second provider through it in the last quarter. An abstraction nobody has exercised is a diagram.
This is the same discipline as any other dependency. We have argued before that vendor lock-in is a reversibility problem, not a vendor one, and model providers are the clearest current example: the switching cost is almost never the API, it is the accumulated prompt behaviour, the evaluation set and the downstream expectations.
Model weights are rented. The record of how your organisation actually does its work is not, and it is the asset that gets more valuable as models improve rather than less. Capture the inputs, the chosen action and the outcome for the workflows that matter, with enough structure to evaluate against later.
Most firms do not have this because nobody logged the outcome, only the event. That is a data-model decision, and it is the kind of work we describe in building AI-ready software architecture. It is also cheap to start and expensive to backfill, which is the signature of work you should do early.
The control that survives every forecast is the boring one: an agent may draft, propose, retrieve and summarise freely, and must have a human decision in front of anything that moves money, deletes data, writes to a customer, or changes access. Note that this list is a business artefact, not a technical one. Most teams have never written it down, which means the boundary is whatever the last integration happened to implement.
Short terms, exportable data, no exclusive dependence on one provider's proprietary feature for a core workflow. You will pay slightly more. Treat the difference as the price of being wrong cheaply — which, given the spread of expert opinion the safety report documents, is the only thing anyone can honestly price.
Here is the best version of the opposing argument, and it is not weak. If capability improves discontinuously rather than smoothly, incrementalism loses badly. Reversibility is only valuable if you get time to reverse; a firm that spent 2026 building careful abstractions while a competitor rebuilt its entire operating model around agents will have optimised for an option it never gets to exercise. On this view, the asymmetry runs the other way: the downside of over-preparing is wasted budget, and the downside of under-preparing is irrelevance. Serious people hold this position, and the proposal by Scher, Abecassis, Barnett and Abeyta for an international agreement to prevent the premature creation of artificial superintelligence — compute thresholds, chip tracking, verification — exists because they take a discontinuity seriously enough to want it governed.
Two things answer it. First, the preparation we are describing is not a hedge against a capability jump; it is independently profitable at today's capability, so it does not trade off against the aggressive strategy — it is the substrate the aggressive strategy runs on. A firm with a clean model boundary and labelled outcome data can rebuild around agents faster than one without, not slower.
Second, the discontinuity argument does not tell you what to build, only that you should hurry. Every concrete proposal we have seen that claims to be ASI-specific turns out, on inspection, to be an evaluation harness, a data pipeline, a permissions model or an abstraction layer — all of which are on the list above. If the argument cannot name a distinct artefact, it is not an argument about what to build.
We sell custom software development, and we use coding agents daily in the systems we build. We do not sell ASI readiness, and we would turn down three engagements specifically.
The regulatory picture deserves the same discipline. Most of what circulates as superintelligence governance is proposal rather than enacted law, and we have set out which is which in our piece on superintelligence regulation: law versus proposals. If you are not sure the term itself is being used consistently in your own planning documents — and it frequently is not — start with what superintelligence actually means before you budget against it.
The next decision is small and you can take it this week: pull your roadmap, mark every AI-adjacent commitment as reversible or not, and for each one ask whether it pays off at today's capability. Anything that is irreversible and only pays off later is the item to argue about in your next planning session. Everything else you were going to build anyway.
Businesses should prepare for superintelligence by making reversible decisions: abstract the model provider behind an interface you own, capture outcome data from your own workflows, and keep a human decision in front of any irreversible action. Each of those pays off at today's capability, independent of any forecast.
ASI readiness is not a service Creuto sells, because there is no defensible way to score a company against a capability nobody can specify or date. Every concrete artefact such programmes propose is an evaluation harness, a data pipeline, a permissions model or an abstraction layer you should own anyway.
Build the four things that pay off regardless: a model-boundary abstraction you have actually run a second provider through, instrumented outcome data for workflows that matter, a written list of irreversible actions that require human approval, and commercial terms short enough to unwind.
The International AI Safety Report 2026 names an evaluation gap: evaluation methods do not reliably reflect real-world performance, and systems that look strong pre-deployment do worse in production. Benchmarks also saturate quickly, so a score that separated models last year may not this year.
No. The International AI Safety Report 2026 records that expert opinion on loss-of-control scenarios varies greatly, that some experts consider them implausible, and that the likelihood, nature and timing remain unusually ambiguous. It gives planners a risk framing rather than a date.
A reversible AI architecture decision is one you can walk back in weeks rather than quarters, because nothing downstream has been built to depend on its specifics. Swapping a model provider behind an interface you own is reversible; replicating a provider's proprietary data model into four systems is not.
Ready to take the first step towards unlocking opportunities, realizing goals, and embracing innovation? We're here and eager to connect.
11th Floor, O-Hub, Chandaka Industrial Estate, Infocity, Bhubaneswar, Odisha 751024
Level 4, 11 York Street Sydney Startup Hub Sydney, NSW – 2000
30 N. Đinh Nghệ, Phước Mỹ Sơn Trà, Đà Nẵng / Da Nang City – 550000
Level 25, AIDP Business Tower, Dubai Marina, United Arab Emirates
50 Beauchamp Street, Wellington, WGN 5028, New Zealand