Creuto is now an OpenAI Select Partner Read More
AI compute governance now has real numbers: 10^24, 10^25 and 10^26 FLOP, plus a TPP export rule. What each says, and what it means for model access.

The most concrete proposal on the table for AI compute governance would cap new training runs at 10^24 FLOP. The largest runs already in existence have likely passed 10^26 FLOP — a hundred times more. That gap, not the politics, is the thing to understand first: every serious governance lever aims at the chips and the buildings, because a training run itself leaves almost nothing a regulator can inspect.
This post covers the numbers that have actually been written down, who wrote them, how verification is meant to work, and what it changes for a company whose product depends on frontier model access.
An International Agreement to Prevent the Premature Creation of Artificial Superintelligence, by Aaron Scher, David Abecassis, Peter Barnett and Brian Abeyta, was submitted to arXiv on 13 November 2025 and last revised on 8 May 2026. It proposes a coalition led by the United States and China, and it names two numbers:
The authors chose 10^24 because it sits "slightly below that used to train models near the state of the art as of August 2025 (such as DeepSeek-V3, trained with 3 × 10^24 FLOP)", and because it "provides some breathing room and a buffer against algorithmic progress". They put DeepSeek-R1 at around 4 × 10^24 FLOP and gpt-oss-120B at around 5 × 10^24 FLOP, and note that as of mid-2025 there were between 50 and 100 models already trained above 10^24 FLOP. Those models would not be recalled; the agreement prohibits new training, not existing weights.
The paper is explicit about its own status. The authors write that the proposal "would be technically sufficient to forestall the development of ASI if implemented today", and in the same abstract that "there does not yet exist the political will to put such an agreement in place". That is a proposal, not a forecast and certainly not a schedule.
The mechanism is physical. The paper's unit of control is a "covered chip cluster", defined around 16 H100-equivalent chips networked at more than 25 Gbit/s — chosen, the authors say, because 25 Gbit/s "is faster than non-data center internet connections" and it is "very rare and expensive for an individual to own more than 16 H100-equivalents".
The arithmetic behind that choice is the useful part. With 16 H100s at FP8 and 50% utilisation — parameters the authors call "realistic but optimistic" — reaching 10^22 FLOP takes 7.3 days, and reaching 10^24 FLOP takes two years. Small clusters can quietly cross the lower threshold; nobody crosses the upper one without a building.
That is why observability concentrates at the hardware layer. The paper notes that the most advanced logic chips in AI accelerators "are almost all fabricated by TSMC — accounting for around 90 percent of market share", mostly on five-nanometre-class nodes supported by "two or three manufacturing plants". Verification is then proposed as a stack: supply chain tracking, mandatory reporting of clusters larger than 16 H100s, sales records from distributors, state intelligence, open-source intelligence, power consumption monitoring, challenge inspections and whistleblower channels. Large clusters, the authors argue, are "the easiest to detect due to their physical footprint, power consumption, extensive personnel requirements, and reliance on standard chip supply chains".
Two of the three layers already have live instruments, and they use different units.
| Instrument | Threshold | Status |
|---|---|---|
| EU AI Act, Article 51(2) | Training compute greater than 10^25 FLOP presumes "high impact capabilities" | In force; thresholds amendable by delegated act |
| US BIS reporting rule (11 September 2024) | Training runs above 10^26 operations; clusters above 10^20 OP/s connected at over 300 Gbit/s | Proposed rule, comments closed 11 October 2024 |
| US BIS licence policy (15 January 2026) | TPP under 21,000 and total DRAM bandwidth under 6,500 GB/s | Final rule, effective on publication |
The EU AI Act's Article 51(2) states that a general-purpose AI model "shall be presumed to have high impact capabilities … when the cumulative amount of computation used for its training measured in floating point operations is greater than 10^25". The US proposed reporting rule of 11 September 2024 would have required quarterly notification for "any AI model training run using more than 10^26 computational operations", and for possession of a cluster "transitively connected by data center networking of greater than 300 Gbit/s and having a theoretical maximum greater than 10^20 computational operations … per second (OP/s) for AI training, without sparsity". It was issued under the October 2023 executive order on AI, and it is a proposed rule — treat it as a drafted threshold, not an obligation, and confirm its current status with your adviser before relying on it.
Export control, by contrast, is live and specific. The BIS final rule published on 15 January 2026 moved exports from the United States to China and Macau of advanced computing commodities "with a TPP less than 21,000 … and a 'total DRAM bandwidth' less than 6,500 GB/s, such as the NVIDIA H200 or AMD MI325X" from a presumption of denial to case-by-case review, subject to certification conditions. For reexports and in-country transfers to those destinations, the rule says the policy "remains a presumption of denial". Note what the unit is: not FLOP, but a chip performance metric defined in Technical Note 2 to ECCN 3A090.
The best argument against compute thresholds is that they measure the wrong thing over time. The paper concedes it: "historical trends in AI algorithmic progress indicate that the number of operations used to train an AI to a given capability level drops by 3× each year", which means "the threshold size of AI chip clusters that must be monitored … would also decrease by 3× " annually. A number fixed in a treaty gets looser every year it is not revised.
The International AI Safety Report 2026, published in February 2026 under Yoshua Bengio's chairmanship, puts a range around that figure: studies "suggest efficiency improvements of 3x per year based on previous data points" but "are unable to rule out rates ranging from 2–6x per year", and it calls estimates of algorithmic efficiency "highly uncertain". The same report records training compute for the most compute-intensive models growing about 5x per year, and states that the largest training runs "have now likely exceeded 10^26 FLOP" (Epoch AI, 2025). So the proposal's answer to the objection is the one it states openly: pick a threshold below the danger zone, and accept that it needs revising. Anyone quoting 10^24 FLOP as a permanent line is quoting it wrongly.
Nothing in any of these instruments restricts you from calling an API. Every threshold above governs training or hardware, and a deployed model sitting behind an endpoint is the least observable layer in the stack — which is precisely why governance has not settled there. The practical exposure is narrower and more boring than the headlines suggest.
Three things follow for architecture. First, chip export policy is jurisdictional, so where your inference runs can change availability and price independently of your contract; if you are choosing regions on AWS, GCP or Azure, treat capacity in a given country as a variable, not a constant. Second, a model behind an interface you own can be replaced; a model wired through your codebase cannot, which is the same reversibility test we apply to vendor lock-in generally. Third, open-weight models you have already downloaded are not affected by a future training cap, because weight releases cannot be recalled — the Safety Report's phrasing is that "open-weight model releases are irreversible". That is a real, if limited, continuity option, and it is one reason open-weight models already carry a large share of tokens at a small share of spend.
We do not sell compliance advice, and we would not sell you an architecture justified by a treaty that does not exist. What we do build is the boundary: a model interface with its own contract tests, evaluation data you own, and a documented fallback path. That work is the same work described in building AI-ready software architecture, and it pays off whether or not any of these thresholds is ever enforced.
If you want one action from this post: find out which country your inference calls are served from, and whether anyone has tested the fallback.
Compute is the most observable input to AI development, which is why proposals target it. Chips are made by a handful of fabs, large clusters draw enough power and staff to be visible, and cluster reporting rules already exist in draft. Regulating a trained model after deployment is much harder.
A FLOP threshold is a limit expressed in floating point operations used to train a model. The EU AI Act presumes systemic risk above 10^25 FLOP. The 2026 arXiv agreement proposal would prohibit new training above 10^24 FLOP and monitor training above 10^22 FLOP.
The arXiv proposal does not verify runs directly. It tracks the hardware: mandatory reporting of clusters larger than 16 H100-equivalents, distributor sales records, power consumption monitoring, challenge inspections, intelligence gathering and whistleblower channels. Large training clusters are detectable through physical footprint, power draw and staffing.
The arXiv proposal states that the most advanced logic chips in AI accelerators are almost all fabricated by TSMC, at around 90 percent market share, on five-nanometre-class nodes supported by only two or three plants. That concentration is what makes supply-chain tracking a credible verification route.
No current instrument restricts calling a commercial model API. Training thresholds apply to developers, and export controls apply to hardware shipments and destinations. The practical risk for a product team is regional capacity and pricing shifting, not losing access to inference outright.
They answer different questions. The EU's 10^25 FLOP mark flags models needing extra obligations, the US draft reporting rule used 10^26 operations to catch only the largest runs, and the arXiv proposal's 10^24 FLOP cap is set deliberately below current frontier scale to halt progress.
Ready to take the first step towards unlocking opportunities, realizing goals, and embracing innovation? We're here and eager to connect.
11th Floor, O-Hub, Chandaka Industrial Estate, Infocity, Bhubaneswar, Odisha 751024
Level 4, 11 York Street Sydney Startup Hub Sydney, NSW – 2000
30 N. Đinh Nghệ, Phước Mỹ Sơn Trà, Đà Nẵng / Da Nang City – 550000
Level 25, AIDP Business Tower, Dubai Marina, United Arab Emirates
50 Beauchamp Street, Wellington, WGN 5028, New Zealand