Creuto is now an OpenAI Select Partner Read More

AI & Machine Learning

Frontier AI safety framework: the if-then rules labs publish

A frontier AI safety framework maps capability thresholds to responses. We compare OpenAI, Google DeepMind and Anthropic, and name who verifies compliance.

Frontier AI safety framework: the if-then rules labs publish

A frontier ai safety framework is a document in which an AI lab states, in advance, which capability thresholds would trigger which responses — and then judges for itself whether a threshold has been crossed. Three are worth reading in full. All three leave the final deployment decision inside the company, and as of 1 October 2026 none of them is audited against a common standard.

These documents are public, short and almost never read by the people who most need them: buyers doing vendor diligence on a model provider. This post compares what OpenAI's Preparedness Framework, Google DeepMind's Frontier Safety Framework and Anthropic's Responsible Scaling Policy actually commit to, and is specific about who checks.

What a frontier ai safety framework commits to, structurally

The International AI Safety Report 2026 counts 12 companies that published or updated one of these documents during 2025, and defines their core mechanism precisely: if-then commitments are "conditional protocols that trigger specific responses when AI models and systems reach predefined capability thresholds". The chain is short.

  1. Evaluate. Run capability evaluations against a defined set of risk domains, usually CBRN, cyber and some form of autonomy or AI self-improvement.
  2. Decide whether a threshold is crossed. An internal body reads the results and rules.
  3. Apply the mapped response. Security controls, deployment restrictions, monitoring, or halting development.
  4. Make the deployment decision. Sign off, or do not ship.

The report also notes what the structure is missing. These frameworks "typically do not use explicit quantitative risk thresholds", unlike risk management in aviation or nuclear power, and they vary in their definitions of risk tiers, evaluation frequency, and the buffer between an evaluation and a threshold. They are commitments about process, not numbers you can check.

What the three frameworks actually say

Each lab uses its own vocabulary for the same idea. Reading them side by side is the fastest way to see where the promises differ.

FrameworkVersion readThreshold languageWho decides and who checks
OpenAI Preparedness FrameworkVersion 2, dated 15 April 2025High and Critical capability thresholds in three Tracked Categories: biological and chemical, cybersecurity, AI self-improvementThe Safety Advisory Group recommends; OpenAI Leadership can approve or reject; the Board's Safety and Security Committee oversees. Third-party evaluation happens "if we deem that a deployment warrants" it
Google DeepMind Frontier Safety FrameworkVersion 3.1, published 17 April 2026Critical Capability Levels for severe risk, plus new Tracked Capability Levels for significant risk, across CBRN, cyber, harmful manipulation, ML R&D and misalignment"The appropriate governance function" determines residual risk is acceptable before external deployment. External parties are involved "where required or appropriate" and "as needed"
Anthropic Responsible Scaling PolicyVersion 3.4, effective 8 July 2026Capability thresholds mapped to mitigations in an industry-wide recommendations table; ASL standards retained in an appendixThe Responsible Scaling Officer approves development and deployment decisions. Published Risk Reports get external review under defined conditions, with reviewer selection approved by the Long Term Benefit Trust

Two details in that table are more interesting than the tier names.

OpenAI's framework contains an explicit competitive escape hatch. Section 4.3 states that if another developer releases a system with High or Critical capability without comparable safeguards, OpenAI "could adjust accordingly the level of safeguards that we require in that capability area" — but only if it assesses that doing so does not meaningfully increase overall risk, only if it "publicly acknowledge[s] that we are making the adjustment", and only while keeping its safeguards "at a level more protective than the other AI developer", sharing information to validate that claim. That is an unusually candid clause, and it is a conditional commitment about a conditional commitment.

Google's framework introduced, in version 3.1, a threshold that reads like it was written after the safety report's evaluation chapter: the Stealth and Situational Awareness Tracked Capability Level, defined as the point at which a model's situational awareness — its "ability to discover and use relevant details of its deployment setting" — and stealth are sufficient that "we cannot rule out the model significantly undermining human control". Crossing it triggers periodic residual risk assessments, including for high-risk internal deployments, with chain-of-thought monitoring named as a possible added safeguard.

Anthropic rewrote the shape of the promise, and the report predates it

The safety report's comparison table lists Anthropic's Responsible Scaling Policy at version 2.2, with the familiar ASL-1 to ASL-4 ladder. That was accurate when the report closed. Version 3.0, effective 24 February 2026, was a comprehensive rewrite, and version 3.4 took effect on 8 July 2026 — four revisions in under five months. If you cited the report's table in a vendor questionnaire today, you would be describing a policy that no longer exists in that form.

The rewrite's reasoning is stated openly and it is a collective action argument. The previous policy committed to mitigations that reduced Anthropic's own absolute risk "without regard to whether other frontier AI developers would do the same", and the new one separates company plans from industry-wide recommendations because, in its words, "if one AI developer paused development to implement safety measures while others moved forward with training and deploying AI systems without strong mitigations, that could result in a world that is less safe". The industry-wide column is explicitly not something the company commits to following unilaterally; a set of competitor-contingent commitments in an appendix covers the case where rivals move together.

Two of the three new instruments are worth knowing by name. Frontier Safety Roadmaps are described as "not hard commitments but rather public goals against which we will openly grade our progress". Risk Reports are the substantive artefact: published assessments of whether the risks of training or deploying a model are justified, with redactions whose existence must now be disclosed in the public version.

Who verifies compliance: the specific answer

The safety report's verdict is one sentence and it should end most vendor-diligence conversations: external assessments of developers' compliance "remain limited, in part because most frameworks are recent, publicly available information is scarce, and there are no standardised external audits". It adds that a study of earlier voluntary commitments found uneven fulfilment across measures.

Reading the frameworks themselves confirms the shape of the gap, step by step.

  • Evaluation: partly external. OpenAI describes deep dives that may include "assessments by independent third party evaluators". Google says early warning evaluations involve "internal and external experts as needed".
  • Threshold determination: internal, in all three. The Safety Advisory Group, the appropriate governance function, the Responsible Scaling Officer.
  • Response adequacy: mixed. Anthropic's external review of a Risk Report covers "the strength of the Risk Report's reasoning and analysis" and whether the reviewer disagrees with its key claims, with the reviewer free to publish concerns and required to have no financial interest in Anthropic. That is the strongest external commitment of the three by a wide margin.
  • Deployment decision: internal, in all three. No framework gives an outside party a veto.

Anthropic's external review is also conditional rather than routine: the guaranteed trigger is a Risk Report that covers a model crossing the automated AI R&D threshold and is "significantly redacted", with reviewers receiving the report within a week of the Board and Trust and asked for public commentary within 30 days. Google's governance section runs to four sentences and names no body at all, describing "a well-established and comprehensive internal governance structure". None of this is an audit in the sense a finance team would recognise.

The case for reading them anyway

The strongest argument against treating any of this as theatre is that the frameworks have already changed behaviour. The safety report records that in 2025 several developers announced that new models triggered early warning alerts, or that they could not rule out that further evaluation would show a threshold had been crossed, and applied heightened safeguards as a precaution. Multiple companies shipped models with extra safeguards specifically because pre-deployment testing could not exclude meaningful uplift for novices developing biological weapons. A voluntary commitment that costs something and gets honoured is not nothing.

Regulation is also starting to bite on the paperwork rather than the thresholds. The report notes California's SB 53 setting transparency requirements on safety frameworks and incident reporting, and the EU's General-Purpose AI Code of Practice providing guidance on evaluations, risk assessment, information security and serious incident reporting for the most advanced models. The direction of travel is disclosure first, verification later.

Still, a framework is a promise the promiser also gets to interpret, amend and date-stamp. Four Anthropic revisions in five months is a sign of a living document, and also a reminder that the version you diligenced is not the version in force.

What to ask a model provider, and what to skip

The questions that get useful answers are the narrow ones. Which version of your framework is in force today, and what changed in the last revision? Which risk domains are in scope, and which are explicitly out? Who signs the deployment decision by role? Under what conditions does an external party see the underlying assessment, and can it publish without your approval? What happens to a customer's deployment if a threshold is crossed after launch?

Skip the questions whose answers are already public and identical across vendors: whether they do red-teaming, whether they have an internal review body, whether they take safety seriously. Every framework says yes. As with any partner evaluation, the signal is in what a document declines to promise.

The architectural consequence matters more than the answers. Because thresholds are self-assessed and frameworks are self-amendable, a provider's safety posture is a dependency you do not control — and the mitigation is the ordinary one for an uncontrolled dependency. Keep model choice behind an interface, so that a change in a provider's deployment terms is a configuration change rather than a rewrite; an AI gateway that routes on identity rather than keys is the cheap version of that. Keep your own evaluation harness, because the vendor's threshold is about catastrophic capability and yours is about whether the feature works — which is why we score an agent's trajectory rather than its final answer. And read the system card for the model you actually deploy, not the framework, since that is where a capability rating like the first Critical cybersecurity classification shows up first.

If you are writing the AI section of a vendor questionnaire this quarter, the version number and the sign-off role are the two fields worth adding. Our AI engineering practice ends up filling those in for clients more often than it changes the model.

Frequently asked questions

A frontier AI safety framework is a voluntary document in which an AI developer states which capability thresholds would trigger which safety responses, before those capabilities exist. The International AI Safety Report 2026 counts 12 companies that published or updated one during 2025, built around conditional if-then commitments.

A capability threshold is a described level of ability at which a model would meaningfully raise the risk of severe harm, such as materially helping a novice build a weapon. Crossing it triggers mapped responses: extra security controls, deployment restrictions, monitoring, or halting further development.

No. Every framework compared here is voluntary, self-assessed and self-amendable. California's SB 53 and the EU General-Purpose AI Code of Practice are beginning to impose transparency and incident-reporting duties around them, but no regulator currently certifies that a lab met its own published thresholds.

Mostly the labs themselves. The International AI Safety Report 2026 states that external assessments of compliance remain limited and that there are no standardised external audits. Anthropic commits to external review of its published Risk Reports under defined conditions, which is currently the strongest third-party commitment of the three.

OpenAI uses High and Critical thresholds across three tracked categories. Google DeepMind uses Critical and Tracked Capability Levels with a safety case before deployment. Anthropic's version 3 separates its own plans from industry-wide recommendations and centres on published Risk Reports with external review.

OpenAI's Preparedness Framework says it could adjust the level of safeguards it requires if a competitor releases a High or Critical capability system without comparable safeguards, but only if it publicly acknowledges the adjustment and keeps its own safeguards more protective than that developer's.

Ask which version is in force today and what the last revision changed, which risk domains are in and out of scope, who signs the deployment decision by role, when an external reviewer sees the underlying assessment, and whether that reviewer can publish without the provider's approval.

Written by

Akash Mohapatra

Akash Mohapatra

Co Founder & Director

1 Oct 2026

·

9 min read

Share

LET'S CONNECT

Connect with Creuto!

Ready to take the first step towards unlocking opportunities, realizing goals, and embracing innovation? We're here and eager to connect.

We don't just aim to fit in – we strive to stand out. Experience the perfect blend of innovation, excellence, and trust that makes us truly unforgettable. Discover the difference with Creuto.

© 2026 Creuto All Rights Reserved