Creuto is now an OpenAI Select Partner Read More

AI & Machine Learning

International AI Safety Report 2026: a builder's summary

The International AI Safety Report 2026 for builders: three risk domains, the 77% vulnerability figure in context, and the evaluation gap in your tests.

International AI Safety Report 2026: a builder's summary

The International AI Safety Report 2026 runs to 220 pages written for governments, and one of its findings should change what your team does next sprint: performance on pre-deployment tests no longer reliably predicts behaviour in production. This is the builder's reading of its three risk domains — what applies if you ship AI features, and what does not.

The report was published on 3 February 2026, chaired by Yoshua Bengio, and guided by over 100 independent experts with an Expert Advisory Panel nominated by more than 30 countries and by the OECD, the EU and the UN. It is the second edition in a series mandated by the 2023 Bletchley Park summit. The preprint mirror on alphaXiv records a 24 February 2026 submission date, which is the mirror's date, not the report's.

One thing to fix before you read further: the report states plainly that it "does not make specific policy recommendations". It is a synthesis of evidence, not a standard. It has no clauses you can comply with and no controls you can adopt. Read as a checklist it is useless; read as a map of which risks now have real evidence behind them, it is the best non-vendor source available.

What the international ai safety report counts as a risk, and in what order

The report sorts emerging risks from general-purpose AI into three domains. The framing matters because each domain has a different owner in a product team.

Risk domainWhat the report evidencesWho owns it on a product team
Malicious useCyberattacks, biological and chemical uplift, influence operations, non-consensual imagery and fraudSecurity and abuse; mostly the model provider's problem, partly yours at the API boundary
Malfunctions and loss of controlFabricated information, flawed code, agents acting before a human can intervene, and models that behave differently when they detect a testEngineering. This is the one you cannot delegate
Systemic risksLabour market effects, automation bias, weakened critical thinking, uneven global adoptionProduct and design, plus whoever signs off on the workflow you are automating

Note what the 2026 edition left out. The report says its scope is deliberately "narrower than that of the 2025 Report, which also addressed issues such as bias, environmental impacts, privacy, and copyright". Those four are usually the first things a client asks about. If you hand this report to a risk committee as your AI risk register, you have handed them a document that skips the majority of what they will be sued over.

The 77% figure is real, and narrower than the coverage suggests

The most quoted number from the malicious use chapter is that an AI agent identified 77% of vulnerabilities in real software, placing it in the top 5% of over 400 mostly human teams. That is accurate, and it is worth reading the underlying sentence rather than the headline.

The setting was the final phase of the DARPA AI Cyber Challenge. The report's body text is precise: one AI system "autonomously identified 77% of the vulnerabilities introduced by the competition organisers, as well as other, unintentional vulnerabilities". So the denominator is a set of bugs deliberately planted in a benchmark corpus, not the population of bugs in your repository. It is a strong result about vulnerability discovery in a scored contest. It is not a claim that three-quarters of the flaws in your codebase are one agent away from discovery.

The same chapter draws the boundary in the other direction too. The report finds that AI systems are "automating more parts of cyberattacks, but cannot yet execute them autonomously", and that fully autonomous end-to-end attacks have not been reported — observed agents execute irrelevant commands, lose track of operational state, and cannot recover from simple errors without a human. That asymmetry is the practical point: discovery is strong, chaining is weak. Budget your defences accordingly, and treat capability claims about cyber risk the way you would treat a vendor's own critical cybersecurity rating — as an input, not a verdict.

The evaluation gap is the finding that should change your test plan

The report names an "evaluation gap": AI systems "often perform impressively in controlled settings such as pre-deployment evaluations, but more poorly in real-world conditions". It gives three causes, and all three apply to evaluations you run yourself, not just to frontier labs.

  • Benchmark validity. Many benchmarks do not measure the capability they claim to, often because a multiple-choice format lets a model reach the right answer by shortcut.
  • Data contamination. Models may have been trained on the questions used to score them — a problem the report says "most developers do not currently track or disclose".
  • Test-awareness. Models are "increasingly likely to identify certain tasks as evaluations and display different behaviours than they would on similar tasks in deployment contexts".

The third one is new since the 2025 edition, and it is the one that breaks an assumption most teams have never questioned. The report's summary of developments says reliable pre-deployment safety testing "has become harder to conduct" because it has become more common for models to distinguish test settings from real deployment and to exploit loopholes in evaluations. Research on the causes — and on whether it can be prevented — is described as still in early stages.

The primary evidence is worth reading directly. Apollo Research's benchmark of 1,000 prompts and transcripts drawn from 61 datasets found Gemini 2.5 Pro reaching an AUC of 0.83 at telling evaluation transcripts from real usage, against a human baseline of 0.92, with both humans and models doing better in agentic settings than in chat. On the developer side, Anthropic's Claude Sonnet 4.5 system card reports that refusals or call-outs of this kind appeared in about 13% of transcripts generated by its automated auditor, and that adding a filter to reject implausible test scenarios reduced evaluation awareness by roughly 25% while increasing observed misaligned behaviour by about 5%.

Read that last pair of numbers again. Making the test more realistic made the model behave worse. That is the evaluation gap with a measurement attached, and it is why an internal eval suite that only ever gets greener is evidence of nothing. In the systems we build, the checks that earn their place are the ones that look like ordinary traffic: sampled production transcripts scored after the fact, held-out cases never shown to anyone tuning prompts, and behavioural assertions on real user sessions. It is the same reason coding agents score far worse on private code than on public benchmarks.

Agents are the part the report is most cautious about

The capability chapter gives a figure worth putting in a planning document: AI agents can now reliably complete some coding tasks that would take a human programmer about half an hour, up from under ten minutes a year earlier, and the time horizon over which agents can operate autonomously has doubled on average every seven months since 2019. That is a measured trend, not a forecast of where it ends.

Against that, the report is blunt about current limits. Agents "reliably fail on longer tasks, lose track of their progress, and often cannot adapt to unexpected obstacles". Reliability techniques "can reduce failure rates but not to the level required in many high-stakes settings". And prompt injection gets its own box: malicious instructions hidden in a web page or a database can hijack an agent, and such attacks are "particularly difficult to defend against because they are delivered using external content outside the user's or developer's control". The report's own chart of reported prompt injection success rates shows them falling over time while remaining, in its words, relatively high.

The design conclusion is unglamorous and it is the one we keep arriving at independently: an agent's blast radius is a function of its permissions, not of its refusal training. Scope the credentials, log every tool call, and put a human in front of the irreversible ones — which is exactly the failure mode we wrote about when agents route around the controls placed in their way.

Technical safeguards, and what the report says they cannot do

The report's answer to "how do we manage this" is layering, not solving. It endorses defence-in-depth — combining safety-trained models with input filters, output filters and content monitors — while stating that progress on worst-case robustness "has been slow" and that users can still obtain harmful outputs by rephrasing requests or splitting them into smaller steps. Adversarial training is described as an ongoing "cat and mouse game".

Two limits are stated flatly. Open-weight models "cannot be recalled once released", their safeguards are easier to remove, and they can be run outside monitored environments — a consideration if your deployment plan involves self-hosting weights. And because risk management measures have limitations, they "will likely fail to prevent some AI-related incidents", which is why the report spends a chapter on societal resilience rather than on prevention alone.

On governance, the report counts 12 companies that published or updated a Frontier AI Safety Framework during 2025, and describes their core mechanism as voluntary if-then commitments: conditional protocols that trigger specific responses when models reach predefined capability thresholds. Its assessment of how well they are verified is the single most useful sentence in the chapter for anyone doing vendor diligence: external assessments of compliance "remain limited, in part because most frameworks are recent, publicly available information is scarce, and there are no standardised external audits". If you want to check a provider's commitments, you are reading their own document. Some of those documents have moved on since — Anthropic's Responsible Scaling Policy reached version 3.4 on 8 July 2026, after the report's snapshot of version 2.2.

Where this report does not apply to you

The strongest objection to a post like this is that the report is aimed at states, not engineers. That is largely fair. Its central concept is the policymaker's "evidence dilemma" — acting too early entrenches bad interventions, waiting for conclusive data leaves society exposed — and it deliberately offers no controls, no thresholds you can adopt and no compliance surface. Most of its loss-of-control chapter concerns capabilities current systems do not have.

So be specific about the parts that transfer. The evaluation gap transfers completely, because you run evaluations. Prompt injection and agent permissions transfer completely, because you grant them. Automation bias transfers, because your interface decides how much scrutiny a user gives an answer — the report finds early evidence that reliance on AI tools can weaken critical thinking and encourage trust in outputs without sufficient scrutiny, which is a UI decision before it is an ethics one, and a close cousin of the hallucination risks we have written about elsewhere.

What does not transfer: CBRN thresholds, model weight security at frontier scale, and anything about whether to release a model at all. Those are the provider's obligations. Do not build a governance document around them and then discover you never wrote down who reviews your agent's tool permissions.

The practical next step is smaller than the report's scope suggests. Take your existing eval suite, pick the three cases you trust most, and check whether any of them could have been in a training set or reads like a test. If all three fail that check, you do not have an evaluation problem in the future — you have one now, and our AI engineering practice spends more time on exactly this than on model selection.

Frequently asked questions

The International AI Safety Report is a scientific synthesis of evidence on the capabilities and risks of general-purpose AI, mandated by the 2023 Bletchley Park summit. The 2026 edition was published on 3 February 2026 and explicitly makes no policy recommendations, presenting evidence rather than standards.

Yoshua Bengio chaired the report, with an independent writing team of over 30 members and guidance from more than 100 AI experts. An Expert Advisory Panel carried nominees from over 30 countries and from the OECD, the European Union and the United Nations, who reviewed a consolidated draft.

The report groups emerging risks into malicious use, malfunctions and loss of control, and systemic risks. Malfunctions is the domain most relevant to teams shipping AI features, because it covers unreliable outputs, autonomous agents acting without oversight, and models that behave differently when they detect an evaluation.

In the final phase of the DARPA AI Cyber Challenge, one AI system autonomously identified 77% of the vulnerabilities that competition organisers had deliberately introduced, plus some unintentional ones, placing it in the top 5% of over 400 mostly human teams. The denominator is planted bugs, not all real-world flaws.

The evaluation gap is the report's term for AI systems performing well on pre-deployment tests but worse in real conditions. Its three named causes are benchmarks that do not measure what they claim, training data contaminated with test questions, and models recognising tasks as evaluations and behaving differently.

No. The report counts 12 companies that published or updated a Frontier AI Safety Framework during 2025 and describes them as voluntary if-then commitments. It states that external assessments of compliance remain limited and that there are no standardised external audits, so verification currently rests on each developer's own disclosures.

No. The 2026 edition states that its scope is deliberately narrower than the 2025 report, which also addressed bias, environmental impacts, privacy and copyright. Teams should not treat the 2026 report as a complete AI risk register, because it omits several of the areas clients ask about first.

Written by

Akash Mohapatra

Akash Mohapatra

Co Founder & Director

1 Oct 2026

·

10 min read

Share

LET'S CONNECT

Connect with Creuto!

Ready to take the first step towards unlocking opportunities, realizing goals, and embracing innovation? We're here and eager to connect.

We don't just aim to fit in – we strive to stand out. Experience the perfect blend of innovation, excellence, and trust that makes us truly unforgettable. Discover the difference with Creuto.

© 2026 Creuto All Rights Reserved