Creuto is now an OpenAI Select Partner Read More
ARC-AGI-2 scores rose from 27.6% to 95% by September 2026 on the official ARC Prize leaderboard. What it tests, the human baseline, and the saturation caveat.

ARC-AGI-2 is the benchmark where a frontier model went from 1.6% to 95% while humans stayed where they always were. As of the official ARC Prize leaderboard data generated on 30 September 2026, the top ARC-AGI-2 score on the Semi-Private Evaluation set is 95.00%, held by GPT-6 Astra (Max). In May 2025, the best score in the benchmark's own paper was 3.0%. This post dates every figure, explains what the benchmark measures, and says what a 95% means and does not mean.
Two things to fix before the numbers. First, the dates on the leaderboard are model release dates, not verification dates, so treat them as the earliest a score could have been set. Second, the leaderboard itself carries a footnote that some ARC-AGI-2 figures are a “score estimate based on partial testing”. Neither caveat changes the shape of the curve; both should stop you quoting a number to two decimal places as if it were a census.
ARC-AGI-2 tests fluid reasoning: inferring an unstated rule from a handful of examples and applying it to a new case. Each task is a set of input–output grid pairs, sized from 1×1 to 30×30 with ten possible cell colours, and the test-taker gets typically two to five demonstration pairs before being asked to produce the output for an unseen input.
The benchmark paper by François Chollet, Mike Knoop, Gregory Kamradt, Bryan Landers and Henry Pinkard — submitted 17 May 2025 and revised on 15 January 2026 — names four things the second version was designed to demand that the first did not:
The practical effect is that ARC-AGI-2 tasks take deliberate thought. Average human completion time in the paper's sample was 2.7 minutes, against ARC-AGI-1 tasks that human testers often solved almost instantly.
ARC-AGI-2 resists memorisation by construction rather than by policy, and the paper is unusually frank about why the first version needed replacing.
Every ARC-AGI-2 task is, to the authors' knowledge, entirely novel. Tasks were also deliberately designed to be less susceptible to brute-force program search: in ARC-AGI-1's 2020 competition the top individual submission scored 20%, but a meta-analysis found 49% of the Private Evaluation set was solved by at least one team, mostly with variations of brute-force search. Roughly half the benchmark's signal was measuring compute, not abstraction.
The leakage problem was worse. The same 100 ARC-AGI-1 Private Evaluation tasks were reused unchanged across four competitions from 2020 to 2024, and the paper estimates around 10,000 scores derived from that hidden set were disclosed over time. Each disclosed score is a channel through which participants can infer the set's characteristics. ARC-AGI-2 instead calibrates Public, Semi-Private and Private subsets so mean human accuracy differs by about one percentage point between them, and allocates newly authored tasks preferentially to private sets.
All figures below are from the official ARC Prize leaderboard's Semi-Private Evaluation data, generated 30 September 2026, with each model's leaderboard release date. Only record-setting entries are listed.
| Date | System | ARC-AGI-2 | Cost per task |
|---|---|---|---|
| 3 Nov 2024 | NVARC | 27.64% | $0.20 |
| 24 Nov 2025 | Opus 4.5 (Thinking, 64K) | 37.64% | $2.40 |
| 11 Dec 2025 | GPT-5.2 Pro (High) | 54.16% | $15.72 |
| 12 Feb 2026 | Gemini 3 Deep Think (2/26) | 84.58% | $13.62 |
| 9 Jun 2026 | Claude Fable 5 (Max) | 89.17% | $5.45 |
| 2 Sep 2026 | GPT-6 Astra (Max) | 95.00% | $1.12 |
Three figures circulating in secondary coverage need correcting against that data. The “around 54%” milestone is usually placed in early 2026; the leaderboard puts it in December 2025 — 54.00% for Gemini 3 Pro (Refinement) dated 4 December and 54.16% for GPT-5.2 Pro (High) dated 11 December. The 84.6% figure is 84.58% and belongs to a model dated 12 February 2026, so mid-February rather than late. Only the third holds exactly: 95% in September 2026 is right, and it is 95.00% for GPT-6 Astra (Max) dated 2 September.
The cost column is the more interesting trend. The 54.16% score in December 2025 cost $15.72 per task. The 95.00% score in September 2026 cost $1.12, and GPT-6.1 Sol (Max), dated 29 September 2026, reached 94.17% at $0.25 per task. Capability rose and the price per solved task fell by more than an order of magnitude.
Three different human numbers get quoted as “the human baseline”, and they measure different things.
The leaderboard separately lists a Human Panel at 100% and $17 per task, which is a panel result and not comparable to a single person's accuracy. If you want one sentence: individual humans get roughly two thirds of attempted test pairs right, and a small group gets all of them.
One finding from the paper deserves more attention than it gets. Across 407 participants in 515 sessions, no self-reported demographic factor — occupation, industry, technical experience, programming proficiency, mathematical background, puzzle-solving aptitude — showed a statistically significant relationship with performance. That is the strongest available evidence that the tasks measure general problem-solving rather than trained skill.
No, and this is where the honest caveat belongs: a benchmark approaching saturation tells you about the benchmark at least as much as about the model.
The paper's own history makes the point. ARC-AGI-1 was listed by its authors as having saturated “below human-level fluid intelligence”, because humans at the higher end of the distribution could solve over 97% of its tasks without much effort. Its designers built version two precisely because a number climbing towards the ceiling had stopped discriminating between systems. The ARC Prize Foundation has since moved on again — the leaderboard now carries a third dataset.
There is a second, narrower caveat in the paper: ARC-AGI-2 accuracies below 5% are not treated as meaningful, because they likely come from noise-level heuristics or incidental pattern fits. Consistent signal, in the authors' experience, starts above 5%. The same logic applies at the top of the range. A 95.00% against a 100% human-solvable set leaves twelve tasks' worth of headroom on a 240-pair set, and differences between the leaders now sit inside a couple of tasks.
The strongest version of the opposing argument is worth stating: if a model solves 95% of tasks specifically designed to be novel, resistant to brute force and drawn from a leak-controlled set, then something general is being measured and dismissing it as benchmark decay is motivated reasoning. That is fair. The answer is that ARC-AGI-2 was built to measure efficient adaptation on grid puzzles with minimal prior knowledge, and its authors never claimed a high score constitutes general intelligence — they claimed the benchmark provides “a more granular signal”. A score is evidence about a capability, not a verdict on a system.
Very little directly, and that is the useful conclusion. No production workload looks like a 30×30 coloured grid. What the trajectory does tell you is that abstraction on genuinely novel inputs, which was a reliable model weakness eighteen months ago, is no longer one you can design around by assumption — and that per-task reasoning cost is falling fast enough to change build-versus-buy arithmetic within a quarter.
What it cannot tell you is whether a model works on your problem. That still requires your own evaluation set, which is why we score the whole trajectory rather than the final answer when evaluating agents, and why we treat model-graded evaluation as something to correct for bias rather than trust, as in our notes on LLM-as-a-judge biases. A public benchmark leaderboard is also a poor guide to model selection on its own: a newer model can be cheaper and faster without being more accurate on your task, which is the pattern we found comparing Opus 5.5 against Opus 5.
If you are choosing a reasoning model this quarter, run your own dated evaluation on tasks from your own domain before the benchmark decides for you. That is the work we do first in AI engineering engagements, and it is the only number that survives the next leaderboard update.
ARC-AGI-2 is a reasoning benchmark from the ARC Prize Foundation in which each task presents a few input-output grid pairs and the test-taker must infer the unstated transformation rule and apply it to a new grid. Grids run from 1x1 to 30x30 with ten possible colours.
Aggregated across all human evaluations and attempts, humans completed 66% of attempted test pairs, rising to 75% when aggregated by task. Every task in the benchmark was solved by at least two independent non-expert testers, so the full set is human-solvable.
On the official ARC Prize leaderboard data generated 30 September 2026, the top ARC-AGI-2 score on the Semi-Private Evaluation set is 95.00% for GPT-6 Astra (Max), a model dated 2 September 2026, at about $1.12 per task.
No. The benchmark measures efficient adaptation on novel grid puzzles requiring minimal prior knowledge. Its authors describe it as providing a more granular signal, not as a test whose completion constitutes general intelligence, and a saturating benchmark says as much about the benchmark.
Every ARC-AGI-2 task is novel to the authors' knowledge, tasks are designed to resist brute-force program search, and the Public, Semi-Private and Private subsets are difficulty-calibrated to within about one percentage point of mean human accuracy so scores on one predict the others.
ARC-AGI-1 saturated below human-level fluid intelligence, nearly half its private set proved vulnerable to brute-force search, and its 100 private tasks were reused unchanged across four competitions with an estimated 10,000 scores disclosed, creating an information leakage channel.
Ready to take the first step towards unlocking opportunities, realizing goals, and embracing innovation? We're here and eager to connect.
11th Floor, O-Hub, Chandaka Industrial Estate, Infocity, Bhubaneswar, Odisha 751024
Level 4, 11 York Street Sydney Startup Hub Sydney, NSW – 2000
30 N. Đinh Nghệ, Phước Mỹ Sơn Trà, Đà Nẵng / Da Nang City – 550000
Level 25, AIDP Business Tower, Dubai Marina, United Arab Emirates
50 Beauchamp Street, Wellington, WGN 5028, New Zealand