A leading product engineering company, creating adaptive software solutions to improve operations, providing businesses with expert development services from across domain.
A leading product engineering company, creating adaptive software solutions to improve operations, providing businesses with expert development services from across domain.
Real-SWE tested coding agents on licensed private codebases. The best AI coding agent success rate was 38.8%, and the average about 27%. What it means.

What is a realistic AI coding agent success rate on your own company's code? A new benchmark offers the first direct answer we have seen, and it is sobering. On tasks cut from private, licensed production codebases, the best agent-and-model combination solved 38.8% of attempts. The average across eight combinations was about 27%. Public leaderboards, built on open-source code the models may have seen in training, suggest far higher numbers.
Real-SWE comes from Specific Labs, an evaluation start-up whose co-founder described it as a Y Combinator Fall 2025 company. Instead of scraping issues from public GitHub repositories, each task comes with a private production codebase licensed from a real business. The New Stack reports that these include a consumer product with more than 200,000 users and a fintech platform that has processed more than 100,000 bank statements.
Because neither the code nor the solutions are public, they are much less likely to have appeared in a model's training data. Specific Labs cannot guarantee that, but the design removes the most common objection to public benchmarks. The tasks are also broad: solutions touch a median of 11 files, nearly double the six-file median of other recent benchmarks.
The scale is small — ten tasks, eight models, eight attempts per task, 640 runs in total — and each model ran in its maker's own coding tool, so the scores measure the full setup rather than the model alone.
Those eight figures average 26.9%, a number an independent analysis by Pebblous confirmed from the per-task results: 172 successful runs out of 640. Six of the ten tasks had success rates below 15%. A tax-jurisdiction bug was fixed 3.1% of the time. No model solved an analytics stream reducer in 64 attempts. Some agents went eight for eight on one task and zero for eight on the next.
The failure analysis is the most useful part for anyone planning delivery. According to The New Stack, Fable 5.1's failures were mostly missed requirements (36.7%) and integration errors (34.7%). GPT-6 Astra's split evenly between integration errors and unverified assumptions, at 34% each. Integration errors appeared in nearly half of Gemini 3.8 Flash's failed runs.
None of these is a failure to write code. They are failures to understand the system: not noticing a requirement, not wiring a change into the existing flow correctly, or assuming how something works instead of checking. Pebblous notes that Real-SWE's task instructions deliberately withhold where in the codebase a change belongs — much as a real ticket rarely tells a new engineer which file to open.
That matches what we see on client projects. Agents are strong at well-bounded work with clear acceptance criteria and good tests. They struggle with cross-cutting changes in unfamiliar code, where the hard part is knowing what the change touches. It is also why we argue for scoring an agent's trajectory, not just its answer: the wrong assumption usually shows up in what the agent read, or failed to read, long before the final diff.
Probably, on the right tasks. Real-SWE does not say agents are useless; it says an unsupervised agent handed a vague ticket on unfamiliar production code will fail most of the time. Several practical conclusions follow for clients planning AI-assisted delivery:
Real-SWE is new, single-source and small. Ten tasks cannot characterise all enterprise software, and the task designers chose deliberately hard problems. It does not prove that public benchmarks are inflated by contamination. The ranking between the top three is close enough that a different task set could reorder it, and changing the tool around a model can shift its score.
What it does establish is a floor for expectations: on unfamiliar private code, with minimal guidance, even the best current agent fails more often than it succeeds. That is a planning input, not a verdict on the technology. Our AI engineering services team uses agents daily on client work, and the gains come from pairing them with specifications, tests and engineers who know where the change belongs.
AI coding agents are far less accurate on unfamiliar private code than public benchmarks suggest. On Specific Labs' Real-SWE benchmark in September 2026, the best agent solved 38.8% of attempts on licensed production codebases, and eight agents averaged about 27%.
On Real-SWE, Fable 5.1 running in Claude Code scored highest at 38.8%, followed by GPT-6 Astra in Codex CLI at 33.8% and Gemini 3.8 Flash in Gemini CLI at 31.2%. The benchmark has only ten tasks, so the ranking is indicative rather than definitive.
Coding agents on Real-SWE failed mainly through missed requirements, integration errors and unverified assumptions rather than an inability to write code. They struggled to understand how a change fits into an unfamiliar system, especially when tasks touched many files.
AI coding agents can speed up well-specified tasks backed by good tests, but on unfamiliar production code with vague tickets they fail most of the time. Clear acceptance criteria, integration tests and experienced reviewers determine how much of the speed-up is real.
Ready to take the first step towards unlocking opportunities, realizing goals, and embracing innovation? We're here and eager to connect.
11th Floor, O-Hub, Chandaka Industrial Estate, Infocity, Bhubaneswar, Odisha 751024
Level 4, 11 York Street Sydney Startup Hub Sydney, NSW – 2000
30 N. Đinh Nghệ, Phước Mỹ Sơn Trà, Đà Nẵng / Da Nang City – 550000
Level 25, AIDP Business Tower, Dubai Marina, United Arab Emirates
50 Beauchamp Street, Wellington, WGN 5028, New Zealand