Creuto is now an OpenAI Select Partner Read More
Will AI replace software engineers? Measured agent results on real private code say 38.8% at best, against a 90% forecast. The work is moving, not vanishing.

In March 2025 Dario Amodei told the Council on Foreign Relations that in three to six months AI would be writing 90 percent of the code. Eighteen months later, the best measured result on real private production code is a 38.8% task resolve rate. That gap is the whole subject. Will AI replace software engineers is a question the measurements can answer better than the forecasts, and the answer is that the work is moving rather than disappearing.
This post uses only measured evidence: published agent benchmark results on real tasks, dated, with the leaderboard linked. Forecasts appear too, but labelled as forecasts and attributed to the person who made them.
No, not on current evidence — but the composition of the job is changing faster than the headcount. Agents now complete a meaningful share of well-scoped coding tasks without supervision, and fail most often not at writing code but at understanding the system they are writing into. That shifts effort from production to verification, which is a different job rather than no job.
Keep the two claims separate, because almost all the confusion comes from merging them. "AI writes most of the code" is a statement about typing. "AI resolves most real engineering tasks end to end" is a statement about finishing. Both can be true at once and they measure different things; only the second one bears on employment.
Here is the argument at its best, and it is not a straw man. The people with the most direct visibility into frontier capability are the ones making the boldest claims. At the Council on Foreign Relations on 10 March 2025, Amodei said "I think we'll be there in three to six months — where AI is writing 90 percent of the code", and that in twelve months "we may be in a world where AI is writing essentially all of the code". He was not predicting mass redundancy; he argued that engineers move up to specifying conditions, overall app design and security implications.
The supporting evidence is real. Benchmarks that separated frontier models a year ago no longer do. Agents operate for hours on multi-file changes. And the economic logic is brutal: if a task that cost a day of salary now costs a few dollars of tokens, demand for that task at the old price does not survive. A reasonable person can hold this position, and many do.
The counter is not that the capability is fake. It is that the measurements of finished work, on code the model has never seen, look nothing like the measurements of code production.
Three dated figures, from official sources, each measuring something different.
| Measurement | Result | What it measures |
|---|---|---|
| Real-SWE, September 2026 | 38.8% best, 26.9% average | Tasks on licensed private production codebases, unseen by any model |
| SWE-Bench Pro public, read 1 October 2026 | 61.5% top entry | Long-horizon tasks in public open-source repositories |
| Amodei at CFR, March 2025 | 90% of code (forecast) | A prediction about share of code written, not tasks resolved |
The Real-SWE report is the one to read closely, because it is the closest published proxy for the code you are actually paid to change. It ran eight models against tasks drawn from a private production codebase licensed from a real company. The best configuration resolved 38.8% of tasks; the average across all eight was 26.9%; the weakest scored 16.2%. One task went unsolved across all 64 attempts.
Its most useful finding is not the headline. The report attributes 94.9% of failures to misreading the system rather than to writing code poorly — the model understood the language and misunderstood the codebase. We wrote about that result in more detail when it landed, in our piece on AI coding agent success rate on private code.
The public-versus-private spread is the same story. On SWE-Bench Pro, the top public-repository entry sits at 61.5% as of 1 October 2026, well above the private-code numbers, and the Real-SWE authors note drops of 4.8 to 15.7 percentage points for the same model moving from public to private sets. Familiarity with the repository is doing a large share of the work that gets reported as capability.
None of this is an argument that agents are weak. A 38.8% resolve rate on unseen production code with no human in the loop would have been an extraordinary claim three years ago. It is an argument that the number you need for a staffing decision is not the number in the announcement.
The strongest recent demonstration is the CodeScene case study reported by InfoQ in September 2026: coding agents refactored a 300,000-line C codebase over three weeks for roughly $4,000 in tokens, producing 2,903 commits across 726 files and moving the Code Health score from 5.6 to 10.0. We have already written that one up in full, including what it would take to reproduce, in our analysis of the 300k-line agentic refactor.
What matters for this question is the practitioner criticism, which InfoQ reported alongside the result and which is sharper than the usual scepticism. The target was an open-source decompilation of a fighting game, chosen partly because the team are its users; one engineer on the project said a healthier production codebase he had worked on was rejected as too healthy to be useful. The merge happened on a fork, through 54 pull requests. Tracy Bannon objected that calling the outcome perfect is pretty bold. Asko Nõmm raised the methodological point that, with agents tuned to their own models, it is unclear how much of the result measures the model and how much measures the harness — and that architecture remains unmeasured, so code can look healthy while fundamental problems surface later.
And two figures that circulate as results of the project are not results of it. InfoQ is explicit that the roughly 70% reduction in AI-induced defects and roughly 45% less token waste are extrapolated from CodeScene's earlier research rather than measured here. The $4,000 and the three weeks were measured.
The deciding detail is the harness. The work succeeded because a decompiled game offers deterministic frame-by-frame replay, so a trace comparison could verify behaviour after every change. Most legacy systems have no equivalent oracle — which is precisely why refactoring them is risky in the first place. The capability demonstrated was real; the condition it required is the thing most codebases lack.
Put the three results together and a consistent shape appears. Generation is cheap and improving. Judging whether generated work is correct against a system nobody fully holds in their head is expensive and has not improved at the same rate. That is a bottleneck moving, not a job class ending — we made the same argument from delivery data in AI coding agent productivity moves the bottleneck to review.
The International AI Safety Report 2026 lands in the same place from a different direction. It records that agents have demonstrated the ability to complete a variety of software engineering tasks with limited human oversight, and in the same paragraph that they cannot yet complete the range of complex tasks and long-term planning required to fully automate many jobs. It also names an evaluation gap: systems perform impressively in pre-deployment evaluation and more poorly in real-world conditions, which is exactly the public-to-private drop the benchmarks show.
In the systems we build, the practical consequence is mundane. Review capacity, not authoring capacity, sets how much work a team can absorb, and the review has to be structured — acceptance criteria written before the agent runs, deterministic checks where an oracle exists, and a human decision in front of anything irreversible. Teams that automate parts of that review path do so narrowly and on measured criteria, as in AI code review automation.
The sharpest version of the worry is about entry-level work, and it deserves a direct answer rather than reassurance. The tasks a junior was traditionally handed — the well-specified function, the framework migration, the boilerplate — are exactly the ones the measurements show agents handle most reliably. That pipeline is genuinely narrowing.
What it is narrowing into is a role that starts at verification rather than production. Reading a diff against an unfamiliar system, finding the acceptance criterion nobody wrote down, and deciding whether a passing test means the right thing are all junior-accessible skills, and they are now the scarce ones. The honest caution is that nobody has published a measurement of how well that substitution works at scale, so treat any confident claim about it, including an optimistic one, as a forecast.
The skills that gained value are the ones that reduce the cost of being sure. Reading an unfamiliar system quickly, since misreading the system is where 94.9% of measured agent failures come from. Writing executable specifications and acceptance criteria, because an agent will satisfy whatever you actually wrote. Building verification infrastructure — deterministic replays, property tests, trace comparison — because a demonstration is only as good as its oracle. And system design, which the CodeScene critics correctly noted no current metric captures.
The skills that lost value are narrower than the discourse suggests: producing boilerplate, mechanical translation between frameworks, and writing a first draft of a well-specified function. Those were never the expensive part of custom software development, which is why removing them has not removed the work.
If you are making a hiring decision on this, the number to argue about is not 90% and not 38.8%. It is your own: what share of agent-written changes your team currently merges without a human finding something, and how long that finding takes. Measure that for a month, and you will have a better basis for the decision than any forecast in this post.
Measured evidence does not support wholesale replacement of developer jobs. The best published agent result on unseen private production code resolved 38.8% of tasks in September 2026, and most failures came from misreading the system. Verification work grows as generation gets cheaper.
AI coding agents resolve roughly 60% of long-horizon tasks in public open-source repositories and far fewer on private code: the Real-SWE benchmark recorded 38.8% at best and 26.9% on average across eight models. Familiarity with the repository explains much of the difference.
AI can write production code, but the measured failure mode is misunderstanding the surrounding system rather than bad syntax, so unreviewed merges carry real risk. Teams we work with keep a human decision in front of anything irreversible and automate review only on narrow, measured criteria.
Software engineering is not a dead career on current evidence. The International AI Safety Report 2026 records that agents complete many software tasks with limited oversight yet cannot handle the complex, long-horizon work needed to fully automate jobs. The expensive part has shifted to verification.
AI in software development reliably automates boilerplate, mechanical translation between frameworks and first drafts of well-specified functions. It does not reliably automate understanding an unfamiliar system, deciding what correct means, or verifying behaviour where no deterministic test oracle exists, which is where measured agent failures concentrate.
Developers should learn to read unfamiliar systems quickly, write executable acceptance criteria, and build verification infrastructure such as property tests and trace comparison. System design also stays valuable, since practitioners reviewing the 300,000-line agentic refactor noted that no current metric captures architecture.
Yes, measured: CodeScene's case study refactored a 300,000-line C codebase in three weeks for about $4,000 in tokens. It relied on deterministic frame-by-frame replay for verification, an oracle most legacy systems lack, and two widely quoted improvement figures were extrapolated rather than measured.
Ready to take the first step towards unlocking opportunities, realizing goals, and embracing innovation? We're here and eager to connect.
11th Floor, O-Hub, Chandaka Industrial Estate, Infocity, Bhubaneswar, Odisha 751024
Level 4, 11 York Street Sydney Startup Hub Sydney, NSW – 2000
30 N. Đinh Nghệ, Phước Mỹ Sơn Trà, Đà Nẵng / Da Nang City – 550000
Level 25, AIDP Business Tower, Dubai Marina, United Arab Emirates
50 Beauchamp Street, Wellington, WGN 5028, New Zealand