A leading product engineering company, creating adaptive software solutions to improve operations, providing businesses with expert development services from across domain.
A leading product engineering company, creating adaptive software solutions to improve operations, providing businesses with expert development services from across domain.
Duolingo's AI code review automation approves about 10% of pull requests with no defect rise. The guardrails and training behind it matter more.

Duolingo now lets an AI system approve about one in ten pull requests with no human reviewer. The interesting part is not the model. It is everything around it: a risk classifier with hard exclusions, a training programme that came first, and the willingness to measure reverts rather than assume safety. For anyone weighing AI code review automation, it is the most concrete published account of what it takes.
In an InfoQ presentation, Duolingo software engineer Sarah Deitke described the results:
That last figure is worth dwelling on. A quarter of engineers choosing a human reviewer when they do not have to is not resistance. It is a healthy signal that the system is optional, and that people use judgement about when a second pair of eyes matters.
The system classifies each pull request as low, medium or high risk, or as undetermined. It draws on the pull request title, the verification steps the author describes, the diff itself and an evaluation by a language model. Only low-risk changes that also meet additional criteria are approved automatically.
The categories are concrete:
Notice what the classifier is really doing. It is not judging whether code is correct — that remains the job of tests. It is judging whether the consequence of the code being wrong is small enough to catch after merge. That is a far easier question, and it is the right one.
The exclusions matter more than the model. Duolingo describes restricting the feature to repositories with a small number of code owners, allowing teams to include or exclude directories, disabling auto-approval for AWS resource changes entirely, excluding repositories subject to SOX or ISO regulation, blocking new engineers during their onboarding period, and offering optional daily summaries of what was approved.
Each of those encodes a decision humans made about where automated approval is unacceptable regardless of how confident the model is. Regulated code, infrastructure and the work of people still learning the codebase are exactly where a confident wrong answer is most expensive. None of those boundaries depends on the model behaving well.
The part of the talk most teams will skip is the part Deitke emphasised most. Before automating review, Duolingo invested in AI literacy: lab-style workshops on MCP servers, editor rules, batch requests and evaluations, which 95% of the engineering organisation said taught them something new; observability dashboards tracking tool usage, cost and model families; short office-hour slots; shared channels and regular meetups; and early access through vendor partnerships.
Her framing is the line to keep: before adding agents and automation, ask whether you have built AI literacy in the organisation, not just AI access. The reason is practical. Engineers who understand how the system decides are the ones who notice when it decides badly and report it — and those reports became a dataset that improves the risk-scoring prompt over time.
This is the most useful answer yet to a problem we have written about before: coding agents generate change faster than teams can review it, so the bottleneck moves to review. Duolingo's response is not to review faster. It is to stop spending human review on changes where review adds little, and to protect the categories where it adds a lot.
If you are considering something similar, the order matters.
Most teams will not reach ten per cent in six months, and should not try to. But the structure transfers to a company of any size: tests prove correctness, a classifier decides consequence, and humans spend their attention where consequence is high. That is how we would approach review automation on any AI engineering engagement, and why it starts as a quality engineering question rather than a tooling one.
Duolingo classifies each pull request as low, medium or high risk, or undetermined, using the title, the author's verification steps, the diff and a language model evaluation. Only low-risk changes that meet additional criteria are approved automatically without a human reviewer.
About 10% of all pull requests within six months. Median time to merge fell from roughly 18 hours to 12, with around one to two reverts a day across roughly 200 auto-approved pull requests and no significant rise in defect rate.
Duolingo never auto-approves user permission changes, large infrastructure modifications or core app feature changes. It also excludes AWS resource changes, SOX and ISO-regulated repositories, and pull requests from engineers still in their onboarding period.
Yes. Between 20% and 30% of Duolingo developers request a manual review even when their change qualifies for auto-approval, which keeps human review a free choice rather than something the system removes.
Engineers who understand how the system makes decisions notice and report its mistakes. Those reports became a dataset used to improve the risk-scoring prompt. Duolingo's workshops were rated useful by 95% of the engineering organisation.
Start with the list of repositories and change types that are never auto-approved, define risk by the consequence of a mistake, measure reverts from the first day, keep manual review optional, and treat every misclassification as a new evaluation case.
Ready to take the first step towards unlocking opportunities, realizing goals, and embracing innovation? We're here and eager to connect.
11th Floor, O-Hub, Chandaka Industrial Estate, Infocity, Bhubaneswar, Odisha 751024
Level 4, 11 York Street Sydney Startup Hub Sydney, NSW – 2000
30 N. Đinh Nghệ, Phước Mỹ Sơn Trà, Đà Nẵng / Da Nang City – 550000
Level 25, AIDP Business Tower, Dubai Marina, United Arab Emirates
50 Beauchamp Street, Wellington, WGN 5028, New Zealand