A group of engineers at a large company did something most organisations cannot: they measured what happens to a codebase when a lot of it starts being written by a model. Their paper, Characterizing the Quality Profile of AI-Generated C++ in Production, looks at 3.52 million code changes and 10.46 million lines of old, messy C++ between April 2025 and April 2026.
That is rare. Most of what we hear about AI and code is either a sales pitch or a story someone tells at a conference. This is a full year of production data.
The main finding is not that AI-generated code is bad. It mostly works. The finding is that the cost of it turns up somewhere else than where it is produced. And if you only look at where it is produced, you will think things are going very well.
That makes it a question about how you organise, not just about which tool you buy.
The researchers sorted static analysis findings into categories and compared AI-written changes against human-written ones. Two numbers stand out:
In other words, the model tends to write code that knows a bit too much about other code, and that copies more than it needs to. The paper also describes “a stylistic reliance on explicit loops over optimized standard APIs”: the model solves things by hand, in small steps, instead of using the library function that already exists.
Now the part that surprises people. Correctness and safety came out at 0.94x, slightly better than human code. Build failures were higher (a median of about 1.3x), but once the code was in, it stayed in: “AI-generated code demonstrated a consistently lower revert rate than human-written code.”
So nothing is on fire. Nothing is breaking. The code is simply getting heavier. By early 2026, compute for AI-heavy functions had grown to about 1.31x of the baseline against 1.25x for human code, and the abstract puts the overall effect at a 5 to 8% increase in compute resource consumption.
The authors say it plainly:
The true cost of AI generation, therefore, shifts from acute reliability failures to chronic maintenance and operational burdens.
Chronic, not acute. Nothing sets off an alarm when your system gets slightly harder to work with each week. You can see it in their measurements: the gap between AI-heavy and human-written code does not appear at a single moment, it simply widens month after month.
This is where the study stops being about code and starts being about how work moves through an organisation.
AI-generated changes received 1.92x as many blocking review threads, which the paper defines as “reviewer feedback that must be resolved before submission”. They also drew 1.39x as many comments in total, needed 1.24x as many reviewer iterations, and took 1.19x as long to merge.
Look at that as a system, not as a statistic. Writing code got much cheaper. Reviewing it did not. So the work piles up in front of the reviewers, and a queue in front of a busy resource is exactly where delivery time disappears.
From the outside it looks like the reviewers got slow. They did not. They are now absorbing, one change at a time, work that used to happen while the code was being written.
What helps here is putting systems thinking before local efficiency, which happens to be one of the natural habits of a LeSS organisation. Making one part faster (how quickly a single developer produces a change) pushed cost into the whole. And if your organisation measures output per developer, that shift is not just invisible; it reads as progress.
Of everything measured here, interface and coupling burden is the number a manager should pay attention to, because what it costs you depends entirely on how the organisation is set up.
If each team owns its own component, extra tangle in the code turns into coordination between teams: more handovers, more waiting, more meetings that exist only because two pieces of code know too much about each other. The structure quietly absorbs the mess and everybody calls it “dependencies”.
Where teams work across the whole product instead, as feature teams do in LeSS, the same tangle stays what it is: a code problem, in reach of the people who just found it. That is a much cheaper kind of problem. It also keeps the maintenance burden the study warns about from hardening into an organisational one.
AI that produces coupling faster does not create any of this. It raises the price of how you were already organised.
The most hopeful result in the paper is also the most familiar one.
The researchers did not fix quality by adding a tougher check at the end. They gave the model specific feedback about the exact categories of problem it was producing. The result: the efficiency score went up 31%, and static findings dropped 11.1% compared to the prompt without feedback (12.5% against the baseline).
Short loop, specific signal, measurable change. That is inspect and adapt, aimed at a model instead of a team, and it works for the same reason it works with teams: the feedback arrived close to where the work happened, and it was concrete enough to act on.
The version of this for people is not a longer review queue. It is pairing and mob programming, where review happens in the same hour as the writing instead of three days later in a comment thread. It is a Definition of Done that names the qualities you actually care about, rather than hoping a reviewer spots them. And it is collective code ownership, so that whoever notices the coupling is allowed to fix it.
There is a pleasing symmetry here. What worked on the model is what has worked with teams for twenty years, and what a LeSS adoption tends to arrive at on its own: shorten the loop, make the signal concrete, and let the people doing the work improve how the work is done. Do that and the chronic maintenance cost the study describes stops quietly accumulating.
If you take one thing from the study, take this: your current metrics will not show you this shift. Throughput looks fine. Defects look fine, better even. Reverts improve. Meanwhile review queues grow, compute creeps up, and the codebase gets a little harder to change every month.
Three questions worth asking:
None of these are AI questions. They are the ordinary questions to ask about any product group: where is the queue, what does the system reward, and are people allowed to fix what they find. They are also, more or less, the questions a LeSS adoption starts with.
The study is careful, specific and refreshingly free of hype. It is worth reading in full: Characterizing the Quality Profile of AI-Generated C++ in Production.
If it describes something you recognise (output up, delivery no faster, reviewers underwater, a system slowly getting stiffer) then the useful conversation is not about the tooling. It is about how the organisation is put together, and that conversation needs the people who can change the structure.
Send it to them.