Back To All

Generation is up 180%. Review is up 30%. That gap is your org chart.

September 15th, 2026
Topics
AI
AI ADOPTION
AI Agents
AI SDLC
Share Article

Generation is up 180%. Review is up 30%. That gap is your org chart.

Bessemer surveyed 175 functional leaders across more than 100 portfolio companies and concluded that engineering org design is now a competitive variable. I think that is right, and I think the report understates how much of the work it implies is measurement work that almost nobody has started.

180% / 30%Code generation growth against code review growth, from the research they cite
49%Of their portfolio companies already delivering more without adding headcount
90% vs 24%Engineering teams actively deploying AI, against finance leaders doing the same

Alex Circei, CEO of Waydev  /  September 2026  /  11 minute read

Bessemer’s talent team published Talent trends for the AI-native C-suite, built on a survey of nearly 175 functional leaders across more than 100 portfolio companies, with input from Artisanal Talent and a set of operating advisors. It covers five functions. I want to spend most of this on the engineering section, because it contains a number that should reorganize how you think about hiring.

Code generation is up 180%. Code review is up 30%.

Everything else in the engineering half of that report is downstream of those two numbers.

The gap is structural, and it lands on the org chart

A six-fold difference in growth rates between two adjacent stages of the same pipeline is not a tooling gap. It is a capacity mismatch, and capacity mismatches resolve themselves as queues.

+180% Code generation +30% Code review Six times Two adjacent stages of the same pipeline, growing at rates that differ by a factor of six.

Figure 1. The figures Bessemer cites from research on productivity effects across generations of AI coding tools. Writing code and shipping code are not the same activity, and only one of them got faster.

The report’s framing is that the question engineering leaders now face is whether they have actually redesigned the organization to take advantage of AI, or whether they are running a 2022 team structure inside a 2026 product environment. That is the right question, and the honest answer in most companies is the second one.

Look at what a typical org chart optimizes for. Ratios of engineers to managers, sized around how much implementation one person could produce. Scoping practices built around multi-week estimates of typing effort. Review treated as a peer courtesy rather than a staffed function with its own capacity plan. Every one of those assumptions was calibrated against the constraint that just moved.

The 2022 shape Specification and scoping Implementation the expensive middle Review and verification Coordination and management The shape the numbers imply Specification and scoping ambiguity now compiles instantly Implementation Review and verification the new expensive middle Coordination Same total capacity, different allocation.

Figure 2. My reading of what the 180 and 30 figures imply for org shape. Bessemer’s point about coordination overhead shrinking is in here too, and it is the one most leaders will find politically hardest.

You cannot keep a 2022 ratio of implementers to reviewers and call the result an AI-native engineering org.

The best practice in the report, and the step everyone skips

The single most copyable thing in the whole report is one paragraph about method. The engineering leaders getting the clearest results start with a deliberately small team, sometimes an undersized scrum of three or four engineers, measure the output that team can actually hit, and only then scale the model across the organization. They track tooling and token cost against output gain, and roll the practice out further only when the return is positive.

That is a controlled experiment. It is also, in my experience, the part that gets dropped first, because the pilot produces an enthusiastic team and a good demo, and nobody wants to be the person who asks for the denominators.

Small pod three or four engineers Measure both sides output gain, and tooling plus token cost Positive ROI? against a fixed baseline Yes: scale the workflow to the next set of teams No: change it or stop it before it becomes policy The gate is the whole method. Without it you are not running a pilot, you are running a rollout with a smaller first cohort.

Figure 3. The controlled rollout pattern described in the report. Note that both the numerator and the denominator have to be measured, which is why most organizations quietly skip to the last box.

One of their operating advisors, Jessica Popp, puts the obligation plainly: “Leaders can’t just evangelize AI.” The responsibility is to show that each use case delivers a positive and measurable return. I would go further and say that in 2027 this stops being a leadership virtue and becomes a reporting requirement, because boards have started asking and the answers so far have been embarrassing.

Super IC is right, and it has a failure mode

The archetype the report describes for engineering leadership is the player-coach builder, sometimes called the Super IC: someone deeply involved in the codebase, using the tools personally, shipping alongside the team, while still making the architectural and organizational calls. Shopify’s Farhan Thawar is quoted on the urgency of harnessing agents this year, and the report argues the leader who can do both people-scaling and this transition will pull ahead.

I agree with the archetype. I run a company where I am still close enough to the work to have opinions about it, and I would not hire an engineering leader in 2026 who could not open a terminal.

But there is a specific failure mode worth naming, because it is going to produce some spectacular flameouts over the next eighteen months. A leader who is personally 10x more productive with agents can mistake their own throughput for the system’s throughput. They ship impressively, their pull requests pile into a review queue their team is already struggling with, and the organizational work of restructuring ratios and rebuilding quality gates gets postponed because the leader is busy demonstrating that the new way works.

Personal fluency is necessary and it is not the job. The job is the system. The tell in an interview is whether a candidate talks about what they built with agents, or about what changed in their team’s numbers after they did.

Adoption is wildly uneven, and that is a governance problem

The cross-functional numbers in the report are as interesting as the engineering ones. Ninety percent of their portfolio’s engineering teams are actively deploying AI. Only 24% of finance leaders are, with data quality, system fragmentation, security and compliance cited as the blockers. Forty-nine percent of companies say they are already delivering more without adding headcount, and 86% of leaders expect AI to meaningfully change how their team operates within twelve months.

From the Bessemer portfolio survey Engineering teams deploying AI 90% Expect their team to change in 12 months 86% Delivering more without adding headcount 49% Finance leaders deploying AI 24% The 90 and the 24 sit in the same companies, and the second one is the reason the first one is hard to fund.

Figure 4. Adoption by function, from the report. The spread between engineering and finance is the practical reason so many AI programmes cannot produce a credible return number.

That spread has a consequence the report does not draw out. If engineering is at 90% and finance is at 24%, then the function that would normally validate an ROI claim is the function least equipped to evaluate it. Engineering makes the claim, finance lacks the instrumentation to check it, and the number that reaches the board is unverified by anyone independent of the programme that produced it.

Bessemer’s own recommendation, that the return on tooling and token spend should be treated as a leadership metric rather than an IT line item, only works if somebody outside engineering can audit it.

The job description, rewritten

 The 2022 engineering leaderWhat the report describes
Primary constraintEngineering hours available to write codeReview, verification and specification capacity
Unit of planningThe quarter, the sprint, the estimateThe day or the week, with roadmaps compressed accordingly
Org designInherited ratios, mid-level management for coordinationRatios as a live variable, smaller autonomous pods, less coordination overhead
Measured onDelivery against planDelivery, plus a defensible return on tooling and token spend
Personally doesReviews the architecture, reads the dashboardsShips in the codebase while making the structural calls
Customer contactEscalations and the occasional QBRExpected, and increasingly formalized through forward deployed roles

The blurring the report describes is real and it runs in both directions. Product leaders are expected to understand model performance. Engineering leaders are expected to spend more time with customers. The Forward Deployed Engineer, along with the Field CTO and Field CISO variants, is the structural expression of that: technical people moved closer to the front line because the front line now asks technical questions a traditional sales engineer cannot answer.

One of their operating advisors, Barak Turovsky, makes the point that the most consequential AI decisions a CEO makes are about operating model design rather than model selection. I think that will read as obvious in three years and it is not obvious yet.

What I would actually ask a candidate

The report suggests turning interviews into working sessions, which is right. Here is my version for an engineering leadership hire.

  1. What was your review-to-generation ratio, and what did you do about it? If a candidate has led an AI-heavy team and has never looked at this, they were managing a 2022 org. If they have, the answer tells you how they think about constraints.
  2. Walk me through a rollout you stopped. Anybody can describe a successful pilot. The interesting question is whether they have ever measured something honestly enough to kill it, which is the only evidence that their successful pilots mean anything.
  3. How did you calculate the return, and what was in the denominator? Tooling, tokens, review burden, rework. A candidate who answers with time saved from a survey and nothing else has not done the work. This is the AI ROI question and it is now a board-level one.
  4. What got worse? Every real transformation has a cost column. Review latency, batch sizes, junior ramp time, incident volume. A candidate with no cost column either did not measure or will not tell you.
  5. How do you know your adoption number is real? Seat counts are not adoption. Ask how they distinguished eligible work that ran through AI from work that merely could have.
  6. What did you change about how juniors learn? If implementation was the training ground and implementation is now automated, the answer to this determines whether their org still has senior engineers in 2031.

Where I would push back

The report is written for CEOs hiring executives, so it naturally frames the answer as a person. Find the right leader and the org design follows. In my experience the causality runs the other way at least as often. Leaders inherit measurement systems, review cultures and ratio assumptions that they cannot see well enough to change, and a brilliant hire dropped into an uninstrumented organization spends their first year arguing from anecdote like everyone else. The hire matters enormously. Give them a baseline and an instrument panel on day one and they will matter twice as much.

The sentence I keep coming back to is the framing about running a 2022 team structure in a 2026 product environment. It is a good question precisely because most leaders cannot answer it with evidence. They know their tools changed. They can feel that the shape of the work changed. Very few can show you what happened to their ratios, their review capacity, their rework rate or their cost per shipped change over the same period.

Org design is a competitive variable now, as Bessemer says. Variables are things you measure. That is the part of this that is not a hiring problem.

Source: Bessemer Venture Partners Talent Team, Artisanal Talent and Atlas Editors, Talent trends for the AI-native C-suite, September 2026, based on a survey of nearly 175 functional leaders across more than 100 portfolio companies. The 180% and 30% figures are cited in that report from research published by CEPR on productivity effects across generations of AI coding tools. Quoted and paraphrased contributors include Bessemer Operating Advisors Jessica Popp and Barak Turovsky, and Shopify’s Farhan Thawar. Figures 1 to 4 are my own; Figure 2 is my interpretation rather than a model proposed by Bessemer.

Ready to unlock your SDLC productivity?

Request a Demo Call