Tobi Lütke was asked whether he should go start an AI lab. His answer explains why most engineering organizations are measuring the wrong half of their software lifecycle.
A schematic, not a benchmark. Green marks the stage AI compresses fastest. Red marks the stages that absorb the overflow. The other three barely move, and that is the whole problem.
Somewhere in the middle of a conversation with David Senra, Shopify founder Tobi Lütke gets asked the question that has been put to every serious technologist since late 2022: given everything happening in AI, shouldn’t you go start a lab, or go work at OpenAI?
His answer is no. Not out of stubbornness, and not because he thinks the technology is overrated. His answer is no because the question misunderstands what he has been doing for the last twenty years.
Lütke’s framing of his own job is unusually plain. Get very good at something. Build it into products. Share it with as many people as possible. His version of that is helping millions of entrepreneurs and small businesses use technology and retail to build their own work. AI does not replace that objective. It gives him a faster, stranger, more powerful way to pursue it.
He is pursuing the same problem in a changed landscape, reinventing what he already built. The landscape moved. The problem did not. On Tobi Lütke, in conversation with David Senra
That distinction sounds small. It is not. And for anyone running an engineering organization right now, it has a very specific operational consequence, which is the part I want to spend this piece on.
I have sat in a lot of rooms over the past eighteen months where a version of this conversation happens. A VP of engineering, a CTO, sometimes a CEO, all circling the same anxiety: are we the incumbent in this story? Should we be building something else entirely? Is the thing we spent nine years getting good at now a liability?
The anxiety is real. Categories are collapsing. Products that took three years to build are being approximated in a weekend. Pricing models built on seat count are being undermined by the fact that the seats now do several times the work.
But the response is almost always wrong in the same specific way. Teams reach for a new mission because the new technology feels like it demands one. They confuse the tool with the purpose.
The mission test
Write down what your company is for, in one sentence, without naming a single piece of technology. No “AI-powered.” No “platform.” Just the outcome you exist to produce for a specific person.
If you cannot write that sentence, you do not have a mission. You have a product category, and product categories get eaten.
If you can write it, read it again and ask whether AI changes the sentence, or changes how you deliver on it. For almost every company I know, it is the second one.
Here is where the Lütke framing becomes concrete rather than philosophical.
The AI SDLC is not a new lifecycle. It is the same lifecycle with radically uneven acceleration applied to it. Planning, review, testing, release and production ownership are all still there, still governed by the same constraints they had in 2019. One station in the line got dramatically faster. The rest did not.
This is the oldest result in operations research, arriving in a new outfit. Speed up one station in a line and you do not speed up the line. You move the queue. Anyone who has run a factory, a hospital, or a build pipeline knows exactly what happens next, and it is not throughput. It is inventory piling up in front of the next station.
In an engineering organization the queue shows up in a predictable order.
Plan
Largely unchanged. Deciding what is worth building is still a judgment problem with humans and customers on both ends of it. Ambiguity here now costs more, because a vague ticket turns into a large volume of confidently written code faster than it used to.
Build
Compressed hard. This is the stage everyone measures, because it is the stage where the tool licence lives and where the vendor dashboard points. It is also the stage that was rarely the actual constraint.
Review
The first place the queue forms. Reviewer capacity did not increase. Diff volume did. Reviewers who used to read code written by a colleague whose reasoning they could reconstruct are now reading code with no author intent behind it, which is slower per line, not faster.
Test
The second place it forms. Generated tests are plentiful and often shallow. Coverage numbers can rise while real defect escape gets worse, which is a genuinely dangerous combination because the metric moves in the reassuring direction.
Release
Unchanged in mechanics, worse in risk profile. Bigger, more frequent changes flowing through the same deployment safety net produces a higher change failure rate unless something else was rebuilt to absorb it.
Operate
Where the bill arrives. Rework, incident load, and the slow accumulation of code nobody in the building understands well enough to modify with confidence.
None of this is an argument against AI in the lifecycle. It is an argument that the gain is real and is currently being spent in the wrong place, because the instrumentation only watches stage two.
Adoption metrics are station-level. Mission metrics are system-level. Almost every AI reporting pack I see is entirely the first kind.
Counted today
Worth counting instead
That last one is where most of the value quietly leaks out. Teams get a real productivity gain and reinvest all of it into internal work no customer will ever notice. The technology delivered. The mission did not move an inch.
There is a prerequisite underneath all of this that most organizations have not solved: provenance.
If you cannot tell which changes were AI-originated, you cannot compare their review cost, their survival rate, or their incident profile against anything. You are left with vendor telemetry describing tool usage, and a set of delivery metrics describing outcomes, with no reliable join between them. Every conclusion drawn across that gap is a guess wearing a chart.
Getting provenance right is unglamorous work. It is also the difference between an AI programme you can steer and one you can only defend.
There is a personal layer to what Lütke is describing, and I think it is the part that actually matters.
Building anything worthwhile takes long enough that the world will change underneath you at least twice. If your identity as a founder is attached to a technology, each of those shifts is an existential crisis. If it is attached to a problem and to the people who have it, each shift is just a new set of tools arriving in the shop.
Nine years into Waydev, I have watched the second thing happen repeatedly, and the reframe never stops being useful. The mission was never the stack. The mission was always the person on the other end who needs something they cannot get anywhere else.
New technology gives you a new way to solve it. Your job is to make sure the new way actually reaches them, and not just the part of the pipeline that was easiest to accelerate.
Source: Tobi Lütke in conversation with David Senra. Clip via Agile Academy.
Ready to unlock your SDLC productivity?