Five years after SPACE changed how we talk about developer productivity, its authors met in person for the first time. Here is what they said, and how the framework needs to change now that agents write, review and ship a growing share of the code.
This article builds on Brian Houck’s recap of a panel at the first Developer Experience Research Forum at UC Irvine, where all six SPACE authors met in person for the first time: Nicole Forsgren, Margaret-Anne Storey, Chandra Maddila, Thomas Zimmermann, Jenna Butler and Houck himself. DX blog
SPACE was published in ACM Queue in February 2021. Its core claim is that developer productivity cannot be captured by a single number: measure across at least three dimensions, include at least one perceptual measure such as a survey, and expect good metrics to pull against each other. ACM Queue One detail from the panel: the working name was FACTS, and trust was in it from the start.
New to the framework? Our guide to the SPACE framework and its metrics walks through each dimension in operational terms, and The SPACE Framework Playbook is the downloadable version for your team.
We read the panel through the lens of the SI SDLC, which we covered in our previous piece: the software lifecycle once agents work for hours across tools and systems, and the hard problems become scope, honesty and proof rather than code generation.
Storey’s view was that the five dimensions survived well, but each is so broad that the next step is to go deeper into them one at a time rather than add new ones.
Butler pointed to headlines about how much code is now AI-generated: lines of code and PR counts are resurfacing as metrics, a decade after researchers warned against using them in isolation. Maddila added that one developer running a swarm of agents can produce a huge number of PRs, and the count says little about quality or impact. We made the same case in what AI changed about the Activity dimension: activity moves first and most visibly, while the cost lands later in review queues and rework.
Houck’s example is Time-To-First-PR for new hires. Even a deliberately trivial first PR leads to better long-term outcomes, because the value is in learning the environment. Good metric design means choosing measures where gaming them still produces the outcome you want.
Zimmermann noted that five dimensions make it harder to play politics with a single number. Butler stressed protecting individual data from managers, since fearful people are not productive people. Some organizations bucket metrics together so nobody can push one number up without answering for the rest.
Developers now ask AI tools questions they used to ask colleagues. In Houck’s SPACE of AI research, communication and collaboration was the only dimension most developers did not believe AI had improved. Microsoft Research
“Make C one of the first things you look at.”
Margaret-Anne Storey, SPACE co-author, via the DX recap
Asked what they would add today, the panel named trust, cognitive and intent debt (losing understanding of codebases as AI writes more of them), deskilling, and addiction-like use of AI tools. Their conclusion: keep the structure, rebalance the rubric. DX blog
The panel spoke mostly about AI assistants. Agents raise the stakes: they act for hours, touch many systems and summarize their own work, and recent safety testing showed those summaries can be wrong. Cybersecurity News Here is how we would rebalance each dimension.
Before agentsAre developers happy with their tools and workload?
NowDo developers still understand the code they own, trust the agents they work with, and keep their skills sharp?
Before agentsDid the code do what it was supposed to do?
NowDid agent-produced work hold up in production, and was it worth what it cost?
Before agentsCommits, PRs, reviews: easy to count, easy to misuse.
NowAgent activity can grow without limit. Count it only alongside quality, and treat out-of-scope actions as a signal of their own.
Before agentsHow well do people share knowledge and review each other’s work?
NowAgent output lands in human review queues. Strain shows up as review overload, a few people reviewing everything, and fewer conversations between colleagues.
Before agentsCan developers work without interruptions and handoff delays?
NowAgents add a new kind of interruption: approvals, check-ins and clean-up. The bottleneck moves to review, CI and deploy.
The panel did not propose a sixth dimension, and neither do we. But in the SI SDLC, trust covers two questions: do people trust each other and their tools, and can the organization verify what agents did? We explored this in Agents can build software without limit. Your ability to trust it has a limit.
Waydev measures AI adoption, impact and ROI across engineering organizations. SPACE is a good test of whether that measurement is balanced, and the SI SDLC raises the bar further.
Waydev’s analytics are built on Git, pull requests, CI and issue trackers. That covers Performance, Activity, Efficiency and much of Collaboration through review data, from records agents cannot rewrite. Under the WAY Framework, Waydev implements DORA, SPACE-consistent views and AI measurement on the same commit-level data.
We are focused on showing AI and agent activity next to quality and outcomes, never as a standalone count to optimize. Our How to Measure AI Agents Playbook covers the metrics in detail.
Review queues are where agent output meets human attention. Making review load and its distribution visible is one of the most useful things an engineering leader can do right now.
We agree with the SPACE authors: productivity data should help teams improve, not rank people.
Book a demoSee SPACE-aligned AI and agent measurement in Waydev.
Ready to unlock your SDLC productivity?