Back To All

SPACE in the SI SDLC era

October 11th, 2026
Topics
AI
AI ADOPTION
AI ROI
AI SDLC
SI
SI SDLC
SI SDLC Impact
SI SDLC ROI
SPACE
Share Article

SPACE in the SI SDLC era

Five years after SPACE changed how we talk about developer productivity, its authors met in person for the first time. Here is what they said, and how the framework needs to change now that agents write, review and ship a growing share of the code.

  • SSatisfaction and wellbeing
  • PPerformance
  • AActivity
  • CCommunication and collaboration
  • EEfficiency and flow
  • TTrust, the layer we add for agents

In short

  • SPACE still holds. Its authors agree the five dimensions have aged well. What needs work is how each one is measured once agents join the team.
  • Activity metrics are back, and dangerous again. Agent swarms can generate PRs at a volume that makes raw counts meaningless on their own.
  • Collaboration is the blind spot. It is the one dimension most developers say AI has not improved.
  • Trust is the missing layer. In the SI SDLC, trust covers people and agents: did the work stay in scope, and is the record of it accurate?

Where this comes from

This article builds on Brian Houck’s recap of a panel at the first Developer Experience Research Forum at UC Irvine, where all six SPACE authors met in person for the first time: Nicole Forsgren, Margaret-Anne Storey, Chandra Maddila, Thomas Zimmermann, Jenna Butler and Houck himself. DX blog

SPACE was published in ACM Queue in February 2021. Its core claim is that developer productivity cannot be captured by a single number: measure across at least three dimensions, include at least one perceptual measure such as a survey, and expect good metrics to pull against each other. ACM Queue One detail from the panel: the working name was FACTS, and trust was in it from the start.

New to the framework? Our guide to the SPACE framework and its metrics walks through each dimension in operational terms, and The SPACE Framework Playbook is the downloadable version for your team.

We read the panel through the lens of the SI SDLC, which we covered in our previous piece: the software lifecycle once agents work for hours across tools and systems, and the hard problems become scope, honesty and proof rather than code generation.

What the authors said, five years on

The dimensions held up, but they are huge

Storey’s view was that the five dimensions survived well, but each is so broad that the next step is to go deeper into them one at a time rather than add new ones.

AActivity is newly important, and newly risky

Butler pointed to headlines about how much code is now AI-generated: lines of code and PR counts are resurfacing as metrics, a decade after researchers warned against using them in isolation. Maddila added that one developer running a swarm of agents can produce a huge number of PRs, and the count says little about quality or impact. We made the same case in what AI changed about the Activity dimension: activity moves first and most visibly, while the cost lands later in review queues and rework.

Pick metrics where gaming still helps

Houck’s example is Time-To-First-PR for new hires. Even a deliberately trivial first PR leads to better long-term outcomes, because the value is in learning the environment. Good metric design means choosing measures where gaming them still produces the outcome you want.

Measurement is political, so design for it

Zimmermann noted that five dimensions make it harder to play politics with a single number. Butler stressed protecting individual data from managers, since fearful people are not productive people. Some organizations bucket metrics together so nobody can push one number up without answering for the rest.

CCollaboration is the most underinvested dimension

Developers now ask AI tools questions they used to ask colleagues. In Houck’s SPACE of AI research, communication and collaboration was the only dimension most developers did not believe AI had improved. Microsoft Research

“Make C one of the first things you look at.”

Margaret-Anne Storey, SPACE co-author, via the DX recap

Asked what they would add today, the panel named trust, cognitive and intent debt (losing understanding of codebases as AI writes more of them), deskilling, and addiction-like use of AI tools. Their conclusion: keep the structure, rebalance the rubric. DX blog

SPACE, rebalanced for the SI SDLC

The panel spoke mostly about AI assistants. Agents raise the stakes: they act for hours, touch many systems and summarize their own work, and recent safety testing showed those summaries can be wrong. Cybersecurity News Here is how we would rebalance each dimension.

Satisfaction and wellbeing

Before agentsAre developers happy with their tools and workload?

NowDo developers still understand the code they own, trust the agents they work with, and keep their skills sharp?

Measure: survey questions on confidence in owned code, trust in agent output, and time spent building skills versus supervising agents.

Performance

Before agentsDid the code do what it was supposed to do?

NowDid agent-produced work hold up in production, and was it worth what it cost?

Measure: change failure rate, rework and incidents, split by human-authored and agent-assisted changes, plus cost per shipped change. More on this in Every AI gain has a cost attached.

Activity

Before agentsCommits, PRs, reviews: easy to count, easy to misuse.

NowAgent activity can grow without limit. Count it only alongside quality, and treat out-of-scope actions as a signal of their own.

Measure: PR volume bucketed with review depth and rework, so nobody can push volume without answering for quality.

Communication and collaboration

Before agentsHow well do people share knowledge and review each other’s work?

NowAgent output lands in human review queues. Strain shows up as review overload, a few people reviewing everything, and fewer conversations between colleagues.

Measure: review load and how evenly it is spread, time PRs wait for a human, and survey signals on knowledge sharing. Waydev’s review collaboration report shows whether reviews are spread across the team or concentrated in a few people.

Efficiency and flow

Before agentsCan developers work without interruptions and handoff delays?

NowAgents add a new kind of interruption: approvals, check-ins and clean-up. The bottleneck moves to review, CI and deploy.

Measure: wait time in review and CI, interruptions from agent approvals, and focus time across the team.

Trust, the layer that runs through all five

The panel did not propose a sixth dimension, and neither do we. But in the SI SDLC, trust covers two questions: do people trust each other and their tools, and can the organization verify what agents did? We explored this in Agents can build software without limit. Your ability to trust it has a limit.

Measure: scope adherence of agent work, how closely agent summaries match the systems of record, and developer trust from surveys.

Five rules for measuring with agents on the team

  • Separate human and agent signals. Mixing them hides both the gains and the risks.
  • Never rank individuals. SPACE was built for teams and systems. Protecting individual data keeps measurement honest.
  • Bucket activity with quality. Volume only counts when paired with review depth, rework and incidents.
  • Keep a perceptual measure. Telemetry can’t tell you whether people still understand their codebase. Ask them.
  • Trust records over reports. Measure from Git, CI, tickets and deploys, not from what an agent says it did. See The artifact chain is the measurement chain.

What this means for Waydev

Waydev measures AI adoption, impact and ROI across engineering organizations. SPACE is a good test of whether that measurement is balanced, and the SI SDLC raises the bar further.

The systems-of-record side of SPACE

Waydev’s analytics are built on Git, pull requests, CI and issue trackers. That covers Performance, Activity, Efficiency and much of Collaboration through review data, from records agents cannot rewrite. Under the WAY Framework, Waydev implements DORA, SPACE-consistent views and AI measurement on the same commit-level data.

Activity that comes with context

We are focused on showing AI and agent activity next to quality and outcomes, never as a standalone count to optimize. Our How to Measure AI Agents Playbook covers the metrics in detail.

Collaboration under agent load

Review queues are where agent output meets human attention. Making review load and its distribution visible is one of the most useful things an engineering leader can do right now.

Teams, not leaderboards

We agree with the SPACE authors: productivity data should help teams improve, not rank people.

Book a demoSee SPACE-aligned AI and agent measurement in Waydev.

Ready to unlock your SDLC productivity?

Request a Demo Call