Waydev Research · Nine years in the field

The WAY Framework

Nine years of research and every pioneer’s best ideas, DORA, SPACE, Core 4, and the AI measurement models, combined under one roof, with none of the lock-in.

Request a Demo
or
See a Live Demo

Backed by Y Combinator · Trusted by FORTUNE 500 enterprise companies

Why the WAY Framework, why now

Every vendor suddenly has a framework. We took the opposite position.

In the span of about eighteen months, nearly every vendor in engineering intelligence shipped its own measurement framework. Cortex published DRIVE. LinearB published APEX. Uplevel published WAVE. Each comes with a book or a guide, a maturity assessment, and a platform that happens to be the only one purpose-built to run it.

The proximate cause has a name: DX. Between 2024 and 2025, DX published the Core 4 framework and an AI measurement framework, both developed with researchers behind the field’s foundational work. The frameworks traveled, leaders adopted the vocabulary, and in late 2025 Atlassian bought the company for one billion dollars. The rest of the market drew a conclusion: the framework was the moat. Own the language leaders use to think about measurement, and you own the shortlist.

Waydev has spent nine years in this market, since before most of these companies existed, and we are taking the opposite position. The WAY Framework is not another acronym competing for your board deck. It is our standing commitment to combine the best of the field’s research, keep it current as the science evolves, and never lock your organization’s vocabulary to our product. But the intellectual history did not start in 2024, and understanding where it did start is the best protection against mistaking marketing for methodology. So this guide takes the frameworks in order: the research generation first, the synthesis generation second, the vendor generation third, and WAY at the end, held to the same tests as everyone else.

Part I · The research generation

Where the science actually comes from

DORA

2014–PRESENT · FORSGREN, HUMBLE, KIM · NOW AT GOOGLE

Official source: dora.dev ↗

Everything in this field descends from DORA, the DevOps Research and Assessment program. It began as an academic-grade research effort: multi-year surveys of tens of thousands of professionals, statistical modeling, and peer-reviewable methods, culminating in the book Accelerate. Its landmark finding was that software delivery performance predicts organizational performance, and that it can be measured with four metrics:

Metric What it measures Dimension
Deployment frequency How often you ship to production Throughput
Lead time for changes How long from commit to production Throughput
Change failure rate What share of changes cause failures Stability
Failed deployment recovery time How fast you restore service after a bad change Stability

The deep insight, often lost in dashboard implementations, is that speed and stability are not a trade-off. Elite performers in DORA’s data are fast and stable, because the same capabilities, small batches, automation, loosely coupled architecture, and a generative culture produce both. DORA was never just four numbers; it is a validated causal chain from capabilities to delivery performance to organizational outcomes.

Strengths: a decade of evidence, industry-standard definitions, benchmarkable across companies, vendor-neutral. Limits: the four keys describe the delivery pipeline, not the whole organization. They are lagging indicators, they say little about individual or team experience, and they were defined before AI-assisted development existed. DORA’s own recent research has moved into AI, defining the adoption capabilities that make AI investment pay off, which the vendor frameworks now borrow liberally.

SPACE

2021 · FORSGREN, STOREY ET AL. · GITHUB / MICROSOFT RESEARCH

Original paper: The SPACE of Developer Productivity, ACM Queue ↗

SPACE was the field’s course correction. By 2021, DORA’s four keys were being abused as productivity scores for individuals, exactly what the research warned against. SPACE, from many of the same researchers, is not a metric set at all. It is a rubric arguing that developer productivity is multidimensional and cannot be captured by any single number. The five dimensions:

Dimension Example signals
Satisfaction and wellbeing Developer satisfaction, burnout risk, retention intent
Performance Outcomes of work: quality, reliability, customer impact
Activity Volume counts: commits, PRs, reviews, deployments
Communication and collaboration Review quality, knowledge flow, discoverability
Efficiency and flow Cycle time, handoffs, interruptions, wait states

SPACE’s operating rules matter more than its letters: measure across at least three dimensions at once, mix system data with perceptual data, and never let one metric stand alone, because of Goodhart’s law: any measure that becomes a target stops being a good measure. Strengths: the intellectual guardrails of the entire field. Limits: it deliberately does not tell you which metrics to pick, which is why every later framework, vendor or otherwise, positions itself as “SPACE, operationalized.”

DORA Core

ONGOING · DORA / GOOGLE CLOUD

Official source: DORA research & Core model, dora.dev ↗

DORA Core is the program’s answer to a decade of annual reports piling up: a consolidated, stable model of the findings that have held up across years of data. It formalizes the chain the research always implied: capabilities (technical practices like version control and CI, process practices like small batches, cultural conditions like psychological safety) drive performance (software delivery and operational performance, the four keys plus reliability), which drives outcomes (organizational performance and team wellbeing).

Why it matters in this discussion: DORA Core is the closest thing the industry has to settled science, it is public, and it belongs to no vendor. When a vendor framework tells you it goes “beyond DORA,” DORA Core is the baseline it should be measured against, and the vocabulary your board reporting stays legible in if you ever change platforms.

Part II · The synthesis generation

DX: the frameworks that started the arms race

DX Core 4

2024 · DX, WITH RESEARCHERS BEHIND DORA AND SPACE

Official source: DX Core 4, getdx.com ↗

Core 4’s contribution was ruthless simplification with a research pedigree. It unified DORA, SPACE, and developer experience research into four dimensions, each with one headline metric an executive can hold in their head:

Dimension Headline metric Instrument
Speed Throughput of shipped work (e.g. PRs merged per engineer) System data
Effectiveness Developer Experience Index (DXI) Survey
Quality Change failure rate System data
Impact Share of engineering time on new capabilities Mixed

Two things made it land. First, credibility: it was framed as a unification of DORA and SPACE, with names from that lineage attached, not a replacement for them. Second, balance by construction: speed is always counterweighted by quality and experience, so gaming one dimension shows up in another. Limits: the DXI, the Effectiveness metric, is proprietary to DX, which is the quiet lock-in inside an otherwise portable framework, a point that matters more now that DX belongs to Atlassian. And a survey-based centerpiece inherits survey mechanics: quarterly cadence, response decay, self-report bias.

DX AI Measurement Framework

2025 · DX RESEARCH

Official source: Measuring AI code assistants and agents, getdx.com ↗

The companion framework answered the question boards started asking: is the AI spend working? Its structure is a three-lens funnel, and it has become the de facto template everyone else riffs on:

Utilization. Who is using AI tools, how often, and for what: adoption breadth and depth, before anything else, because impact analysis on a team that barely uses the tools measures nothing.

Impact. What AI usage does to the metrics that already matter: the Core 4 dimensions, plus self-reported time savings. The design principle is the right one: AI should be evaluated against your existing performance model, not a parallel scoreboard of AI-specific vanity metrics.

Cost. Spend against value: licenses, tokens, and the ROI arithmetic the CFO will run whether or not engineering runs it first.

Utilization, impact, cost is a genuinely sound skeleton, and readers should notice that every vendor framework in Part III contains some rotation of it. Limits: the impact lens leans on self-reported time savings, which are systematically optimistic, and the framework’s home is now inside Atlassian, which sells an AI coding agent of its own. The methodology remains useful. The neutrality of its steward is a fair question.

Part III · The vendor generation

DRIVE, APEX, WAVE: three answers, three products

DRIVE

2026 · CORTEX · AUTHORED BY ITS CTO

Official source: DRIVE, cortex.io ↗

DRIVE measures what Cortex calls organizational health in the age of AI, explicitly positioned at a higher altitude than developer productivity: not “are developers effective” but “can the organization sustainably turn customer needs into reliable software.” Five pillars, each phrased as a leadership question:

Pillar The question Sample metrics
Delivery Are we shipping fast, sustainably? Deploy frequency, lead time, on-call pager volume
Reliability Are we keeping promises to customers? SLO pass/fail, Sev0/Sev1 incident counts
Initiatives Are org-wide investments progressing? Initiative milestone completion, action item completion
Vigilance Are we managing security risk? Open critical CVEs, assets below the compliance bar, orphaned assets
Efficiency Are resources on the right problems? Cloud spend vs budget, LLM token costs, innovation capacity share

Attached to the pillars is the OpEx review, a recurring leadership ritual borrowed from manufacturing operational excellence: treat the org as one observable system, review it against DRIVE, reallocate people and money to close gaps. That ritual is arguably the framework’s best idea, and it costs nothing to adopt with or without Cortex.

What it gets right: it is honest about its altitude. Vigilance and Initiatives are real leadership concerns that DORA and SPACE never claimed to cover, and the AI-era framing (velocity outpacing controls) is a fair diagnosis. What to watch: DRIVE is a governance framework, not a productivity or AI ROI framework: it can tell you the factory is up to code, not what the factory produced per dollar of AI spend. Its pillars map one-to-one onto Cortex’s catalog, scorecards, and workflows, several of its metrics (orphaned assets, compliance-bar coverage) are only measurable if you maintain a service catalog, and its vocabulary is one vendor’s. Adopting DRIVE as your organization’s definition of engineering health is, in practice, adopting Cortex’s data model.

APEX

2026 · LINEARB

Official source: The APEX framework, linearb.io ↗

APEX is the most focused of the three: an operating model for validating whether AI increases throughput without breaking delivery confidence or developer experience. Its design principle is “north stars over metric volume”: four pillars, one guiding metric each.

Pillar North star Diagnostics
AI leverage AI-assisted pull requests Human-AI contribution ratio by phase
Predictability Planning and capacity accuracy CFR, rework, defects as leading indicators
Efficiency (flow) Cycle time, plus CFR Pickup time, review time, PR size
X (DevEx) Developer satisfaction (survey) DORA’s seven AI readiness capabilities

Its sharpest ideas are real contributions to the conversation. Measure AI in the critical path: if AI is not visible at the PR level, you cannot connect it to delivery, and PR-level attribution beats seat counts by a wide margin. Bottlenecks shift downstream: if AI cuts coding time but review queues balloon, the system did not improve, the constraint moved, so decompose cycle time. Predictability before velocity: speed that breaks commitments is chaos, not progress. And it prescribes an operating cadence, weekly for AI leverage, per sprint for predictability, monthly for flow, quarterly for DevEx, which is more practical guidance than most frameworks bother with.

What it gets right: PR-level AI attribution and constraint thinking. It also borrows honestly, crediting DORA’s AI capabilities by name. What to watch: APEX is scoped to the delivery pipeline plus sentiment. There is no business alignment, allocation, or cost dimension, so it answers “is AI making us faster sustainably” but not “what is AI returning per dollar” or “is the work pointed at the right things.” Its predictability pillar presumes sprint-style planning data, which is really Jira hygiene by another name. And LinearB states plainly that it is “the only platform purpose-built to implement APEX at scale,” which tells you what the guide is: a well-made funnel.

WAVE

2025–2026 · UPLEVEL

Official source: The WAVE Framework, uplevelteam.com ↗

WAVE frames engineering organizations as sociotechnical systems: human collaboration and environmental factors matter as much as deployment statistics, so the framework spans both. It explicitly critiques DORA as narrow and lagging, and SPACE as theoretical without a measurement approach. Four interconnected dimensions, each summarized by a lagging outcome metric driven by leading-indicator inputs:

Dimension What it covers Sample measures
Ways of Working Cultural and behavioral enablers Deep work hours, team health, AI maturity
Alignment Effort connected to business value Allocation of effort, planning effectiveness, user feedback cycles
Velocity Flow of work through the system Composite velocity score, handoffs, PR review health
Environment Efficiency System quality and friction Recovery (a DORA subset), code quality, flow efficiency

WAVE’s genuinely valuable observations: AI inflates activity metrics, so merge and issue velocity can soar while deployment frequency and delivered value stay flat, exactly the trap executives fall into when they read raw output as ROI. Allocation reality: teams believe they spend far more time on new value than objective analysis shows. And knowledge work spends most of its life waiting, so flow efficiency, not activity, is where the leverage hides.

What it gets right: the sociotechnical framing is correct, and the leading-versus-lagging structure is more thoughtful than most. What to watch: WAVE’s signature inputs, deep work hours, meeting patterns, team health, come from calendars, collaboration tools, and surveys, which means the framework operationalizes workday measurement that works councils and privacy reviews increasingly refuse, and several of its dimensions are only measurable with Uplevel’s specific instrumentation and human interpretation layer. It is also the least benchmarkable of the three: composite proprietary scores cannot be compared across the industry the way DORA metrics can.

The pattern

Every vendor framework is a mirror of its vendor

Line them up and the pattern is impossible to miss. Cortex sells a service catalog and governance platform, and DRIVE is a governance framework whose metrics need a catalog. LinearB sells PR workflow instrumentation and automation, and APEX is a PR-centric flow framework. Uplevel sells calendar-and-survey-based insight with a coaching layer, and WAVE is a sociotechnical framework requiring exactly that instrumentation. DX sold a survey platform, and Core 4’s centerpiece is a proprietary survey index.

None of this makes the frameworks dishonest. Each contains ideas worth stealing: DRIVE’s OpEx review ritual, APEX’s PR-level AI attribution and constraint thinking, WAVE’s warning about AI-inflated activity metrics, Core 4’s one-metric-per-dimension discipline. Steal all of it. It is free.

What is not free is adopting a vendor’s framework as your organization’s official definition of engineering performance. Your targets, your board narrative, and your managers’ vocabulary get expressed in terms only one platform natively speaks. Data lock-in is a project. Vocabulary lock-in is a culture change.

“A framework you cannot take with you when you leave the vendor was never a framework. It was an onboarding flow.”

The map

All eight frameworks, one table

Framework Steward Primary question Instrument Portability
DORA (four keys) Research program (Google) How fast and stable is delivery? System data Fully portable, benchmarkable
SPACE Academic authors What dimensions must any measurement cover? Mixed (rubric) Fully portable
DORA Core DORA / Google Cloud Which capabilities drive which outcomes? System + survey Fully portable
DX Core 4 DX (now Atlassian) One balanced scorecard for productivity System + survey (DXI) Mostly portable; DXI is proprietary
DX AI Framework DX (now Atlassian) Is the AI spend working? System + self-report Structure portable; steward sells AI tools
DRIVE Cortex Is the org operationally healthy? Catalog + system data Several metrics require a service catalog
APEX LinearB Is AI improving throughput sustainably? PR data + survey Concepts portable; self-described as LinearB-built
WAVE Uplevel Is the sociotechnical system healthy? Calendar + survey + system Composite scores; needs Uplevel’s instrumentation
WAY Waydev Which of the above answers your current question? Work-level system data, all lenses Fully portable by definition; owns no metrics

The rubric

Five questions to ask any framework. We answer them for WAY below.

1. Whose evidence is it built on? A decade of published research, or one company’s customer anecdotes distilled into an acronym? Both can be useful. Only one is science.

2. Can you take it with you? If every metric definition survives a platform change, it is a framework. If the centerpiece is a proprietary index or requires one vendor’s instrumentation, it is a product feature wearing a framework’s clothes.

3. What instrument does it live on? System data is continuous and objective but blind to feelings. Surveys capture experience but arrive quarterly, decay, and self-report optimistically. Calendars and messages see the workday but trigger privacy reviews. Good measurement mixes instruments deliberately, and knows which claims each can and cannot support.

4. Does it answer your current question? Delivery health is DORA’s question. Balanced productivity is Core 4’s. Governance is DRIVE’s. AI throughput is APEX’s. Sociotechnical health is WAVE’s. AI ROI per dollar is, notably, fully answered by none of them, which is why boards keep asking it.

5. Is it Goodhart-resistant? Every dimension needs a counterweight: speed against quality, throughput against experience, adoption against outcomes. A framework without built-in tension is an invitation to game it.

Introducing the WAY Framework

Yes, we see the irony. Here is why WAY is different by construction.

We just spent an entire guide warning you about vendor frameworks, and now we are naming one. So let us be precise. WAY defines no proprietary metrics, no composite index, and no acronym-shaped pillars competing with the research. It is a meta-framework: a set of commitments about how measurement should be built, distilled from nine years of doing this, since we shipped one of the first Git analytics platforms in 2017 and patented the approach with the USPTO. WAY is the only framework on this page you could, in principle, implement without us. We simply intend to be the best place to run it.

W

Work-first

Every framework here is a projection of the same underlying reality: commits, pull requests, reviews, deployments, and the tickets and spend around them. WAY’s first commitment is to measure at the source, continuously, at the commit and PR level, rather than through proxies: not calendars, not self-reported time savings, not catalog metadata. Get the source right and every projection becomes available.

A

Agnostic

WAY combines the best of the field rather than replacing it. Waydev implements DORA and DORA Core with benchmarks, SPACE-consistent views, Core 4-style balanced scorecards, and an AI measurement layer on the utilization-impact-cost logic, grounded in how AI-assisted work behaves in the delivery path. And we steal the good ideas openly: DRIVE’s OpEx review, APEX’s constraint analysis, WAVE’s activity-inflation warning. The ideas were never the lock-in. The instrumentation was.

Y

Yours

Your metric definitions stay industry-standard and portable. Your board narrative stays legible to any executive, auditor, or future platform. Your history stays intact, rebuilt directly from your Git provider, so even changing vendors costs no baselines. And your framework choice stays reversible: change the lens, not the platform, and not the culture. Your organization’s language for engineering performance belongs to your organization.

The warranty

When the science evolves, your platform evolves. You re-platform nothing.

The research will keep moving. DORA reinvented itself around AI capabilities. SPACE corrected the field’s excesses. Core 4 compressed a decade into four numbers. Something will follow, and when it does, the vendor-framework companies face a conflict: absorbing the new science means admitting their acronym was not the answer. WAY has no such conflict, because WAY owns no answer.

Our standing guarantee, backed by nine years of doing exactly this, is that whatever the field’s best validated thinking is, Waydev will implement it, benchmark it, and hand it to you in your own vocabulary. Not a fifth acronym in the war, but the position that ends it: measure the work, combine the best, and keep it yours.

Ready to measure the WAY?

See your organization through every lens, on your own data

Connect your repos and see DORA, SPACE, Core 4, and a full AI ROI view on the same data, in days. Pick the language your board understands best, knowing you can change it any time. That is the warranty.

Request a Demo
or
See it in action