A new framework from DX and the researchers behind SPACE, DevEx, and DORA
Five properties of context your agents can actually use
The first four ask whether context lets an agent do good work. The fifth asks whether the agent should have it at all.
Ask an engineering leader how to get more out of AI coding agents and the conversation almost always turns to models: newer ones, bigger context windows, better reasoning. A new paper in ACM Queue argues that this is the wrong place to look first. The same model can perform very differently depending on what it is given to work with, and most organizations have no way to evaluate that input.
Where CAFE(S) sits in the AI system
Every agent output is shaped by four layers. Models and harnesses get most of the engineering attention. The layer in between gets almost none.
CAFE(S) evaluates one layer: the translation from intent into context. It doesn’t grade the model or the harness. It asks whether what you handed the agent was fit for the job.
The paper is CAFE(S): Your Agent Is Only As Good As Its Context, by Brian Houck (distinguished scientist at DX, now part of Atlassian), Max Kanat-Alexander (Capital One), Eirini Kalliamvakou (GitHub), Margaret-Anne Storey (University of Victoria), and Nicole Forsgren (Google). If those names sound familiar, it’s because several of them are behind SPACE, DevEx, and DORA, the frameworks most engineering organizations already use to measure delivery. DX, the developer intelligence company behind the DX Core 4, is closely tied to this work: Houck leads research there, and Storey collaborates with the DX team.
Their new contribution is a vocabulary for something every team running agents has felt but few can name: the quality of the context those agents work from. Below we break down the framework, the failure patterns it identifies, and what it means for engineering leaders who need to prove their AI investment is working.
The core claim: context quality is now a first-order driver of what agentic software costs, how safely it behaves, and whether developers can rely on it.
None of this is a new problem. Software teams have struggled for decades to document decisions, keep knowledge current, and help people find what they need. The authors’ point is that AI doesn’t create the need for good knowledge management. It dramatically raises the cost of bad knowledge management, and makes that cost visible immediately and at scale.
They identify three forces that have turned context from an abstract concern into an engineering problem.
This is the least visible cost because it shows up as friction, not a line item. When context is unclear, the agent guesses. When it’s incomplete, the agent produces something plausible but wrong. When it’s stale, the agent confidently applies a policy that no longer exists. Every one of those failures lands back on a developer, who has to notice, diagnose, re-explain, and retry. The agent was supposed to remove toil. Poor context hands it right back.
Context is where token spend accumulates, and the paper cites two striking examples of what engineering it well is worth:
Organizations that repeatedly send agents large, redundant, poorly structured context are simply paying more for the same work, and that gap widens as agents take on longer-running tasks.
Organizations remain accountable for what their agents say and do. The paper points to the 2024 tribunal ruling that held Air Canada liable for a refund its chatbot promised but its policy didn’t provide. It also cites EchoLeak and CamoLeak, two prompt-injection attacks against Microsoft 365 Copilot and GitHub Copilot Chat in which untrusted content became part of the agent’s working context and the agent faithfully acted on it. Neither system was hacked in the traditional sense. The context was compromised, so everything built on it was too.
The authors define good context by what it enables: a human and agent working together can accomplish the task correctly, efficiently, and safely. A useful mental test they offer is to ask what a competent colleague would need to do the same task well if they couldn’t lean over and ask a follow-up question.
From there, they name five properties. The first four describe whether context lets an agent do good work. The fifth, set in parentheses, asks whether the agent should have that context at all.
Can the agent interpret the request the way it was intended?
Breaks when a term the task depends on has more than one reasonable reading, and the agent silently picks one.
Can the agent proceed, and does it know when it’s done?
Breaks when the goal, constraints, or definition of done were never written down because they felt obvious to the author.
Can the agent trust that the context is true?
Breaks when docs, APIs, or policies changed and the context didn’t, so the agent reasons correctly from wrong facts.
Can the agent focus on what matters?
Breaks when the whole codebase gets sent for a question about one module, and the signal drowns in noise.
Should the agent have this context at all?
Breaks when attacker-controlled content, or secrets and personal data the task doesn’t need, end up in the agent’s working context.
Adapted from Table 1 in CAFE(S), ACM Queue (2026).
The dimensions are designed to be independent, which is what makes them useful in practice. Context can be perfectly clear and still describe the system as it was last year. It can be accurate and secure but bury the one detail that matters under thousands of irrelevant tokens. It can be lean and accurate and still never say what finished looks like. Each pillar isolates one thing you can review for and fix on its own.
Two distinctions in the paper are worth calling out. Clarity and actionability are not the same thing: clarity asks whether the context can only be read one way, while actionability asks whether enough was said at all. And efficiency is not brevity for its own sake. The goal is the highest signal-to-noise ratio, not the smallest possible context. A slightly verbose context that’s complete beats a tight one that’s missing a constraint.
The authors also explain why they stopped at five. Timeliness folds into fidelity, since stale context is just context that’s no longer true. Cost is a downstream result of efficiency. Provenance is a way of protecting fidelity. Accessibility, whether information can be found at all, is deliberately treated as a prerequisite rather than a quality dimension.
The question isn’t whether your agent has context. It’s whether the context it has is actually fit for the task.
The most practical part of the paper is a catalog of recurring failure patterns. Just as engineers use “code smell” to name recurring design problems, the authors propose context smells. The key insight is that in many of these cases, the necessary information was already present. It was stale, ambiguous, contradictory, buried, or mixed with content that should never have been trusted.
| Pattern | What happens | Dimension |
|---|---|---|
| Specification ambiguity | A semi-ambiguous request leads the agent to pick an interpretation and build on it, compounding the error at every step | Clarity |
| Referential ambiguity | Several entities match the same reference, so the agent correctly follows the request on the wrong object | Clarity |
| Weekend runaway | No stopping criteria, so the agent retries and expands scope for hours, running up a large bill | EfficiencyActionability |
| Goal drift | The agent gradually optimizes for a proxy of the task instead of what the user actually wanted | Actionability |
| Negative-constraint fragility | The context lists what not to do without saying what to do instead | Actionability |
| Lost in the details | Implementation detail without rationale, so the agent optimizes local mechanics and loses the point | ActionabilityEfficiency |
| Stale guidance | The context was accurate once, but the system moved on | Fidelity |
| Confident hallucination | Unsupported claims get presented as fact because nothing in the workflow treats them as needing verification | Fidelity |
| Schema drift | Tool inputs or outputs change shape over time, breaking downstream steps even though each component still works | Fidelity |
| Contradictory context | Two sources disagree, and the agent has to arbitrate between incompatible versions of reality | ClarityFidelity |
| Lost in the middle | The right information is present but overlooked because of volume or placement | Efficiency |
| Overspecification | Prescribing the how in more detail than needed, displacing what the model would have done better unguided | EfficiencyActionability |
| Indirect prompt injection | Retrieved content carries instructions that compete with or override the real task | Security |
| Paste-in-the-prompt leak | Users paste confidential data into external AI tools | Security |
Where the 18 failure patterns land
Number of patterns in the paper that involve each dimension. Some patterns involve two.
Actionability is the most common failure. Missing goals, missing constraints, and no definition of done cause more of these patterns than any other dimension. It’s also the cheapest to fix: write down what finished looks like.
Read that list as an engineering leader and a pattern jumps out. Almost every one of these smells shows up in delivery data before anyone names it. A weekend runaway is a token spike. Specification ambiguity is a PR that gets reworked after review. Stale guidance is an agent that keeps proposing a deprecated approach. The failures are context problems, but their fingerprints land in your engineering metrics.
One principle in the paper is easy to overlook and very practical: the same words can be good context in one place and bad context in another.
The authors use the example of an AGENTS.md file at a repository’s root, which gets loaded into every agent session. An instruction about running the test suite belongs there because nearly every session needs it. An instruction about one rarely touched file does not. Placed there, it gets injected into thousands of sessions that have nothing to do with it, costing tokens and attention every time. Move the same sentence into a comment inside that file and it becomes high-quality context, available exactly when it’s relevant. Nothing about the words changed. Only their location did.
## Testing Run `make test` before opening a PR. Don't change the rounding in billing/legacy_tax.py. Finance reconciles against it.
# Don't change the rounding below.
# Finance reconciles against it.
def round_tax(amount):
...This has a direct implication for review. Because shared context is read by every agent session that loads it, its defects are amplified. A stale architecture doc isn’t misleading one developer. It’s misleading every agent that reads it. The authors recommend that teams review important context artifacts against CAFE(S) the same way they review code, with effort proportional to how widely the context is reused.
The paper is refreshingly practical about where to start. Before any of the five properties matter, three preconditions have to be met: the agent has to be able to reach the context, it has to be able to find the relevant slice, and the knowledge has to exist in writing at all. A lot of organizational knowledge still lives only in hallway conversations and people’s heads, and no amount of context engineering can retrieve what was never captured.
State the goal, how you’d know it’s solved, what’s off limits and why, where the relevant code lives, and when the agent should come back to you. Treat everything an agent might read, from READMEs to code comments, as context held to the same standard.
Specs, AGENTS.md files, skills, glossaries, and architecture docs degrade like any other software unless they’re maintained. Review them before they’re broadly reused, and agree on where each kind of context lives so humans and agents can find it.
Every major shared source should have an owner, ideally a team rather than a person so it survives reorgs. Invest in retrieval that keeps restricted information protected, teach context writing as an engineering skill, and stop rewarding only short-term feature delivery.
The authors add one important caveat. Good context enables autonomy but doesn’t guarantee it. Satisfying CAFE(S) is a prerequisite for trusting an agent with more independence, not a license. Whether to grant that independence still depends on the stakes, your risk tolerance, and who is accountable when something goes wrong.
The authors are explicit about scope. CAFE(S) is a definition, not a measurement system. They describe the properties clearly enough for teams to discuss and review them, and leave the work of measuring context quality and its link to outcomes to future research. They also position it alongside, not in place of, the frameworks engineering organizations already run.
| Framework | What it covers | What CAFE(S) adds |
|---|---|---|
| DORA | Throughput and stability of software delivery | The quality of the information feeding that delivery system |
| SPACE, DX Core 4, EngThrive | The human side of productivity and experience | The shared information environment shaping both human and agent effectiveness |
| RAG evaluation | Retrieval and generation quality | Adds clarity, actionability, and security to a single quality frame |
| OWASP LLM Top 10 | LLM application security risks | Evaluates security alongside usefulness instead of separately |
| Context engineering guidance | How to build context and harnesses | Criteria for judging the result, however it was built |
That framing matters for anyone measuring AI in engineering. If context quality drives agent cost and reliability, then it should show up in the metrics you already track. You may not be able to score an AGENTS.md file directly yet, but you can watch the downstream signals each dimension produces.
| Dimension | When it fails, you’ll see | Signals worth tracking |
|---|---|---|
| CClarity | Agents solving the wrong problem, then humans redoing it | Rework and churn on AI-assisted PRs, review rounds per AI-assisted change |
| AActionability | Runaway sessions, tasks that never close, scope creep | Token spend per task, session length outliers, AI-assisted work that stalls before merge |
| FFidelity | Code built on deprecated APIs or outdated decisions | Change failure rate and defect escape for AI-assisted versus human-authored changes |
| EEfficiency | High cost for routine work, inconsistent results on simple tasks | Tokens per merged PR, cost per AI-assisted change by team and by tool |
| SSecurity | Sensitive data or untrusted content in agent workflows | Security findings on AI-assisted code, which tools and repos agents can reach |
The value of watching these together is comparison. If two teams use the same model and the same tools but one spends twice the tokens per merged PR and reworks twice as often, the difference probably isn’t the model. It’s the context. That’s the gap between teams that CAFE(S) gives you language to describe, and that engineering data lets you find.
Now available in Waydev
Waydev now uses the CAFE(S) framework. It groups the signals above under the five dimensions, so you can see which teams, repos, and tools have context problems, and whether those problems are about clarity, actionability, fidelity, efficiency, or security.
Models will keep getting better. Harnesses will keep getting more sophisticated. But as the paper argues, both are bounded by the quality of the context they receive, and the need for context that’s clear, actionable, accurate, focused, and safe won’t go away with the next model release. Those are properties of good communication, not of any particular technology.
The organizations that get the most out of agents won’t be the ones that buy the most tokens. They’ll be the ones that treat context as a first-class engineering artifact, and measure its effects with the same rigor they apply to code.
Waydev tracks AI impact from code to production and quantifies AI ROI down to every token spent, so you can find which teams are getting value from their agents and which are paying for poor context.
Book a demoSource: Houck, B., Kanat-Alexander, M., Kalliamvakou, E., Storey, M.-A., and Forsgren, N. CAFE(S): Your Agent Is Only As Good As Its Context. ACM Queue, 2026. DOI: 10.1145/3847288. Licensed under Creative Commons Attribution 4.0. The Notion, Meta, Air Canada, EchoLeak, and CamoLeak examples are drawn from sources cited in the paper. Atlassian figures are from The Agentic Pivot (2026). The signals table reflects Waydev’s recommendations and is not part of the CAFE(S) framework.
Ready to unlock your SDLC productivity?