Back To All

C A F E (S) : Five properties of context your agents can actually use

October 1st, 2026
Topics
AI
AI ADOPTION
AI Agents
AI IMPACT
AI ROI
AI SDLC
Share Article

Ask an engineering leader how to get more out of AI coding agents and the conversation almost always turns to models: newer ones, bigger context windows, better reasoning. A new paper in ACM Queue argues that this is the wrong place to look first. The same model can perform very differently depending on what it is given to work with, and most organizations have no way to evaluate that input.

Where CAFE(S) sits in the AI system

Every agent output is shaped by four layers. Models and harnesses get most of the engineering attention. The layer in between gets almost none.

IntentWhat a person actually wants done
ContextThe information that represents the task, constraints, and background
HarnessOrchestration, tools, memory, retrieval
ModelThe raw intelligence

CAFE(S) evaluates one layer: the translation from intent into context. It doesn’t grade the model or the harness. It asks whether what you handed the agent was fit for the job.

The paper is CAFE(S): Your Agent Is Only As Good As Its Context, by Brian Houck (distinguished scientist at DX, now part of Atlassian), Max Kanat-Alexander (Capital One), Eirini Kalliamvakou (GitHub), Margaret-Anne Storey (University of Victoria), and Nicole Forsgren (Google). If those names sound familiar, it’s because several of them are behind SPACE, DevEx, and DORA, the frameworks most engineering organizations already use to measure delivery. DX, the developer intelligence company behind the DX Core 4, is closely tied to this work: Houck leads research there, and Storey collaborates with the DX team.

Their new contribution is a vocabulary for something every team running agents has felt but few can name: the quality of the context those agents work from. Below we break down the framework, the failure patterns it identifies, and what it means for engineering leaders who need to prove their AI investment is working.

The paper at a glance

The core claim: context quality is now a first-order driver of what agentic software costs, how safely it behaves, and whether developers can rely on it.

  1. 1Good context has five properties: clarity, actionability, fidelity, efficiency, and security.5independent dimensions you can review and fix one at a time
  2. 2Poor context now shows up on the invoice, in legal liability, and in daily developer friction.~90%cost drop Notion saw by engineering context for reuse
  3. 3Agent failures follow repeatable patterns, which the authors call context smells.18failure patterns mapped to the five dimensions
  4. 4Where context lives matters as much as what it says. The same instruction can help in one place and hurt in another.1AGENTS.md line can reach thousands of sessions
  5. 5CAFE(S) is a definition, not a measurement system. The authors leave measurement to future work.+DORAdesigned to sit alongside DORA, SPACE, and DevEx, not replace them

Why context quality matters now

None of this is a new problem. Software teams have struggled for decades to document decisions, keep knowledge current, and help people find what they need. The authors’ point is that AI doesn’t create the need for good knowledge management. It dramatically raises the cost of bad knowledge management, and makes that cost visible immediately and at scale.

They identify three forces that have turned context from an abstract concern into an engineering problem.

Developer experienceEvery bad guess an agent makes lands back on a developer to diagnose and redo
CostContext is where tokens accumulate, and redundant context is paid for on every call
LiabilityCompanies stay accountable for what their agents say, and for what gets into their context

Developer experience

This is the least visible cost because it shows up as friction, not a line item. When context is unclear, the agent guesses. When it’s incomplete, the agent produces something plausible but wrong. When it’s stale, the agent confidently applies a policy that no longer exists. Every one of those failures lands back on a developer, who has to notice, diagnose, re-explain, and retry. The agent was supposed to remove toil. Poor context hands it right back.

Cost

Context is where token spend accumulates, and the paper cites two striking examples of what engineering it well is worth:

~90%lower costs at Notion after engineering context for reuse through prompt caching, with output quality maintainedCited in CAFE(S)
~40%fewer tool calls and tokens at Meta after precomputing concise context files for a pipeline spanning 4,100+ filesCited in CAFE(S)
48%fewer tokens, with 44% better answer quality, for agents using Atlassian’s Teamwork Graph context in internal testingAtlassian, The Agentic Pivot

Organizations that repeatedly send agents large, redundant, poorly structured context are simply paying more for the same work, and that gap widens as agents take on longer-running tasks.

Liability

Organizations remain accountable for what their agents say and do. The paper points to the 2024 tribunal ruling that held Air Canada liable for a refund its chatbot promised but its policy didn’t provide. It also cites EchoLeak and CamoLeak, two prompt-injection attacks against Microsoft 365 Copilot and GitHub Copilot Chat in which untrusted content became part of the agent’s working context and the agent faithfully acted on it. Neither system was hacked in the traditional sense. The context was compromised, so everything built on it was too.

The CAFE(S) framework

The authors define good context by what it enables: a human and agent working together can accomplish the task correctly, efficiently, and safely. A useful mental test they offer is to ask what a competent colleague would need to do the same task well if they couldn’t lean over and ask a follow-up question.

From there, they name five properties. The first four describe whether context lets an agent do good work. The fifth, set in parentheses, asks whether the agent should have that context at all.

C
ClaritySpecific and unambiguous

Can the agent interpret the request the way it was intended?

Breaks when a term the task depends on has more than one reasonable reading, and the agent silently picks one.

A
ActionabilityDirective and goal oriented

Can the agent proceed, and does it know when it’s done?

Breaks when the goal, constraints, or definition of done were never written down because they felt obvious to the author.

F
FidelityAccurate and up to date

Can the agent trust that the context is true?

Breaks when docs, APIs, or policies changed and the context didn’t, so the agent reasons correctly from wrong facts.

E
EfficiencyScoped and token optimized

Can the agent focus on what matters?

Breaks when the whole codebase gets sent for a question about one module, and the signal drowns in noise.

S
SecuritySafe and compliant

Should the agent have this context at all?

Breaks when attacker-controlled content, or secrets and personal data the task doesn’t need, end up in the agent’s working context.

Adapted from Table 1 in CAFE(S), ACM Queue (2026).

The dimensions are designed to be independent, which is what makes them useful in practice. Context can be perfectly clear and still describe the system as it was last year. It can be accurate and secure but bury the one detail that matters under thousands of irrelevant tokens. It can be lean and accurate and still never say what finished looks like. Each pillar isolates one thing you can review for and fix on its own.

CAFES
Clear, but out of dateDescribes the system as it was last year
CAFES
Accurate and safe, but noisyThe one detail that matters is buried in thousands of tokens
CAFES
Lean and accurate, but no finish lineNever says what done looks like

Two distinctions in the paper are worth calling out. Clarity and actionability are not the same thing: clarity asks whether the context can only be read one way, while actionability asks whether enough was said at all. And efficiency is not brevity for its own sake. The goal is the highest signal-to-noise ratio, not the smallest possible context. A slightly verbose context that’s complete beats a tight one that’s missing a constraint.

The authors also explain why they stopped at five. Timeliness folds into fidelity, since stale context is just context that’s no longer true. Cost is a downstream result of efficiency. Provenance is a way of protecting fidelity. Accessibility, whether information can be found at all, is deliberately treated as a prerequisite rather than a quality dimension.

The question isn’t whether your agent has context. It’s whether the context it has is actually fit for the task.

Context smells: how agent failures actually happen

The most practical part of the paper is a catalog of recurring failure patterns. Just as engineers use “code smell” to name recurring design problems, the authors propose context smells. The key insight is that in many of these cases, the necessary information was already present. It was stale, ambiguous, contradictory, buried, or mixed with content that should never have been trusted.

Common context smells and the dimensions they violate
PatternWhat happensDimension
Specification ambiguityA semi-ambiguous request leads the agent to pick an interpretation and build on it, compounding the error at every stepClarity
Referential ambiguitySeveral entities match the same reference, so the agent correctly follows the request on the wrong objectClarity
Weekend runawayNo stopping criteria, so the agent retries and expands scope for hours, running up a large billEfficiencyActionability
Goal driftThe agent gradually optimizes for a proxy of the task instead of what the user actually wantedActionability
Negative-constraint fragilityThe context lists what not to do without saying what to do insteadActionability
Lost in the detailsImplementation detail without rationale, so the agent optimizes local mechanics and loses the pointActionabilityEfficiency
Stale guidanceThe context was accurate once, but the system moved onFidelity
Confident hallucinationUnsupported claims get presented as fact because nothing in the workflow treats them as needing verificationFidelity
Schema driftTool inputs or outputs change shape over time, breaking downstream steps even though each component still worksFidelity
Contradictory contextTwo sources disagree, and the agent has to arbitrate between incompatible versions of realityClarityFidelity
Lost in the middleThe right information is present but overlooked because of volume or placementEfficiency
OverspecificationPrescribing the how in more detail than needed, displacing what the model would have done better unguidedEfficiencyActionability
Indirect prompt injectionRetrieved content carries instructions that compete with or override the real taskSecurity
Paste-in-the-prompt leakUsers paste confidential data into external AI toolsSecurity
A selection of the 18 patterns in Table 2 of CAFE(S), paraphrased.

Where the 18 failure patterns land

Number of patterns in the paper that involve each dimension. Some patterns involve two.

Actionability6
Security5
Fidelity4
Efficiency4
Clarity3

Actionability is the most common failure. Missing goals, missing constraints, and no definition of done cause more of these patterns than any other dimension. It’s also the cheapest to fix: write down what finished looks like.

Read that list as an engineering leader and a pattern jumps out. Almost every one of these smells shows up in delivery data before anyone names it. A weekend runaway is a token spike. Specification ambiguity is a PR that gets reworked after review. Stale guidance is an agent that keeps proposing a deprecated approach. The failures are context problems, but their fingerprints land in your engineering metrics.

Context quality depends on where it lives

One principle in the paper is easy to overlook and very practical: the same words can be good context in one place and bad context in another.

The authors use the example of an AGENTS.md file at a repository’s root, which gets loaded into every agent session. An instruction about running the test suite belongs there because nearly every session needs it. An instruction about one rarely touched file does not. Placed there, it gets injected into thousands of sessions that have nothing to do with it, costing tokens and attention every time. Move the same sentence into a comment inside that file and it becomes high-quality context, available exactly when it’s relevant. Nothing about the words changed. Only their location did.

AGENTS.md (repo root)Loaded every session
## Testing
Run `make test` before opening a PR.

Don't change the rounding in
billing/legacy_tax.py. Finance
reconciles against it.
Injected into thousands of sessions that never touch billing. Costs tokens and attention every time.
billing/legacy_tax.pyLoaded when relevant
# Don't change the rounding below.
# Finance reconciles against it.
def round_tax(amount):
    ...
Same words, now read only by the agents working on this file. Clear, efficient, and right on time.

This has a direct implication for review. Because shared context is read by every agent session that loads it, its defects are amplified. A stale architecture doc isn’t misleading one developer. It’s misleading every agent that reads it. The authors recommend that teams review important context artifacts against CAFE(S) the same way they review code, with effort proportional to how widely the context is reused.

What to do about it, at every level

The paper is refreshingly practical about where to start. Before any of the five properties matter, three preconditions have to be met: the agent has to be able to reach the context, it has to be able to find the relevant slice, and the knowledge has to exist in writing at all. A lot of organizational knowledge still lives only in hallway conversations and people’s heads, and no amount of context engineering can retrieve what was never captured.

Step 1ReachableThe source is connected to the agent
Step 2FindableThe agent can pull the relevant slice
Step 3Written downThe knowledge exists outside people’s heads
ThenCAFE(S)Quality becomes the constraint
  • Individuals: brief agents like a senior engineer briefs a junior one

    State the goal, how you’d know it’s solved, what’s off limits and why, where the relevant code lives, and when the agent should come back to you. Treat everything an agent might read, from READMEs to code comments, as context held to the same standard.

  • Teams: treat shared context as code

    Specs, AGENTS.md files, skills, glossaries, and architecture docs degrade like any other software unless they’re maintained. Review them before they’re broadly reused, and agree on where each kind of context lives so humans and agents can find it.

  • Organizations: give context owners and review cadences

    Every major shared source should have an owner, ideally a team rather than a person so it survives reorgs. Invest in retrieval that keeps restricted information protected, teach context writing as an engineering skill, and stop rewarding only short-term feature delivery.

The authors add one important caveat. Good context enables autonomy but doesn’t guarantee it. Satisfying CAFE(S) is a prerequisite for trusting an agent with more independence, not a license. Whether to grant that independence still depends on the stakes, your risk tolerance, and who is accountable when something goes wrong.

CAFE(S) defines quality. Someone still has to measure it.

The authors are explicit about scope. CAFE(S) is a definition, not a measurement system. They describe the properties clearly enough for teams to discuss and review them, and leave the work of measuring context quality and its link to outcomes to future research. They also position it alongside, not in place of, the frameworks engineering organizations already run.

How CAFE(S) fits with what you already use
FrameworkWhat it coversWhat CAFE(S) adds
DORAThroughput and stability of software deliveryThe quality of the information feeding that delivery system
SPACE, DX Core 4, EngThriveThe human side of productivity and experienceThe shared information environment shaping both human and agent effectiveness
RAG evaluationRetrieval and generation qualityAdds clarity, actionability, and security to a single quality frame
OWASP LLM Top 10LLM application security risksEvaluates security alongside usefulness instead of separately
Context engineering guidanceHow to build context and harnessesCriteria for judging the result, however it was built
Paraphrased from Table 3 in CAFE(S).

That framing matters for anyone measuring AI in engineering. If context quality drives agent cost and reliability, then it should show up in the metrics you already track. You may not be able to score an AGENTS.md file directly yet, but you can watch the downstream signals each dimension produces.

From CAFE(S) dimension to signals in your engineering data
DimensionWhen it fails, you’ll seeSignals worth tracking
CClarityAgents solving the wrong problem, then humans redoing itRework and churn on AI-assisted PRs, review rounds per AI-assisted change
AActionabilityRunaway sessions, tasks that never close, scope creepToken spend per task, session length outliers, AI-assisted work that stalls before merge
FFidelityCode built on deprecated APIs or outdated decisionsChange failure rate and defect escape for AI-assisted versus human-authored changes
EEfficiencyHigh cost for routine work, inconsistent results on simple tasksTokens per merged PR, cost per AI-assisted change by team and by tool
SSecuritySensitive data or untrusted content in agent workflowsSecurity findings on AI-assisted code, which tools and repos agents can reach
Waydev’s recommendations. These are proxy signals for context quality, not direct measures of it.

The value of watching these together is comparison. If two teams use the same model and the same tools but one spends twice the tokens per merged PR and reworks twice as often, the difference probably isn’t the model. It’s the context. That’s the gap between teams that CAFE(S) gives you language to describe, and that engineering data lets you find.

Now available in Waydev

CAFE(S) is built into Waydev

Waydev now uses the CAFE(S) framework. It groups the signals above under the five dimensions, so you can see which teams, repos, and tools have context problems, and whether those problems are about clarity, actionability, fidelity, efficiency, or security.

CClarity AActionability FFidelity EEfficiency SSecurity
See CAFE(S) in Waydev

The bottom line

Models will keep getting better. Harnesses will keep getting more sophisticated. But as the paper argues, both are bounded by the quality of the context they receive, and the need for context that’s clear, actionable, accurate, focused, and safe won’t go away with the next model release. Those are properties of good communication, not of any particular technology.

The organizations that get the most out of agents won’t be the ones that buy the most tokens. They’ll be the ones that treat context as a first-class engineering artifact, and measure its effects with the same rigor they apply to code.

See what your agents are really costing you

Waydev tracks AI impact from code to production and quantifies AI ROI down to every token spent, so you can find which teams are getting value from their agents and which are paying for poor context.

Book a demo

Source: Houck, B., Kanat-Alexander, M., Kalliamvakou, E., Storey, M.-A., and Forsgren, N. CAFE(S): Your Agent Is Only As Good As Its Context. ACM Queue, 2026. DOI: 10.1145/3847288. Licensed under Creative Commons Attribution 4.0. The Notion, Meta, Air Canada, EchoLeak, and CamoLeak examples are drawn from sources cited in the paper. Atlassian figures are from The Agentic Pivot (2026). The signals table reflects Waydev’s recommendations and is not part of the CAFE(S) framework.

Ready to unlock your SDLC productivity?

Request a Demo Call