Welcome to the latest issue of Engineering Enablement, a weekly newsletter sharing research and perspectives on developer productivity.
Whenever a new evaluation measurement framework shows up, the first reasonable question is which existing one it competes with. That reflex is healthy. New frameworks often overlap with things you’re already measuring, and every addition comes with some cost. Anyone who has watched a measurement program collapse under its own weight has learned to ask what is being displaced before adopting something new.
I wanted to answer that question for my recently published CAFE(S) framework, because it doesn’t compete with what you’re already running. CAFE(S) is a framework for evaluating the quality of the context we give AI agents across five dimensions: Clarity, Actionability, Fidelity, Efficiency, and (Security). It asks whether the information agents have to work with is good enough for the tasks they’re trying to accomplish.
DORA looks at your delivery system. SPACE, DX Core 4, and EngThrive look at developers and the outcomes they’re working toward. Retrieval and model evals assess what the AI retrieved and produced. OWASP looks at the risk surface. Context and harness engineering guidance tells you how to build the environment the agent works in.
What none of your existing frameworks were built to look at is the information the agent was handed in the first place. I’ve argued that AI didn’t create a context problem, but pushed an existing issue past its breaking point.
This isn’t an argument to drop what you’re already running, or to blindly bolt one more framework onto the stack. The rest of this piece looks at where CAFE(S) sits alongside existing frameworks you may be using. In doing so, I hope to make the boundaries between all of them a little clearer.
Delivery and experience metrics tell you something is wrong, not where
DORA, SPACE, DX Core 4, and EngThrive are all anchored to outcomes, which is exactly why they’ve survived several technology cycles and why I don’t expect AI to invalidate any of them.
But an outcome moving in the wrong direction is a problem, not a diagnosis. Change failure rate climbing, DXI sliding, or innovation rate stagnating tells you where to start asking questions, not what caused the change.
Core 4 and EngThrive both explicitly carry diagnostic dimensions alongside the outcome ones, so this isn’t a clean lagging vs. leading split, but those diagnostics aren’t focused on the assembled context as a first-class engineering artifact.
We’re still trying to accomplish the same top-level outcomes in our agentic workflows as we were before. What has changed is the diagnostic surface, requiring us to look in new places for answers when we’re not achieving them.
Ask yourself: when a delivery or experience metric moves the wrong way, can you tell whether the information your agents worked from had anything to do with it?
Retrieval evals tell you whether retrieval worked, not whether the context worked
Retrieval-augmented generation (RAG) evaluation frameworks measure the quality of the information retrieved and the responses generated from it, which touches on the Fidelity and Efficiency dimensions of CAFE(S). If you run RAG evals well, you already have two of the five properties partially covered.
The other three (Clarity, Actionability, and Security) are outside their frame. A retrieval eval does not ask whether the task was stated in a way that admits a single interpretation, whether anyone defined what finished looks like, or whether attacker-controlled content made it into the window.
There’s an important difference in what gets evaluated. A retrieval eval scores the information your retriever fetched. The agent reasons over the whole assembled context, which also includes the human’s prompt, the root AGENTS.md file, tool output, prior turns, and whatever memory the harness decided to carry forward. Most of that never passed through your retriever, so most of it never appears in your eval.
Ask yourself: are you evaluating the passages your retriever returned, or everything the agent actually read?
OWASP evaluates security risk, not context quality
OWASP goes much deeper on security risk than CAFE(S) does. The ‘S’ pillar is one property among five; OWASP provides focused security guidance. If you’re choosing between them for security depth, you shouldn’t choose CAFE(S). But you shouldn’t have to choose between them.
CAFE(S) puts security inside a broader quality model for context. That’s useful because security and usefulness aren’t always separable. Oversharing, for example, can be both a security failure and an efficiency failure. Trimming unnecessary context can simultaneously reduce token spend and shrink the attack surface.
Failures of the first four dimensions of CAFE(S) make agents less effective. Failures of security make them unsafe.
Ask yourself: are you evaluating your context for both usefulness and safety?
Context engineering tells you how to build context, not whether it’s good
Anthropic’s context engineering guidance, Birgitta Böckeler’s harness engineering guidance, and OpenAI’s harness engineering guide all offer practices for building better context and harnesses. CAFE(S) starts at the other end: given the context an agent actually received, was it any good?
That’s the difference between prescriptive and evaluative. You can follow good context engineering practices and still end up with context that is ambiguous, incomplete, stale, bloated, or unsafe. Conversely, CAFE(S) doesn’t tell you which retrieval strategy, memory architecture, or harness design to use. It gives you properties for evaluating the result.
Ask yourself: do you have a definition of good context that is independent of how you built it?
TRUCE evaluates code quality, not context quality
TRUCE is a multidimensional framework for evaluating code quality and is the closest conceptual sibling to CAFE(S).
CAFE(S) applies a similar idea to a different engineering artifact. As context becomes a first-class input to agentic software development, its quality matters independent of the quality of the code that eventually comes out.
TRUCE gives us a language for thinking about the quality of code. CAFE(S) is intended to do the same for context.
Ask yourself: if context is becoming a first-class engineering artifact, do you have a way to evaluate its quality?
Final thoughts #
So, do you need another framework? Maybe that’s the wrong question.
DORA, SPACE, retrieval evals, OWASP, TRUCE, and CAFE(S) aren’t competing ways to measure the same thing. They answer different questions about different parts of the engineering system.
CAFE(S) adds one that has become increasingly important in agentic development: was the context we gave the agent any good?
It doesn’t replace what you’re already using. It fills the gaps between them.
That’s it for this week. Thanks for reading.
-Brian