Building an enterprise AI benchmark changed how I evaluate AI DevRev released Enterprise-Bench, an open-source enterprise AI benchmark built on a synthetic midmarket payments company with 42 customer accounts, 40 product parts, five interconnected enterprise systems and 14 tasks, after finding that context assembly rather than model reasoning was the dominant bottleneck. Scaling surrounding data volume up to 256 times without changing the correct answer dropped the relevant-data share from about 40% at the smallest scale to roughly 0.16% at the largest, and the author argues accuracy losses under that noise indicate retrieval architecture failure rather than model failure. I spent the early part of my career building database systems, managing Oracle’s storage engine group and later helping build Aster Data. That work taught me to watch the gap between benchmark results and production behavior. When the Transaction Processing Performance Council TPC was formed https://www.tpc.org/information/sessions/tpc history.pdf in 1988, it was because vendors, customers and researchers lacked a common language for comparing transaction systems. The benchmarks worked, but they also created an incentive to tune configurations around the test rather than around production. The response was to add oversight through fair-use policies, peer review and independent audits. The lesson has stayed with me: a benchmark result means something when the improvement behind it survives contact with a customer environment. Enterprise AI has reached the same point. So, we built a benchmark around the conditions enterprise AI actually has to survive. The hardest part of enterprise AI is not making a model reason. It is giving that model the right sliver of trusted context without making it reconstruct the company every time. When we began building Enterprise-Bench https://github.com/devrev/enterprise-bench , I expected model reasoning to dominate the work. Instead, we kept returning to the same bottleneck: assembling the right context. The models could usually answer a straightforward business question once they had the right information. Getting them that information, with the right permissions and without drowning them in irrelevant data, was a much harder engineering problem. Stanford’s Holistic Evaluation of Language Models HELM https://crfm.stanford.edu/helm/index.html shows how public benchmarks use standardized datasets to compare model performance consistently. Inside an enterprise, even a straightforward question can be difficult to answer because the relevant data is scattered across systems, inconsistently described and governed by different permissions. “Which customers are affected by this bug and what is its impact?” is not a math Olympiad problem. Yet the answer may require an agent to connect a support ticket to a product component, match that component to an engineering issue, identify the affected accounts in a CRM, calculate the revenue exposure and return only what the user is permitted to see. The model can only reason about what the system can find, connect and safely expose. Some of the most revealing tasks in the benchmark were the ones that looked easy. Asking for the number of open tickets on an account is a lookup. Asking which priority-one tickets correspond to unresolved engineering issues, and how much customer revenue is exposed, is also a lookup. It just crosses more systems and more access boundaries to get there. Enterprise AI evaluations often misdiagnose integration failure as model failure. A model guesses because two applications name the same product differently. It misses a customer because the relationship between a ticket and an opportunity exists only through a third object. It retrieves a stale document because the connector indexed a snapshot rather than maintaining the current state. Buying a stronger model may improve the fluency of the answer. It does not repair the missing relationship. We modeled the benchmark on a synthetic midmarket payments company with 42 customer accounts, 40 product parts, five interconnected enterprise systems and 14 tasks spanning engineering, sales and support. Then we increased the surrounding data volume by up to 256 times without changing the correct answer. At the smallest scale, about 40% of the available data was relevant to a task. At the largest, roughly 0.16% was relevant. The question did not become harder, and the answer did not change. Only the amount of noise grew. Answer-preserving scaling exposes what smaller tests hide. If accuracy falls as irrelevant data grows, the retrieval architecture has become less effective. If an agent repeatedly fetches broad sets of records and discards the context, token costs will climb with the company. Enterprises may hold trillions of tokens across customer, product, employee and service history. Even a million-token context window is a workspace, not organizational memory. The problem is not fitting everything into the prompt. It is identifying the smallest trustworthy set of facts needed for the decision at hand. Most enterprise AI starts federated. A model connects to a CRM, a support platform, an issue tracker and a document repository through APIs, and when someone asks a question, it retrieves from each source and assembles an answer. That works for discovery. It struggles the moment a question spans the operating model of the business, because the system has to reconstruct identity, relationships, permissions and current state on every single request. The word “harness” can make this sound like an implementation detail. It is not. The harness is the system around the model: memory, retrieval, tools, permissions, evidence and the logic that brings context together. In our initial comparison, we held the foundation model, tasks, data and independent judge constant. Our structured-memory system completed 94.3% of the tasks correctly, compared with 63.6% for Claude Code using the same Opus 4.8 model family. It also used about 4.4 times fewer tokens per correct answer at production scale. The same model produced materially different outcomes depending on the architecture around it. For decades, enterprise architecture has separated compute from state. We do not ask an application to rebuild its database every time it receives a request. We index, cache, preserve relationships and enforce access close to the data. I fully believe that enterprise AI needs the same discipline. The split is between what is deterministic and what is not. Resolving which account a contract refers to, whether a permission applies and which version of a fact is current is work that software can do exactly. Interpreting a badly worded question, weighing ambiguous evidence and helping someone decide is work that models do well. Federated architectures blur that line and ask probabilistic inference to rebuild structures a database could simply maintain. Useful memory is current, structured and permission-aware. It understands that the customer in a contract, the account in a CRM and the organization in a support platform may be the same entity. It keeps those relationships current, preserves provenance and knows that shared memory does not mean universally visible memory. This is why I believe enterprises should invest in their own memory and workflow layer before trying to build or tune their own models. An enterprise strategy should outlast any single model https://www.cio.com/article/4204569/the-enterprise-ai-strategy-that-outlasts-any-single-model.html . Models will improve, specialize and change in price, but the relationships, permissions, skills, decisions and accumulated judgment that make up a company’s intellectual property should remain under its control. The benchmark also changed how I think about autonomy. The industry often presents autonomy as a feature that can be switched on. In an enterprise, it should be treated as a promotion that an agent earns. Our framework progresses from reactive retrieval at L1 to analytical reasoning at L2, then to proactive coordination at L3, and finally to extended self-directed operation at L4. The current public suite tests only L1 and L2. We expect L3 systems to begin moving into true production over the next six months, while some more advanced enterprises are already testing them internally. Before an agent is trusted to manage a customer escalation, renewal process or another multi-step workflow, it should repeatedly prove that it can retrieve the right facts, reconcile evidence across systems, preserve access controls, explain the basis of its answer and produce an auditable record as the work progresses. Reliability at one level should become the entry requirement for the next. If an agent cannot read consistently, it has not earned the right to write. CIOs do not need another leaderboard that declares a universally smarter model. They need evaluations that reproduce the conditions under which their own AI will operate. I would test seven things: Most importantly, choose tasks that matter to the business. Ask which product problems are putting renewals at risk, what commitments were made on the last customer call, or which support cases have breached the terms of their actual service agreements. A polished demo can hide weak architecture. A question grounded in your operating reality usually cannot. Building an enterprise benchmark changed the question I ask about AI. Instead of starting with, “How intelligent is the model?” I now ask: “Can the system assemble the right context, at the right moment, for the right person and prove that it did so?” The trillion-dollar enterprise AI bet is not how many tokens we can squeeze into a context window. It is how few of the right tokens we need to deliver a precise, efficient and safe answer.