Kapa tests agent retrieval on 1,000 real company queries Kapa co-founder and CTO Finn Bauer said the company's Kapa Deep retrieval system scored 0.65 on Company Knowledge Bench, a private benchmark published October 2nd built from 1,000 real production queries, edging out Kapa's Default mode and a grep-based agent at 0.61 each. Kapa built all seven compared systems and ingested the documents through its own pipeline, so the results are not an independent ranking; Kapa Deep answered in about five seconds and returned roughly 5,000 tokens, while the grep agent took 13 to 17 seconds and returned about 40,000 tokens. Kapa reports hybrid search scored 0.41, adding a reranker lifted it to 0.50, and query decomposition raised it to 0.56. Kapa tests agent retrieval on 1,000 real company queries Co-founder Finn Bauer says Kapa Deep scored 0.65 on a private benchmark comparing seven Kapa-built systems using data ingested through Kapa. By RuntimeWire Staff https://runtimewire.com/author/runtimewire-staff ยท Published Primary source: Kapa https://www.kapa.ai/blog/company-knowledge-bench Why it matters Retrieval quality affects whether workplace agents can answer accurately, quickly and without flooding another model with irrelevant context. Kapa's private, self-built benchmark makes a concrete engineering case while limiting independent verification. Kapa https://www.kapa.ai/?ref=runtimewire co-founder and CTO Finn Bauer https://twitter.com/finnbauer5?ref=runtimewire says the company's retrieval system beat six alternatives on a new benchmark built from 1,000 real production queries. The benchmark shows how Kapa wants agents to search the untidy records businesses actually use. Bauer's Company Knowledge Bench https://www.kapa.ai/blog/company-knowledge-bench?ref=runtimewire , published October 2nd, puts Kapa Deep at a retrieval score of 0.65 and about five seconds per query. The closest alternatives scored 0.61: Kapa's faster Default mode and a model-driven agent using grep. Kapa says the grep system took roughly five times as long as Deep. The comparison documents Kapa's engineering choices. Kapa built all seven systems, ran them on documents ingested through its pipeline and has not made the benchmark public, so it is not an independent ranking of retrieval products. A benchmark shaped by production questions Bauer comes to the problem from systems engineering. Y Combinator's profile https://www.ycombinator.com/companies/kapa-ai?ref=runtimewire identifies him as Kapa's co-founder and CTO and says he previously worked as a software engineer at Bloomberg, building low-latency trading systems. He and co-founder and CEO Emil Soerensen https://twitter.com/EmilSoehr?ref=runtimewire studied computer science at Imperial College London and finance at the London School of Economics, according to the profile. Kapa's original focus was helping technical companies make complex products easier to use. In its account of the company's founding, the founders said two tech companies approached them in January 2023 with the same problem: support teams were fielding repetitive questions despite having extensive documentation. Kapa has since expanded the retrieval use case it describes, from customer-facing product help to agents searching internal documents, tickets, code and workplace conversations. The benchmark defines a good result as one that returns enough information to answer the question, avoids unnecessary material and favors better sources. A current Slack discussion can outrank an old internal document; a dedicated reference page can outrank a forum post that repeats the same fact. For ambiguous queries, Kapa says its standard requires retrieval to include material for each reasonable interpretation, leaving the agent to choose among them. The test set covers developer questions about documentation, APIs, code and GitHub issues; employee questions drawing on Slack, Confluence, Notion and Google Drive; and support work using tickets, help-center articles and internal handbooks. Kapa is selling a shared retrieval layer for agents that need to search across a company's scattered knowledge. The score has a boundary Kapa reports that hybrid search scored 0.41, adding a reranker lifted the score to 0.50, and query decomposition raised it to 0.56. Its Default mode scored 0.61 in 3.3 seconds. The larger grep-based agent also scored 0.61, but Kapa says it took 13 to 17 seconds per query and returned about 40,000 tokens. Deep scored 0.65 in about five seconds and returned roughly 5,000 tokens. The figures measure retrieval accuracy alongside query time and the amount of material returned. Every unnecessary chunk consumes time and adds material for the answer-generating model to process. Kapa argues that an agent that searches iteratively and prunes irrelevant documents can return a more useful result without making each query slower or more expensive to consume. The labels behind the scores are a second limitation. Kapa says agents generated labels for the 1,000 cases after the team tested the labeling agents against 170 cases marked by humans. The company says it improved the labeling agents until their agreement with human reviewers was high enough, but the post does not report a numerical agreement rate. Because the benchmark is private, outsiders cannot inspect the cases, test the labels or reproduce the comparison. The scores are an internal measurement, not a portable yardstick for buyers choosing retrieval systems. Kapa says its production data is why the benchmark cannot be public. The data may better reflect the company's customers than a synthetic test set, while the private methodology makes it harder for anyone outside Kapa to judge how representative those cases are. For Bauer and Kapa, the benchmark turns the company's own product challenge into an evaluation system: measure how well agents find the right answer across messy company records, then use those measurements to improve retrieval. Its evidence is comparative within that system. Whether its definitions and findings can help others make the same engineering decisions without access to Kapa's underlying data remains an open question.