Kapa releases Company Knowledge Bench, a 1,000-case retrieval eval for agents
Kapa, a platform that indexes company knowledge for AI agents, built an internal benchmark of 1,000 annotated eval cases from real production data to compare retrieval methods. The company tested seven retrievers—including traditional hybrid search, agentic grep-based search, and its own Kapa systems—measuring both accuracy and query time. Results showed Kapa's 'Deep' retriever scoring 0.65 in about five seconds, while a frontier model using only grep matched a tuned pipeline's 0.61 score but took roughly five times longer.
GoKawiil's interpretation of the reporting above, not reported fact.
The comparison suggests that for messy, real-world company data like wikis, tickets, and chat logs, retrieval speed and accuracy tradeoffs vary widely depending on method, which could inform how teams building AI agents choose their search infrastructure. Kapa acknowledges its own bias, since it built all seven retrievers and used its own ingestion pipeline, meaning outside verification would be needed to confirm these results generalize. The post also implies that standard public retrieval benchmarks may not reflect the realities of enterprise knowledge, a gap this benchmark aims to address.
- Kapa created a 1,000-case internal benchmark testing retrieval methods on real company data.
- Its 'Kapa Deep' retriever scored highest (0.65) while taking about five seconds per query.
- A grep-only agentic approach matched a tuned retrieval pipeline's score but was roughly five times slower.
Source: kapa.ai, 2026-10-02
Published there as: “Benchmarking retrieval for agents on messy real-world company knowledge”
Read the original report → The summary and analysis above are GoKawiil's own, written from reporting by the source above. Facts and quotes belong to the original publisher.