{"slug": "keenable-ai-open-sources-needle-a-live-search-benchmark-that-rebuilds-its-query", "title": "Keenable AI Open-Sources NEEDLE: A Live Search Benchmark That Rebuilds Its Query Set Every Hour", "summary": "Keenable AI released NEEDLE, an open-source live search benchmark that regenerates its query set hourly from RSS feeds and Google Trends for news and daily from SEC XBRL, arXiv, Europe PMC, CourtListener, and public agent logs for finance, scholar, legal, and rare-entity queries. The benchmark runs 15 search APIs under one protocol and scores results against a pooled oracle engine called 'ultimate' to measure the ceiling of agentic search quality. Published 7-day means for the window ending 2026-08-28 show finance queries nearly solved, with Exa scoring 0.910 and Keenable 0.872.", "body_md": "How do you benchmark a web search API when the thing being tested can read the answer key? A search agent has a fetch tool. If the gold labels sit in a public dataset, the agent can download them mid-evaluation and skip retrieval entirely. A similar problem arises when the answers are already encoded in the model’s parametric memory: a correct response no longer demonstrates that web search worked. Keenable’s answer is [NEEDLE](https://keenableai.github.io/needle/), a live open-source benchmark that rebuilds its query set from fresh public sources rather than freezing one. News queries are regenerated hourly from RSS feeds and Google Trends; finance, scholar, legal, and rare-entity queries are regenerated daily from SEC XBRL, arXiv, Europe PMC, CourtListener, and public agent logs. Fifteen search APIs run against the same query text under one protocol, and every score is read against *ultimate*, a pooled oracle engine that marks what the whole field managed to find.\n\n**Is it reproducible?**\n\nYes, as an open source evaluation harness rather than a product. [needle](https://github.com/keenableai/needle) is a Python CLI installed with uv sync and driven by two subcommands per benchmark, generate and run. It needs an OpenRouter key for judging and one API key per engine tested, and runs on a laptop or in CI. It allows recreated all query streams that are being used in addition to the ranking quality judgements.\n\n**What NEEDLE measures**\n\n[NEEDLE](https://keenableai.github.io/needle/) stands for News, Everyday, Expert, Deep-tail, and Legal Evaluation. Each vertical models a different agent intent. **News** projects the newest item from ~124 curated RSS feeds and Google Trends into a keyword query. **Finance** asks registry facts from [Wikidata](https://www.wikidata.org/) and GLEIF plus single-quarter 10-Q figures from SEC XBRL. **Scholar** turns one paper into four query styles: a degraded title, a full-text-only detail, a natural-language clue, and a hedged tip-of-the-tongue description. **Deep-tail** samples rare-word queries from public agent-trajectory releases including [DeepResearchGym](https://arxiv.org/abs/2601.17617), [OpenResearcher](https://arxiv.org/abs/2603.20278) and [LRAT](https://arxiv.org/abs/2604.04949). **Legal** pulls recent CourtListener opinions across 14 federal courts and eCFR sections.\n\nScoring splits along the same line. News and deep-tail have no single correct result, so an LLM judge rates each result 0 to 4 and the harness reports nDCG@5 with a duplicate-URL penalty. Finance reports answer-recall@5: does the fact reach the agent inside a top-5 snippet. Scholar and legal are known-item tasks scored by identifier match.\n\n**The ceiling is the interesting part**\n\nEvery engine receives the same query text. The runner issues one call at a time, so latency percentiles are comparable and no engine takes concurrent load. Judging happens on the engine’s own ranking, titles and snippets. Pages are never fetched and results are never re-ranked. Evidence is clipped to 2,000 characters for everyone, and the judge does not see the engine name.\n\nThe more interesting number is the *ultimate ceiling*. For each query, NEEDLE pools the results returned by every engine into a synthetic oracle engine, then orders that combined set by relevance. That creates an empirical ceiling based on what the entire field was able to retrieve.\n\nThe gap to *ultimate* is therefore an upper bound on agentic search quality as it stands today. A large gap means better results existed but every engine failed to surface or rank them well. A weak *ultimate* score means something different: even after pooling every provider, the benchmark found little strong evidence. In other words, NEEDLE can distinguish a ranking problem from a retrieval problem shared by the whole market.\n\n**Where the field actually stands**\n\nNumbers below are published 7-day means for the window ending 2026-08-28.\n\nFinance is close to solved: Exa 0.910, Keenable 0.872, Perplexity 0.871, Google 0.847, against an *ultimate* of 0.965. Scholar spreads out, Keenable 0.774 to Tavily 0.310 against a 0.869 ceiling, because title queries are answerable from metadata and body queries are not. Deep-tail is hardest and closest to real agent traffic: Exa leads at 0.557 of *ultimate*, Keenable follows at 0.470, Bing sits at 0.199. The gap between delivered and achievable quality widens as queries approach how agents actually search.\n\nLatency is another important metric here, because agents call search dozens of times per task. Same window: Keenable-realtime 193 ms p50 / 284 ms p95, Exa 1,876 / 2,955, Bing 2,767 / 9,381.\n\n**Key Takeaways**\n\n- NEEDLE regenerates queries hourly for news and daily for the other four verticals, so there is no fixed set to overfit.\n- Five verticals, 15 search APIs, one protocol: same query text, same 2,000-character evidence cap, one request at a time.\n- Every leaderboard is read against\n*ultimate*, a pooled oracle engine marking the ceiling the whole field reached. - On rare-entity queries from real agent logs the top engine reaches 0.557 of that ceiling; on finance most engines cluster between 0.77 and 0.91.\n- Code is MIT, runs execute in public GitHub Actions, and per-run artifacts ship to a Hugging Face dataset.\n\nCheck out the [ live dashboard](https://keenableai.github.io/needle/), the\n\n[, the](https://github.com/keenableai/needle)\n\n**GitHub repo**[, and the](https://keenable.ai/blog/needle-the-benchmark-your-search-engine-can-t-memorize)\n\n**technical write-up**[. All credit goes to the researchers of this project.](https://huggingface.co/datasets/keenable-ai/needle-results)\n\n**archived artifacts** Also, feel free to follow us on ** Twitter** and don’t forget to join our\n\n**and Subscribe to**\n\n[150k+ML SubReddit](https://www.reddit.com/r/machinelearningnews/)**. Wait! are you on telegram?**\n\n[our Newsletter](https://magic.beehiiv.com/v1/f5e63dd4-5653-4f09-83e2-321a8b1ba526?email={{email}})\n\n[now you can join us on telegram as well.](https://t.me/machinelearningresearchnews)Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? [Connect with us](https://forms.gle/wbash1wF6efRj8G58)\n\nAsif Razzaq is the CEO of Marktechpost Media Inc.. As a visionary entrepreneur and engineer, Asif is committed to harnessing the potential of Artificial Intelligence for social good. His most recent endeavor is the launch of an Artificial Intelligence Media Platform, Marktechpost, which stands out for its in-depth coverage of machine learning and deep learning news that is both technically sound and easily understandable by a wide audience. The platform boasts of over 2 million monthly views, illustrating its popularity among audiences.", "url": "https://wpnews.pro/news/keenable-ai-open-sources-needle-a-live-search-benchmark-that-rebuilds-its-query", "canonical_source": "https://www.marktechpost.com/2026/08/31/keenable-ai-open-sources-needle-a-live-search-benchmark-that-rebuilds-its-query-set-every-hour/", "published_at": "2026-08-31 23:43:26+00:00", "updated_at": "2026-08-31 23:51:33.955237+00:00", "lang": "en", "topics": ["ai-research", "ai-tools", "large-language-models"], "entities": ["Keenable AI", "NEEDLE", "Exa", "SEC XBRL", "arXiv", "Europe PMC", "CourtListener", "Google Trends"], "alternates": {"html": "https://wpnews.pro/news/keenable-ai-open-sources-needle-a-live-search-benchmark-that-rebuilds-its-query", "markdown": "https://wpnews.pro/news/keenable-ai-open-sources-needle-a-live-search-benchmark-that-rebuilds-its-query.md", "text": "https://wpnews.pro/news/keenable-ai-open-sources-needle-a-live-search-benchmark-that-rebuilds-its-query.txt", "jsonld": "https://wpnews.pro/news/keenable-ai-open-sources-needle-a-live-search-benchmark-that-rebuilds-its-query.jsonld"}}