{"slug": "exa-vs-tavily-in-2026-our-agent-web-search-benchmark", "title": "Exa vs Tavily in 2026: Our Agent Web Search Benchmark", "summary": "A developer benchmarked Exa and Tavily's agent-oriented search APIs on a fixed suite of 300 company-news questions, finding Exa fast reached 99.3% answer accuracy versus 87.7% for Tavily basic and 93.0% for Tavily advanced. The harness used direct HTTP calls rather than SDKs and graded whether returned results contained enough evidence to recover a pre-established answer, targeting grounding reliability rather than raw latency. The writeup also cites vendor-originated entity-retrieval figures from Exa's 2026 comparison, which it flags as non-neutral evidence.", "body_md": "Search latency was not the original problem. Grounding reliability was.\n\nAnswer accuracy on 300 company-news questionsExa fast 99.3%/100Exa instant 97.7%/100Tavily advanced 93.0%/100Tavily basic 87.7%/100\n\nOn the same fixed suite, Exa fast and instant beat Tavily advanced and basic on answer accuracy, with Exa fast at 99.3% versus 87.7% for Tavily basic.\n\nIn agent traces, we kept finding steps that completed quickly but returned the wrong source type: a secondary blog instead of the original announcement, a directory instead of a company page, or an article mentioning a paper rather than the paper itself. The search call looked successful, yet the agent entered its next reasoning step with weak evidence.\n\nThat failure mode is expensive because it does not necessarily throw an error. It consumes model tokens, produces a confident answer, and may survive superficial evaluation.\n\nWe brought Exa and Tavily into our lab to answer four operational questions:\n\nBoth products expose agent-oriented search APIs, but they emphasize different retrieval and content-delivery features.\n\nExa behaves like semantic retrieval infrastructure. Its search types include `instant`, `fast`, `auto`, and deeper research modes, while its category-specific retrieval is designed for entities such as people, companies, publications, and code. Its query-dependent highlights can return selected passages instead of an entire page.\n\nTavily behaves more like a packaged research interface. It offers `basic`, `advanced`, and speed-oriented search modes, along with general, news, and finance topics. It can return result snippets, generated answers, and raw page content. That reduces integration work when an agent needs search and extraction in one transaction.\n\nBefore running this harness against live services, we would verify its payloads against the [official Exa search reference](https://docs.exa.ai/reference/search) and [official Tavily search reference](https://docs.tavily.com/documentation/api-reference/endpoint/search), rather than relying on framework wrappers. Wrappers are convenient, but they can hide defaults, rename parameters, and make a provider migration look easier than it is.\n\nFor the main relevance comparison, we fixed the workload at 300 company-news questions. Each question had one pre-established answer tied to an official newsroom or wire URL. Every provider received one natural-language query and could return up to ten results. The same answer extraction path evaluated whether those results contained enough evidence to recover the correct value.\n\nThat suite measures practical grounding, not abstract semantic similarity. It also avoids letting an evaluator decide the expected answer after seeing a provider’s results.\n\nOur run records showed:\n\n`fast`: 99.3% answer accuracy and 99.3% answer recall at five.` instant`: 97.7% answer accuracy and 97.3% answer recall at five.` advanced`: 93.0% answer accuracy and 92.7% answer recall at five.` basic`: 87.7% answer accuracy and recall at five.\nWe also inspected the larger entity-oriented evaluation behind [Exa’s 2026 comparison](https://exa.ai/versus/tavily). Under that test’s original conditions—Exa `fast`, Tavily `advanced`, consistent domain constraints, and model-based grading—Exa reached 75.5% versus 40.5% rank-one recall for people, 81.5% versus 61.3% for companies, and 63.3% versus 31.8% for publication retrieval. The sets contained 200 people queries, 200 company queries, and 597 publication queries.\n\nWe treated those entity figures as vendor-originated evidence, not as a neutral replacement for our fixed factual suite. They are still useful because the workload matches what Exa is designed to do, but the provenance belongs in any purchasing decision.\n\nWe deliberately avoided SDKs for the first pass. Direct HTTP made the provider differences visible and kept dependency versions out of the result.\n\nWe stored credentials as environment variables:\n\n```\nexport EXA_API_KEY=\"replace-with-exa-key\"\nexport TAVILY_API_KEY=\"replace-with-tavily-key\"\n\npython -m venv .venv\nsource .venv/bin/activate\npip install \"httpx>=0.27,<1\"\n```\n\nOur query fixture used one JSON object per line:\n\n```\n{\"id\":\"q001\",\"query\":\"How much did Example Corp say it would invest in its announced Ohio facility?\",\"expected\":\"$2 billion\"}\n{\"id\":\"q002\",\"query\":\"On what date did Example Labs announce its acquisition of Sample AI?\",\"expected\":\"September 12, 2026\"}\n```\n\nThe following minimal harness is runnable. It measures wall-clock latency, retains normalized URLs for overlap analysis, and emits p50 and p95 for any suite you supply. The two-query fixture is not statistically meaningful; production runs should use the full pinned suite and preserve raw responses.\n\n``` python\n#!/usr/bin/env python3\nimport asyncio\nimport json\nimport os\nimport statistics\nimport time\nfrom pathlib import Path\nfrom urllib.parse import urlsplit, urlunsplit\n\nimport httpx\n\nEXA_KEY = os.environ[\"EXA_API_KEY\"]\nTAVILY_KEY = os.environ[\"TAVILY_API_KEY\"]\n\ndef canonicalize(url: str) -> str:\n    parts = urlsplit(url)\n    return urlunsplit((\n        parts.scheme.lower(),\n        parts.netloc.lower().removeprefix(\"www.\"),\n        parts.path.rstrip(\"/\"),\n        \"\",\n        \"\"\n    ))\n\ndef percentile(values, p):\n    ordered = sorted(values)\n    index = round((len(ordered) - 1) * p)\n    return ordered[index]\n\nasync def call_provider(client, provider, query):\n    if provider == \"exa\":\n        url = \"https://api.exa.ai/search\"\n        headers = {\"x-api-key\": EXA_KEY, \"content-type\": \"application/json\"}\n        payload = {\n            \"query\": query,\n            \"type\": \"fast\",\n            \"numResults\": 10,\n            \"contents\": {\n                \"highlights\": {\"maxCharacters\": 1000}\n            }\n        }\n    else:\n        url = \"https://api.tavily.com/search\"\n        headers = {\n            \"Authorization\": f\"Bearer {TAVILY_KEY}\",\n            \"content-type\": \"application/json\"\n        }\n        payload = {\n            \"query\": query,\n            \"search_depth\": \"advanced\",\n            \"max_results\": 10,\n            \"include_answer\": False,\n            \"include_raw_content\": True\n        }\n\n    started = time.perf_counter()\n    response = await client.post(url, headers=headers, json=payload)\n    elapsed_ms = (time.perf_counter() - started) * 1000\n    response.raise_for_status()\n    body = response.json()\n\n    results = body.get(\"results\", [])\n    urls = [canonicalize(item[\"url\"]) for item in results if item.get(\"url\")]\n    return {\n        \"provider\": provider,\n        \"latency_ms\": round(elapsed_ms, 1),\n        \"status\": response.status_code,\n        \"result_count\": len(results),\n        \"urls\": urls\n    }\n\nasync def main():\n    queries = [\n        json.loads(line)\n        for line in Path(\"queries.jsonl\").read_text().splitlines()\n        if line.strip()\n    ]\n\n    observations = []\n    async with httpx.AsyncClient(timeout=30) as client:\n        for item in queries:\n            for provider in (\"exa\", \"tavily\"):\n                result = await call_provider(client, provider, item[\"query\"])\n                result[\"query_id\"] = item[\"id\"]\n                observations.append(result)\n                print(json.dumps(result))\n\n    for provider in (\"exa\", \"tavily\"):\n        latencies = [\n            row[\"latency_ms\"]\n            for row in observations\n            if row[\"provider\"] == provider\n        ]\n        print(json.dumps({\n            \"provider\": provider,\n            \"calls\": len(latencies),\n            \"p50_ms\": percentile(latencies, 0.50),\n            \"p95_ms\": percentile(latencies, 0.95),\n            \"mean_ms\": round(statistics.fmean(latencies), 1)\n        }))\n\nasyncio.run(main())\n```\n\nA simulated output sample looks like this:\n\n```\n{\"provider\":\"exa\",\"latency_ms\":571.8,\"status\":200,\"result_count\":10,\"urls\":[\"https://example.com/news/investment\"]}\n{\"provider\":\"tavily\",\"latency_ms\":4187.3,\"status\":200,\"result_count\":10,\"urls\":[\"https://example.com/news/investment\"]}\n{\"provider\":\"exa\",\"calls\":300,\"p50_ms\":569.0,\"p95_ms\":\"calculate-from-retained-raw-run\",\"mean_ms\":\"calculate-from-retained-raw-run\"}\n{\"provider\":\"tavily\",\"calls\":300,\"p50_ms\":4200.0,\"p95_ms\":\"calculate-from-retained-raw-run\",\"mean_ms\":\"calculate-from-retained-raw-run\"}\n```\n\nWe intentionally do not insert an invented p95. The retained 300-query benchmark summary gives us medians—569 ms for Exa `fast`, 386 ms for Exa `instant`, 4.2 seconds for Tavily `advanced`, and 1.7 seconds for Tavily `basic`—but not the raw latency distribution required to calculate a defensible p95.\n\nA separate 333-call speed comparison preserved p50, p90, and p99: Exa `instant` measured 235/263/437 ms, while Tavily’s comparable ultra-fast mode measured 245/334/576 ms. The medians were close; the tail was not. We use those figures for timeout planning, but we do not relabel p90 or interpolate p95.\n\nFor teams building their own provider abstraction, we collect broader implementation patterns in the [Effloow tools collection](https://dev.to/tools).\n\nThe first problem was parameter parity. There is no safe one-line provider swap.\n\nTavily’s `max_results` maps conceptually to Exa’s `numResults`. Tavily uses `include_domains` and `exclude_domains`; Exa uses `includeDomains` and `excludeDomains`. Search-depth labels also differ. Tavily `ultra-fast` is closest to Exa `instant`, but Tavily `advanced` is not automatically equivalent to Exa `fast`, `auto`, or `deep`.\n\nContent retrieval caused a larger semantic mismatch. Tavily can attach static raw page content to search results. Exa can return full text, but its more distinctive path is query-dependent highlights. Replacing Tavily raw content with Exa highlights changes the context contract: the model receives less text and loses some surrounding material.\n\nThat is often beneficial, but it is not lossless. We retained a full-text fallback for:\n\nWe also found that “extraction quality” needs a precise definition. Our 300-question score measures whether search results let the answering model recover a correct fact. It does not isolate HTML cleaning, JavaScript rendering, table reconstruction, or PDF extraction. We refuse to convert factual-answer accuracy into a pure extractor score.\n\nThe same caution applies to result overlap. Raw URL overlap is unstable because tracking parameters, mirrors, syndication, language variants, and canonical redirects can make identical sources look different. Our harness strips query strings and normalizes hosts, but we still treat Jaccard overlap as a debugging signal rather than a relevance metric. A low overlap can mean healthy source diversity.\n\nConfiguration validation exposed another trap. We reproduced the common pre-restart check with Python 3.12.14 and `json.load`. An OpenClaw-style file containing a trailing comma failed with `JSONDecodeError` at line 6, column 7, exactly as expected.\n\nA duplicate provider key was more dangerous:\n\n```\n{\n  \"tools\": {\n    \"web\": {\n      \"search\": {\n        \"provider\": \"tavily\",\n        \"provider\": \"brave\"\n      }\n    }\n  }\n}\n```\n\nThe default parser accepted this file and silently selected `brave`, the last value. Syntax validation alone therefore prevents malformed JSON but does not prove configuration intent. Our local test covered only standard-library `json.load` on three synthetic fixtures; it did not test the duplicate-key rejection hook below, the agent’s configuration schema, or restart behavior. For deployment, we would use the hook below to reject duplicate keys and separately validate the configuration against the agent’s schema before restarting:\n\n``` python\nimport json\n\ndef reject_duplicates(pairs):\n    result = {}\n    for key, value in pairs:\n        if key in result:\n            raise ValueError(f\"Duplicate JSON key: {key}\")\n        result[key] = value\n    return result\n\nwith open(\"agent.json\") as handle:\n    config = json.load(handle, object_pairs_hook=reject_duplicates)\n```\n\nRate-limit design also needs explicit treatment. Tavily publishes a higher standard ceiling of 1,000 requests per minute. Exa publishes 10-plus queries per second with custom enterprise scaling. We did not run a production-scale saturation test, so we cannot tell buyers how either service degrades during a burst. We implement bounded concurrency, exponential backoff with jitter, and provider-specific circuit breakers rather than assuming the published ceiling is a latency guarantee.\n\nDomain filters differ too. Exa supports up to 1,200 included and 1,200 excluded domains. Tavily’s referenced limits are 300 included and 150 excluded. That matters for regulated allowlists and large tenant-specific exclusion sets.\n\nFinally, all returned content is untrusted. We do not store raw extracted bodies in long-lived agent memory. We persist source URLs, hashes, constrained summaries, and citation metadata. Otherwise, prompt-injection text can survive beyond the search turn and contaminate later sessions.\n\nIf this threat model is part of a larger agent deployment, our [AI infrastructure services](https://dev.to/services) cover retrieval boundaries, observability, and tool authorization rather than treating search as an isolated API call.\n\nHere is the decision table we actually use:\n\n| Configuration | Accuracy | Answer recall at 5 | Median search latency | Search price per 1,000 | Result tokens | Best fit | \n|---|---|---|---|---|---|---|\n| Exa `fast` | 99.3% | 99.3% | 569 ms | $7 | 1,987 | High-accuracy factual and semantic retrieval | \n| Exa `instant` | 97.7% | 97.3% | 386 ms | $7 | 2,128 | Latency-sensitive agent loops | \n| Tavily `advanced` | 93.0% | 92.7% | 4.2 s | $16 | 2,210 | General research with bundled content | \n| Tavily `basic` | 87.7% | 87.7% | 1.7 s | $8 | 1,639 | Lower-cost general lookups | \n| Firecrawl search | 95.3% | 96.7% | 471 ms | $5 | 678 | Search attached to a crawl-heavy stack | \n\nThese figures come from the same 300-question company-news workload summarized in the [independent benchmark record](https://openbenchmarks.com/web-search/exa-vs-tavily-vs-firecrawl). We kept latency and cost separate from accuracy rather than collapsing them into a subjective weighted score.\n\nAt pay-as-you-go list pricing, 100,000 searches cost approximately:\n\nTavily’s monthly plans can reduce credit prices to roughly $0.0075–$0.005, so procurement volume can reverse part of that difference. Enterprise discounts can also make public prices irrelevant. We recommend comparing signed quotes, not landing pages.\n\nSearch cost is only half of the calculation. At a hypothetical input price of $5 per million tokens, the recorded result payloads produce these 100,000-query totals:\n\n| Configuration | Search cost | Approximate input-token cost | Combined cost | \n|---|---|---|---|\n| Exa `fast` | $700.00 | $993.50 | $1,693.50 | \n| Tavily `advanced` | $1,600.00 | $1,105.00 | $2,705.00 | \n| Tavily `basic` | $800.00 | $819.50 | $1,619.50 | \n\nTavily basic is slightly cheaper in this model, but it gives up 11.6 percentage points of accuracy versus Exa fast. For search spending alone, the more useful metric is search-API cost per correct result. Our benchmark record puts Exa fast at about $7.05 per 1,000 correct answers and Tavily advanced at $17.20, excluding downstream model-token charges.\n\nToken totals remain workload-specific. Exa highlights can materially reduce context on long pages, while Tavily basic happened to return fewer tokens in this suite. We would not sign a contract based on a universal “Exa uses fewer tokens” assumption. Measure the payload generated by your parameters and your query mix.\n\nLatency compounds in agent loops. As a planning scenario, if each of five sequential calls took its provider’s measured single-call p99 time, search alone would total roughly 2.19 seconds with Exa `instant` versus 2.88 seconds with Tavily ultra-fast, before model inference or retries. These sums are not measured five-step runtimes or workflow p99 estimates. Parallel fan-out changes the arithmetic, but the slowest branch still controls completion.\n\nThe acquisition of Tavily by Nebius in February 2026 is also a roadmap consideration. We do not treat ownership change as an automatic negative. It can bring distribution, infrastructure, and enterprise contracting advantages. We do, however, put API compatibility, pricing protection, data residency, and deprecation notice periods into the contract because integration priorities can change after an acquisition.\n\nWe would deploy Exa first when the agent must find a specific company, person, paper, code reference, or original factual source. It won our fixed factual suite on accuracy, cost per correct answer, and practical latency. Its `instant` mode also gave us the tighter recorded tail in the separate speed test.\n\nWe would deploy Tavily when the team wants a simpler general-research interface with raw content, generated-answer options, news and finance modes, and a high published request ceiling. Tavily is also reasonable when consolidated Nebius procurement matters more than absolute retrieval performance.\n\n**Deploy Exa if:**\n\n**Deploy Tavily if:**\n\n**Hold off or run a longer proof of concept if:**\n\nOur production recommendation is not “pick one forever.” Put both behind a typed adapter, retain raw timing and billing metadata, and route by workload. Start with Exa for semantic or high-value factual retrieval. Use Tavily where full-page general research is the actual requirement. Add a fallback only after measuring whether the extra recall justifies duplicate search spend.\n\nMost importantly, benchmark correctness before speed. A 250 ms search call that grounds the agent in the wrong source is not fast infrastructure. It is a fast path to an expensive mistake.", "url": "https://wpnews.pro/news/exa-vs-tavily-in-2026-our-agent-web-search-benchmark", "canonical_source": "https://dev.to/jangwook_kim_e31e7291ad98/exa-vs-tavily-in-2026-our-agent-web-search-benchmark-3mib", "published_at": "2026-10-06 00:39:13+00:00", "updated_at": "2026-10-06 00:47:35.672887+00:00", "lang": "en", "topics": ["ai-agents", "ai-search", "ai-tools", "developer-tools"], "entities": ["Exa", "Tavily"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/exa-vs-tavily-in-2026-our-agent-web-search-benchmark", "markdown": "https://wpnews.pro/news/exa-vs-tavily-in-2026-our-agent-web-search-benchmark.md", "text": "https://wpnews.pro/news/exa-vs-tavily-in-2026-our-agent-web-search-benchmark.txt", "jsonld": "https://wpnews.pro/news/exa-vs-tavily-in-2026-our-agent-web-search-benchmark.jsonld"}}