Comparing Markdown Search Results: SerpApi vs. Exa, Tavily, and Firecrawl SerpApi published a comparison of markdown search output across SerpApi, Tavily, Exa, and Firecrawl, running three queries ("coffee", "how to make delicious coffee", and "grok 4.7") through all four providers with generated answers disabled. SerpApi's Google Search markdown output returned the full results page in 13,025 tokens, the smallest response a model could answer from, versus 90,234 tokens for Tavily basic with markdown raw content, 133,343 for Exa text mode, and 130,041 for Firecrawl search plus scrape. Token counts were measured with tiktoken's o200k_base encoding on the full saved response. Tavily, Exa, Firecrawl, and SerpApi all return markdown for LLMs, but they aren't converting the same thing. Tavily, Exa, and Firecrawl convert the pages a search found. SerpApi converts the search itself, so you get the results page from the real search engine as clean markdown. Markdown Output https://serpapi.com/markdown-output works the same way across all of our 100+ search APIs. This blog runs the same three queries through all four providers, using SerpApi's Google Search API https://serpapi.com/search-api as the baseline. It compares where the markdown lives in each response, what it costs in tokens, and what a model can answer from it. By the end, you'll know which output fits your agent: the results page, a list of links, or the pages behind them. We've already compared positioning and pricing in SerpApi vs. Firecrawl https://serpapi.com/blog/serpapi-vs-firecrawl/ , SerpApi vs. Exa https://serpapi.com/blog/serpapi-vs-exa-ai/ , and SerpApi Alternatives: Best Web Search APIs https://serpapi.com/blog/serpapi-alternatives-best-web-search-apis/ . This one is only about the output. How We Compared Every provider got the same query on its playground defaults, on a free plan, with generated answers turned off. The only settings we changed are the content modes named in each table and the location. We set the location as close to Austin, Texas, as each provider allows, which meant the United States for Tavily and Exa, and Austin for Firecrawl. The SerpApi baseline is Google Search https://serpapi.com/search-api with location=Austin, Texas, United States , gl=us , and hl=en , the same defaults as the playground https://serpapi.com/playground . We also set no cache=true https://serpapi.com/search-api api-parameters-serpapi-parameters-no-cache so every SerpApi search was a live fetch. We ran three queries, each one testing something different: | Query | What it tests | |---|---| | coffee | SERP features. A one-word query where Google shows most of them at once | | how to make delicious coffee | Semantic question. The kind of question an agent forwards from a user | | grok 4.7 | Freshness. Announced 45 minutes before capture, so it shows who searches live and who reads from an index | Token counts use tiktoken https://github.com/openai/tiktoken with the o200k base encoding on the full saved response, JSON or markdown, nothing stripped. Each provider section shows a trimmed copy of its response. Test 1: SERP Features Here is every mode side by side for coffee. | Provider | Mode | Tokens | What the tokens buy | |---|---|---|---| | SerpApi | Google Search, JSON | 66,237 | Full SERP as structured objects | | SerpApi | Google Search, markdown | 13,025 | Full SERP as a document | | Tavily | basic , snippets only | 5,283 | 10 titles, URLs, snippets | | Tavily | basic , include raw content=markdown | 90,234 | 8 titles plus 7 page bodies | | Tavily | advanced , include raw content=markdown | 104,792 | 10 titles plus 10 page bodies | | Exa | bare | 701 | 10 titles and URLs | | Exa | highlights | 14,172 | 10 titles plus extracted passages | | Exa | text | 133,343 | 10 titles plus 10 page bodies | | Firecrawl | search only | 8,747 | 10 titles, URLs, descriptions | | Firecrawl | search + scrape, markdown | 130,041 | 10 titles plus 9 page bodies | Two things to read from this table. The cheapest rows carry almost nothing, 10 links and at most a snippet each. Among the responses a model can actually answer from, SerpApi markdown is the smallest at 13,025 tokens. The sections below show one of these responses from each provider. SerpApi SerpApi scrapes the Google results page in real time and returns it as JSON or as markdown. The Markdown Output page https://serpapi.com/markdown-output covers how to use it, and our video on supercharging AI agents with real-time data https://www.youtube.com/watch?v=beqoAVXz7iA shows it inside an agent loop. Here we only look at the coffee response. The markdown is not a dump of the JSON. It is the results page curated for an LLM, with one section per SERP feature in the order Google showed them. The local pack https://serpapi.com/local-pack and organic results https://serpapi.com/organic-results become tables with links, and the knowledge graph https://serpapi.com/knowledge-graph becomes a list of its facts. Nothing inside the linked pages is fetched; what you get is what a person in Austin sees when they type coffee. This is the opening of the Google Search markdown document from the table: --- search metadata: id: 6aa9bc62b9dbc418dd7b473f status: Success json endpoint: "https://serpapi.com/searches/.../6aa9bc62b9dbc418dd7b473f.json" markdown endpoint: "https://serpapi.com/searches/.../6aa9bc62b9dbc418dd7b473f.md" created at: "2026-09-15T21:45:06.935Z" search parameters: engine: google q: Coffee location requested: Austin, Texas, United States hl: en gl: us --- Search Information - Query Displayed: Coffee - Total Results: 214 - Organic Results State: Results for exact spelling Local Map Image https://serpapi.com/searches/6aa9bc62b9dbc418dd7b473f/images/M0vMWN9e2SQWmXNdUKaKlQ.png https://www.google.com/search?q=Coffee&...&udm=1 Local Results Places | Position | Title | Type | Rating | Reviews | Price | Description | Thumbnail | Thumbnail Large | Links Phone | Links Website | Place Id | Gps Coordinates Latitude | Gps Coordinates Longitude | Address | Hours | Phone | | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | | 1 | Houndstooth Coffee | Coffee shop | 4.5 | 1210 | $1–10 | "Excellent coffee, offer singe source and blends, the options always rotate" | Thumbnail https://serpapi.com/searches/6aa9bc62b9dbc418dd7b473f/images/-WgD2tu7k8gkCk2xfEgTRbF9GSQB9Snw18rTTBABMPE7QjI8PtOIbqd0wV7X81n3.jpeg | https://lh3.googleusercontent.com/gps-cs-s/AHRPTWnWTvhbq7JcIcVUxEmUhGnD40dIdSQbygpfH9KvFfmvSN9Hix8E60n3riRiG648DFrJz7Z8raNZfxcVWY5UbcOHkCyE56MPZoVV2rslKvPAVcTTZkyfbbVHBGYpciBgoKjzCcrVFA=w1000-h1000-c-n | tel: 512 394-6051 | http://www.houndstoothcoffee.com/?utm source=google&utm medium=organic&utm campaign=googlebusiness | 11265938073076301333 | 30.266216 | -97.743066 | 401 Congress Ave. 100c, Austin, TX 78701 | Open · Closes 7 PM | 512 394-6051 | < -- 2 more places, then 8 more sections, 13,025 tokens in total -- - This is the results page: SerpApi returns what Google ranked, grouped, and labeled. The local pack, the knowledge graph, and the AI Overview exist here and nowhere else in this comparison. - A fifth of the JSON: Markdown drops the JSON brackets and keys and keeps the text a model reads, which is the reason the feature exists. - Trim what you don't need: The JSON Restrictor https://serpapi.com/json-restrictor keeps only the sections you ask for, and it also works with Markdown output. For example, json restrictor=organic results keeps only the organic results. - Every search has an ID and a URL: search metadata.id and markdown endpoint are in the frontmatter. Open the URL later and you get the same document back. Tavily Tavily is a search API built for agents. A basic search for coffee returns 10 results with a title, a URL, a snippet in content , and a relevance score. There is no markdown yet. The markdown appears when you set include raw content to markdown . It lands inside each result as raw content , the converted body of that result’s page. The response stays JSON and the markdown is a string field inside it. This is the first result of the basic search with include raw content=markdown : { "query": "coffee", "results": { "url": "https://en.wikipedia.org/wiki/Coffee", "title": "Coffee", "content": "Coffee is a beverage brewed from roasted ground coffee beans. Darkly colored and bitter, coffee has a stimulating effect on humans due to its caffeine content, ... ", "score": 0.83742094, "raw content": " Jump to content bodyContent \n\n /static/images/icons/enwiki-25.svg Wikipedia /static/images/mobile/copyright/wikipedia-wordmark-en-25.svg The Free Encyclopedia /static/images/mobile/copyright/wikipedia-tagline-en-25.svg https://serpapi.com/wiki/Main Page \n\n Search /wiki/Special:Search \"Search Wikipedia f \" \n\n Contents\n\n 1 Etymology Etymology \n 2 History History \n + 2.1 Legendary accounts and myths Legendary accounts and myths ", // raw content continues for 242,608 characters "favicon": "https://en.wikipedia.org/static/apple-touch/wikipedia.png", "id": "fe77ba-00" } // 7 more results, raw content present on 7 of 8 , "response time": 1.43, "request id": "6e4056c2-7185-4d66-ba80-44fe823d7392" } - raw content is the page, not the SERP: For coffee the first result is Wikipedia, and raw content starts with Wikipedia’s navigation links and table of contents. This response has no local pack, knowledge graph, or People Also Ask. - The pages are the cost: The snippets-only response is 5,283 tokens. The extra tokens in the other Tavily rows are the page bodies it adds. Exa Exa searches its own neural index of the web. The bare search returns 10 results with id , title , url , publishedDate on some, and no snippet at all. Page content comes through two opt-in fields. The highlights field is an array of passages Exa extracted from each page. The text field is the full page body, which Exa returns as markdown-style text by default. This is the first result of the text search: { "requestId": "6e0760b35bd59e7fd1699bc92570c7ff", "results": { "id": "https://en.wikipedia.org/wiki/Coffee", "title": "Coffee", "url": "https://en.wikipedia.org/wiki/Coffee", "text": "Coffee\n\nCoffee\n\n| Latte and black filtered coffee | |\n| --- | --- |\n| Type | Usually hot; can be iced |\n| Origin | Yemen 1 2 3 |\n| Introduced | 15th century |\n| Color | Black, dark brown, light brown, beige |\n| Flavor | Distinctive, somewhat bitter |\n| Ingredients | Roasted coffee beans |\n| Standard drinkware | Mug |\n\nCoffee is a beverage brewed from roasted ground coffee beans.", // text continues for 102,126 characters "image": "https://upload.wikimedia.org/wikipedia/commons/thumb/1/19/Coffee beans unroasted.jpg/250px-Coffee beans unroasted.jpg" } // 9 more results , "searchTime": 880.4, "costDollars": { "total": 0.007, "search": { "neural": 0.007 } } } - text is the page, not the SERP: The Wikipedia result keeps its infobox as a markdown table, but nothing from Google's results page is in it. - text is the largest response in this comparison: Exa inlines the full body of all 10 pages. Firecrawl Firecrawl is a page scraper with a /search endpoint on top. With scraping off, the search returns 10 web results with url , title , and description . That is 8,747 tokens and no markdown. With “scrape content from search results” on and the format set to markdown , each result gains a markdown field holding the main content of the page. It also gets a metadata object with the page’s own title, description, and status code. This is the first result of search + scrape with markdown format: { "web": { "url": "https://en.wikipedia.org/wiki/Coffee", "title": "Coffee - Wikipedia", "description": "For other uses, see Coffee disambiguation https://en.wikipedia.org/wiki/Coffee disambiguation .", "position": 1, "markdown": " Jump to content https://en.wikipedia.org/wiki/Coffee bodyContent \n\n Page semi-protected https://thumb.wikimedia.org/wikipedia/en/thumb/1/1b/Semi-protection-shackle.svg/20px-Semi-protection-shackle.svg.png https://en.wikipedia.org/wiki/Wikipedia:Protection policy semi \"This article is semi-protected.\" \n\nFrom Wikipedia, the free encyclopedia\n\nBrewed beverage\n\nThis article is about the beverage. For other uses, see Coffee disambiguation https://en.wikipedia.org/wiki/Coffee disambiguation \"Coffee disambiguation \" .", // markdown continues for 249,433 characters "metadata": { "title": "Coffee - Wikipedia", "sourceURL": "https://en.wikipedia.org/wiki/Coffee", "statusCode": 200, "cachedAt": "2026-09-15T15:48:22.840Z", "cacheState": "hit", "creditsUsed": 1 // 19 more keys, including og: tags copied from the page head } } // 9 more results, markdown present on 9 of 10 , "creditsUsed": 11 } - Pages can come from a cache: The Wikipedia page above has cacheState: hit , and its cachedAt shows Firecrawl stored it about seven hours before our request. For coffee , 6 of the 9 pages with markdown came from that cache. - Search plus scrape is the expensive path: It costs about 10 times the tokens of the SerpApi markdown for the same query, because the response includes 9 full pages. Test 2: Semantic Question Now let’s test a question a user would ask an agent, how to make delicious coffee. Google shows a different page for it. The local pack and knowledge graph are gone, and in their place are recipes and videos. The related questions and the AI Overview are still there. The other three providers still return the top pages. Tavily, Exa, and Firecrawl return the same kind of response we saw for coffee, page bodies inside JSON, so here we only show SerpApi. The part to look at is the AI Overview https://serpapi.com/ai-overview , Google’s own answer to the question. It comes in the markdown with its sources: Ai Overview Text Blocks 1. To make tasty coffee at home, start with freshly ground whole beans, a precise coffee-to-water ratio, and filtered water heated to the optimal temperature. 2. Essential Rules for Great Coffee < -- 7 more text blocks -- References | Source | Index | Title | Snippet | Thumbnail | | --- | --- | --- | --- | --- | | thekitchn.com | 2 | https://www.thekitchn.com/best-coffee-brewing-method-22977756 https://www.thekitchn.com/best-coffee-brewing-method-22977756 | | | | coffeechronicler.com | 3 | https://coffeechronicler.com/best-way-to-make-coffee-at-home/ https://coffeechronicler.com/best-way-to-make-coffee-at-home/ | | | < -- 8 references in total, then Organic Results, Related Searches, Pagination -- An agent can quote this answer or open one of its sources. The other providers can’t return it, because the AI Overview only exists on Google's results page, not on any of the pages they fetch. For a how-to question, SerpApi returns Google's answer and its sources in 11,266 tokens. The next smallest full-page response was Exa's, at 32,772 tokens. Test 3: Freshness The grok 4.7 query is the freshness test. xAI announced it on X, and we ran the query 45 minutes later. By then Google had indexed the announcement, but most of the web had not. A provider searching live should return the release. A provider reading from an index returns whatever it crawled last. SerpApi's response is about a third the size of the others, but the bigger difference is the content. SerpApi returned Google's results page with the xAI announcement first and the launch coverage in Top Stories https://serpapi.com/top-stories . All 10 of Tavily's results were written before the launch, so a model reading them would answer that Grok 4.7 is not out yet. Exa ranked the announcement and Vercel's changelog first, then filled the rest with pre-launch speculation. Firecrawl found three posts about the release, and only the two Cursor forum threads came back with content. For breaking news, SerpApi returned the most coverage of the release. Eight of the 11 results across its organic results and Top Stories were about the launch. Results Across All Three Queries On every query, SerpApi's markdown was the smallest response a model could answer from. It used 10,209 to 13,025 tokens, while the full-page modes in the chart used 30,802 to 133,343. The gap is widest on coffee, where all three full-page responses include the Wikipedia article on coffee. Location and Traceability The tests above compared what each provider returned. Two differences don't show up in a single response, but they matter once an agent runs in production. Customizing the Search Each SerpApi engine exposes the main parameters its search engine offers, and the markdown follows them. For Google Search that means you can pin a query to what a specific user would see. | Parameter | What it controls | |---|---| | location | City, state, or neighborhood the search comes from | | gl | Country of the results | | hl | Interface language | | google domain | Which Google domain to query, such as google.co.uk | | device | Desktop, tablet, or mobile layout | The other three providers accept at best a country or a city. Each engine lists its parameters in its docs, for example the Google Search API docs https://serpapi.com/search-api . Trace Every Search Every SerpApi response, JSON or markdown, opens with search metadata . It has the search id, a json endpoint , a markdown endpoint , and the created at timestamp. The Search Archive API https://serpapi.com/search-archive-api keeps each search for 31 days, so the same id returns the same results page later, in either format. Every SerpApi excerpt in this post came from an archived search. For an agent, a log line with a search id is a record of exactly what the model read. If a user asks why the model said something, you can open the search and check. Tavily returns a request id and Exa a requestId , and neither can fetch that search again. Firecrawl returns a search id, plus a scrapeId and cacheState for each page, but its API documents no way to fetch the search again. With those three, the only record is the copy you kept. When to Use Which The question is whether your agent needs the search or the pages behind it. Tavily, Exa, and Firecrawl aren't doing anything wrong. Returning pages is the product, and for a research agent that needs to read every source, it is the right product. The token counts above are the size of those pages, not a flaw. The table below shows each case in more detail, with the reason behind each pick: | You need | Pick | Because | |---|---|---| | What Google shows for a query, including local pack, knowledge graph, and AI Overview | SerpApi markdown | The only response here with those sections, and the smallest one a model can answer from | | Results for something that happened this hour | SerpApi markdown | A live Google search, with Google's relative dates on most rows | | Results as a user in a specific city, country, or device would see them | SerpApi markdown | location , gl , hl , and device pin the search down to the neighborhood | | A record of what the model read | SerpApi markdown | Every search keeps an id and a URL for 31 days | | Fields to parse in code, such as ratings, prices, or positions | SerpApi JSON | Typed fields, no markdown to parse | | A short list of URLs to decide what to fetch next | Tavily basic , Exa bare, Firecrawl search only | Under 9,000 tokens each. Nothing to answer from, but nothing wasted either | | The full body of the pages behind a query | Tavily include raw content , Exa text , Firecrawl search + scrape | They return the pages. Expect 30,000 to 135,000 tokens, and clean the navigation yourself. | Conclusion Any agent that needs to know what the web says right now, where a user would see it, and exactly what the model read needs the search before it needs the pages. SerpApi's Markdown Output returns that search as one document, from Google Search or any of our other 100+ search APIs.