# Comparing Markdown Search Results: SerpApi vs. Exa, Tavily, and Firecrawl

> Source: <https://serpapi.com/blog/markdown-search-results-serpapi-vs-exa-tavily-firecrawl/>
> Published: 2026-10-01 19:17:31+00:00

Tavily, Exa, Firecrawl, and SerpApi all return markdown for LLMs, but they aren't converting the same thing. Tavily, Exa, and Firecrawl convert the pages a search found. SerpApi converts the search itself, so you get the results page from the real search engine as clean markdown. [__Markdown Output__](https://serpapi.com/markdown-output) works the same way across all of our 100+ search APIs.

This blog runs the same three queries through all four providers, using SerpApi's [__Google Search API__](https://serpapi.com/search-api) as the baseline. It compares where the markdown lives in each response, what it costs in tokens, and what a model can answer from it. By the end, you'll know which output fits your agent: the results page, a list of links, or the pages behind them.

We've already compared positioning and pricing in [__SerpApi vs. Firecrawl__](https://serpapi.com/blog/serpapi-vs-firecrawl/), [__SerpApi vs. Exa__](https://serpapi.com/blog/serpapi-vs-exa-ai/), and [__SerpApi Alternatives: Best Web Search APIs__](https://serpapi.com/blog/serpapi-alternatives-best-web-search-apis/). This one is only about the output.

## How We Compared

Every provider got the same query on its playground defaults, on a free plan, with generated answers turned off. The only settings we changed are the content modes named in each table and the location. We set the location as close to Austin, Texas, as each provider allows, which meant the United States for Tavily and Exa, and Austin for Firecrawl. The SerpApi baseline is [__Google Search__](https://serpapi.com/search-api) with `location=Austin, Texas, United States`, `gl=us`, and `hl=en`, the same defaults as the [__playground__](https://serpapi.com/playground). We also set [__no_cache=true__](https://serpapi.com/search-api#api-parameters-serpapi-parameters-no-cache) so every SerpApi search was a live fetch.

We ran three queries, each one testing something different:

| Query | What it tests | 
|---|---|
| `coffee` | SERP features. A one-word query where Google shows most of them at once | 
| `how to make delicious coffee` | Semantic question. The kind of question an agent forwards from a user | 
| `grok 4.7` | Freshness. Announced 45 minutes before capture, so it shows who searches live and who reads from an index | 

Token counts use [__tiktoken__](https://github.com/openai/tiktoken) with the `o200k_base` encoding on the full saved response, JSON or markdown, nothing stripped. Each provider section shows a trimmed copy of its response.

## Test 1: SERP Features

Here is every mode side by side for coffee.

| Provider | Mode | Tokens | What the tokens buy | 
|---|---|---|---|
| SerpApi | Google Search, JSON | 66,237 | Full SERP as structured objects | 
| SerpApi | Google Search, markdown | 13,025 | Full SERP as a document | 
| Tavily | `basic` , snippets only | 5,283 | 10 titles, URLs, snippets | 
| Tavily | `basic` ,`include_raw_content=markdown` | 90,234 | 8 titles plus 7 page bodies | 
| Tavily | `advanced` ,`include_raw_content=markdown` | 104,792 | 10 titles plus 10 page bodies | 
| Exa | bare | 701 | 10 titles and URLs | 
| Exa | `highlights` | 14,172 | 10 titles plus extracted passages | 
| Exa | `text` | 133,343 | 10 titles plus 10 page bodies | 
| Firecrawl | search only | 8,747 | 10 titles, URLs, descriptions | 
| Firecrawl | search + scrape, markdown | 130,041 | 10 titles plus 9 page bodies | 

Two things to read from this table. The cheapest rows carry almost nothing, 10 links and at most a snippet each. Among the responses a model can actually answer from, SerpApi markdown is the smallest at 13,025 tokens. The sections below show one of these responses from each provider.

### SerpApi

SerpApi scrapes the Google results page in real time and returns it as JSON or as markdown. The [__Markdown Output page__](https://serpapi.com/markdown-output) covers how to use it, and our video on [__supercharging AI agents with real-time data__](https://www.youtube.com/watch?v=beqoAVXz7iA) shows it inside an agent loop. Here we only look at the coffee response.

The markdown is not a dump of the JSON. It is the results page curated for an LLM, with one section per SERP feature in the order Google showed them. The [__local pack__](https://serpapi.com/local-pack) and [__organic results__](https://serpapi.com/organic-results) become tables with links, and the [__knowledge graph__](https://serpapi.com/knowledge-graph) becomes a list of its facts. Nothing inside the linked pages is fetched; what you get is what a person in Austin sees when they type coffee.

This is the opening of the Google Search markdown document from the table:

```
---
search_metadata:
  id: 6aa9bc62b9dbc418dd7b473f
  status: Success
  json_endpoint: "https://serpapi.com/searches/.../6aa9bc62b9dbc418dd7b473f.json"
  markdown_endpoint: "https://serpapi.com/searches/.../6aa9bc62b9dbc418dd7b473f.md"
  created_at: "2026-09-15T21:45:06.935Z"
search_parameters:
  engine: google
  q: Coffee
  location_requested: Austin, Texas, United States
  hl: en
  gl: us
---

## Search Information

- Query Displayed: Coffee
- Total Results: 214
- Organic Results State: Results for exact spelling

## Local Map

[![Image](https://serpapi.com/searches/6aa9bc62b9dbc418dd7b473f/images/M0vMWN9e2SQWmXNdUKaKlQ.png)](https://www.google.com/search?q=Coffee&...&udm=1)

## Local Results

### Places

| Position | Title | Type | Rating | Reviews | Price | Description | Thumbnail | Thumbnail Large | Links Phone | Links Website | Place Id | Gps Coordinates Latitude | Gps Coordinates Longitude | Address | Hours | Phone |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| 1 | Houndstooth Coffee | Coffee shop | 4.5 | 1210 | $1–10 | "Excellent coffee, offer singe source and blends, the options always rotate" | ![Thumbnail](https://serpapi.com/searches/6aa9bc62b9dbc418dd7b473f/images/-WgD2tu7k8gkCk2xfEgTRbF9GSQB9Snw18rTTBABMPE7QjI8PtOIbqd0wV7X81n3.jpeg) | https://lh3.googleusercontent.com/gps-cs-s/AHRPTWnWTvhbq7JcIcVUxEmUhGnD40dIdSQbygpfH9KvFfmvSN9Hix8E60n3riRiG648DFrJz7Z8raNZfxcVWY5UbcOHkCyE56MPZoVV2rslKvPAVcTTZkyfbbVHBGYpciBgoKjzCcrVFA=w1000-h1000-c-n | tel:(512) 394-6051 | http://www.houndstoothcoffee.com/?utm_source=google&utm_medium=organic&utm_campaign=googlebusiness | 11265938073076301333 | 30.266216 | -97.743066 | 401 Congress Ave. #100c, Austin, TX 78701 | Open · Closes 7 PM | (512) 394-6051 |

<!-- 2 more places, then 8 more sections, 13,025 tokens in total -->
```

- **This is the results page:** SerpApi returns what Google ranked, grouped, and labeled. The local pack, the knowledge graph, and the AI Overview exist here and nowhere else in this comparison.
- **A fifth of the JSON:** Markdown drops the JSON brackets and keys and keeps the text a model reads, which is the reason the feature exists.
- **Trim what you don't need:** The[__JSON Restrictor__](https://serpapi.com/json-restrictor) keeps only the sections you ask for, and it also works with Markdown output. For example,`json_restrictor=organic_results` keeps only the organic results.
- **Every search has an ID and a URL:**`search_metadata.id` and`markdown_endpoint` are in the frontmatter. Open the URL later and you get the same document back.

### Tavily

Tavily is a search API built for agents. A basic search for coffee returns 10 results with a title, a URL, a snippet in `content`, and a relevance score. There is no markdown yet. The markdown appears when you set `include_raw_content` to `markdown`. It lands inside each result as `raw_content`, the converted body of that result’s page. The response stays JSON and the markdown is a string field inside it.

This is the first result of the basic search with `include_raw_content=markdown`:

```
{
  "query": "coffee",
  "results": [
    {
      "url": "https://en.wikipedia.org/wiki/Coffee",
      "title": "Coffee",
      "content": "Coffee is a beverage brewed from roasted ground coffee beans. Darkly colored and bitter, coffee has a stimulating effect on humans due to its caffeine content, [...]",
      "score": 0.83742094,
      "raw_content": "[Jump to content](#bodyContent)\n\n[![](/static/images/icons/enwiki-25.svg)  ![Wikipedia](/static/images/mobile/copyright/wikipedia-wordmark-en-25.svg) ![The Free Encyclopedia](/static/images/mobile/copyright/wikipedia-tagline-en-25.svg)](https://serpapi.com/wiki/Main_Page)\n\n[Search](/wiki/Special:Search \"Search Wikipedia [f]\")\n\n## Contents\n\n* [1 Etymology](#Etymology)\n* [2 History](#History)\n  + [2.1 Legendary accounts and myths](#Legendary_accounts_and_myths)",
      // raw_content continues for 242,608 characters
      "favicon": "https://en.wikipedia.org/static/apple-touch/wikipedia.png",
      "id": "fe77ba-00"
    }
    // 7 more results, raw_content present on 7 of 8
  ],
  "response_time": 1.43,
  "request_id": "6e4056c2-7185-4d66-ba80-44fe823d7392"
}
```

- **raw_content is the page, not the SERP:** For coffee the first result is Wikipedia, and raw_content starts with Wikipedia’s navigation links and table of contents. This response has no local pack, knowledge graph, or People Also Ask.
- **The pages are the cost:** The snippets-only response is 5,283 tokens. The extra tokens in the other Tavily rows are the page bodies it adds.

### Exa

Exa searches its own neural index of the web. The bare search returns 10 results with `id`, `title`, `url`, `publishedDate` on some, and no snippet at all. Page content comes through two opt-in fields. The `highlights` field is an array of passages Exa extracted from each page. The `text` field is the full page body, which Exa returns as markdown-style text by default.

This is the first result of the text search:

```
{
  "requestId": "6e0760b35bd59e7fd1699bc92570c7ff",
  "results": [
    {
      "id": "https://en.wikipedia.org/wiki/Coffee",
      "title": "Coffee",
      "url": "https://en.wikipedia.org/wiki/Coffee",
      "text": "Coffee\n\nCoffee\n\n| Latte and black filtered coffee | |\n| --- | --- |\n| Type | Usually hot; can be iced |\n| Origin | Yemen [1] [2] [3] |\n| Introduced | 15th century |\n| Color | Black, dark brown, light brown, beige |\n| Flavor | Distinctive, somewhat bitter |\n| Ingredients | Roasted coffee beans |\n| Standard drinkware | Mug |\n\nCoffee is a beverage brewed from roasted ground coffee beans.",
      // text continues for 102,126 characters
      "image": "https://upload.wikimedia.org/wikipedia/commons/thumb/1/19/Coffee_beans_unroasted.jpg/250px-Coffee_beans_unroasted.jpg"
    }
    // 9 more results
  ],
  "searchTime": 880.4,
  "costDollars": {
    "total": 0.007,
    "search": { "neural": 0.007 }
  }
}
```

- **text is the page, not the SERP:** The Wikipedia result keeps its infobox as a markdown table, but nothing from Google's results page is in it.
- **text is the largest response in this comparison:** Exa inlines the full body of all 10 pages.

### Firecrawl

Firecrawl is a page scraper with a `/search` endpoint on top. With scraping off, the search returns 10 web results with `url`, `title`, and `description`. That is 8,747 tokens and no markdown. With “scrape content from search results” on and the format set to `markdown`, each result gains a markdown field holding the main content of the page. It also gets a metadata object with the page’s own title, description, and status code.

This is the first result of search + scrape with markdown format:

```
{
  "web": [
    {
      "url": "https://en.wikipedia.org/wiki/Coffee",
      "title": "Coffee - Wikipedia",
      "description": "For other uses, see [Coffee (disambiguation)](https://en.wikipedia.org/wiki/Coffee_(disambiguation)).",
      "position": 1,
      "markdown": "[Jump to content](https://en.wikipedia.org/wiki/Coffee#bodyContent)\n\n[![Page semi-protected](https://thumb.wikimedia.org/wikipedia/en/thumb/1/1b/Semi-protection-shackle.svg/20px-Semi-protection-shackle.svg.png)](https://en.wikipedia.org/wiki/Wikipedia:Protection_policy#semi \"This article is semi-protected.\")\n\nFrom Wikipedia, the free encyclopedia\n\nBrewed beverage\n\nThis article is about the beverage. For other uses, see [Coffee (disambiguation)](https://en.wikipedia.org/wiki/Coffee_(disambiguation) \"Coffee (disambiguation)\").",
      // markdown continues for 249,433 characters
      "metadata": {
        "title": "Coffee - Wikipedia",
        "sourceURL": "https://en.wikipedia.org/wiki/Coffee",
        "statusCode": 200,
        "cachedAt": "2026-09-15T15:48:22.840Z",
        "cacheState": "hit",
        "creditsUsed": 1
        // 19 more keys, including og:* tags copied from the page head
      }
    }
    // 9 more results, markdown present on 9 of 10
  ],
  "creditsUsed": 11
}
```

- **Pages can come from a cache:** The Wikipedia page above has`cacheState: hit` , and its`cachedAt` shows Firecrawl stored it about seven hours before our request. For`coffee` , 6 of the 9 pages with markdown came from that cache.
- **Search plus scrape is the expensive path:** It costs about 10 times the tokens of the SerpApi markdown for the same query, because the response includes 9 full pages.

## Test 2: Semantic Question

Now let’s test a question a user would ask an agent, how to make delicious coffee. Google shows a different page for it. The local pack and knowledge graph are gone, and in their place are recipes and videos. The related questions and the AI Overview are still there. The other three providers still return the top pages.

Tavily, Exa, and Firecrawl return the same kind of response we saw for coffee, page bodies inside JSON, so here we only show SerpApi. The part to look at is the [__AI Overview__](https://serpapi.com/ai-overview), Google’s own answer to the question. It comes in the markdown with its sources:

```
## Ai Overview

### Text Blocks

1. To make tasty coffee at home, start with freshly ground whole beans, a precise coffee-to-water ratio, and filtered water heated to the optimal temperature.
2. Essential Rules for Great Coffee

<!-- 7 more text blocks -->

### References

| Source | Index | Title | Snippet | Thumbnail |
| --- | --- | --- | --- | --- |
| thekitchn.com | 2 | [https://www.thekitchn.com/best-coffee-brewing-method-22977756](https://www.thekitchn.com/best-coffee-brewing-method-22977756) |  |  |
| coffeechronicler.com | 3 | [https://coffeechronicler.com/best-way-to-make-coffee-at-home/](https://coffeechronicler.com/best-way-to-make-coffee-at-home/) |  |  |

<!-- 8 references in total, then Organic Results, Related Searches, Pagination -->
```

An agent can quote this answer or open one of its sources. The other providers can’t return it, because the AI Overview only exists on Google's results page, not on any of the pages they fetch.

For a how-to question, SerpApi returns Google's answer and its sources in 11,266 tokens. The next smallest full-page response was Exa's, at 32,772 tokens.

## Test 3: Freshness

The grok 4.7 query is the freshness test. xAI announced it on X, and we ran the query 45 minutes later. By then Google had indexed the announcement, but most of the web had not. A provider searching live should return the release. A provider reading from an index returns whatever it crawled last.

SerpApi's response is about a third the size of the others, but the bigger difference is the content. SerpApi returned Google's results page with the xAI announcement first and the launch coverage in [__Top Stories__](https://serpapi.com/top-stories). All 10 of Tavily's results were written before the launch, so a model reading them would answer that Grok 4.7 is not out yet. Exa ranked the announcement and Vercel's changelog first, then filled the rest with pre-launch speculation. Firecrawl found three posts about the release, and only the two Cursor forum threads came back with content.

For breaking news, SerpApi returned the most coverage of the release. Eight of the 11 results across its organic results and Top Stories were about the launch.

## Results Across All Three Queries

On every query, SerpApi's markdown was the smallest response a model could answer from. It used 10,209 to 13,025 tokens, while the full-page modes in the chart used 30,802 to 133,343. The gap is widest on coffee, where all three full-page responses include the Wikipedia article on coffee.

## Location and Traceability

The tests above compared what each provider returned. Two differences don't show up in a single response, but they matter once an agent runs in production.

### Customizing the Search

Each SerpApi engine exposes the main parameters its search engine offers, and the markdown follows them. For Google Search that means you can pin a query to what a specific user would see.

| Parameter | What it controls | 
|---|---|
| `location` | City, state, or neighborhood the search comes from | 
| `gl` | Country of the results | 
| `hl` | Interface language | 
| `google_domain` | Which Google domain to query, such as `google.co.uk` | 
| `device` | Desktop, tablet, or mobile layout | 

The other three providers accept at best a country or a city. Each engine lists its parameters in its docs, for example the [__Google Search API docs__](https://serpapi.com/search-api).

### Trace Every Search

Every SerpApi response, JSON or markdown, opens with `search_metadata`. It has the search id, a `json_endpoint`, a `markdown_endpoint`,  and the `created_at` timestamp. The [__Search Archive API__](https://serpapi.com/search-archive-api) keeps each search for 31 days, so the same `id` returns the same results page later, in either format. Every SerpApi excerpt in this post came from an archived search.

For an agent, a log line with a search id is a record of exactly what the model read. If a user asks why the model said something, you can open the search and check. Tavily returns a `request_id` and Exa a `requestId`, and neither can fetch that search again. Firecrawl returns a search id, plus a `scrapeId` and `cacheState` for each page, but its API documents no way to fetch the search again. With those three, the only record is the copy you kept.

## When to Use Which

The question is whether your agent needs the search or the pages behind it.

Tavily, Exa, and Firecrawl aren't doing anything wrong. Returning pages is the product, and for a research agent that needs to read every source, it is the right product. The token counts above are the size of those pages, not a flaw.

The table below shows each case in more detail, with the reason behind each pick:

| You need | Pick | Because | 
|---|---|---|
| What Google shows for a query, including local pack, knowledge graph, and AI Overview | SerpApi markdown | The only response here with those sections, and the smallest one a model can answer from | 
| Results for something that happened this hour | SerpApi markdown | A live Google search, with Google's relative dates on most rows | 
| Results as a user in a specific city, country, or device would see them | SerpApi markdown | `location` ,`gl` ,`hl` , and`device` pin the search down to the neighborhood | 
| A record of what the model read | SerpApi markdown | Every search keeps an `id` and a URL for 31 days | 
| Fields to parse in code, such as ratings, prices, or positions | SerpApi JSON | Typed fields, no markdown to parse | 
| A short list of URLs to decide what to fetch next | Tavily `basic` , Exa bare, Firecrawl search only | Under 9,000 tokens each. Nothing to answer from, but nothing wasted either | 
| The full body of the pages behind a query | Tavily `include_raw_content` , Exa`text` , Firecrawl search + scrape | They return the pages. Expect 30,000 to 135,000 tokens, and clean the navigation yourself. | 

## Conclusion

Any agent that needs to know what the web says right now, where a user would see it, and exactly what the model read needs the search before it needs the pages. SerpApi's Markdown Output returns that search as one document, from Google Search or any of our other 100+ search APIs.
