Perplexity Introduces Photon: A Rust-Based Retrieval Engine That Cuts p99 Latency From 800 ms to 65 ms Perplexity released Photon, an in-house Rust-based retrieval and ranking engine that cut p99 retrieval and ranking latency from about 800 ms to about 65 ms and now handles all production traffic for its AI-native search stack. Photon also powers a new Fast Search mode in the Perplexity Search API, priced at $1 per 1,000 requests, which scored 64.3% across 3,554 tasks on 6 benchmarks at $59.73 in estimated model-plus-search cost versus 64.0% at $187.60 for the default preset, about 68% cheaper, while internal long-tail relevance (DCG) fell from 2.45 to 2.21 and answer availability dropped 2.9 percentage points from 0.596 to 0.567. Photon runs on about 20% fewer serving machines, stores about 2.5x as much data per document, and is not open source, so it cannot be self-hosted. Perplexity has released Photon https://www.perplexity.ai/hub/blog/photon , an in-house retrieval and ranking engine written in Rust. It replaces an open-source engine Perplexity had forked for its AI-native search stack. Photon now handles retrieval and ranking for all production traffic. It also powers a new Fast Search https://docs.perplexity.ai/docs/search/fast-search mode in the Perplexity Search API. Perplexity reports single-call latency of 160 ms at p50 and 230 ms at p95. Is it deployable? Yes, as a hosted API. Set search type: "fast" on POST /search and pay $1 per 1,000 requests. Photon itself is not open source, so the engine cannot be self-hosted. Why Perplexity Replaced its Old Engine The old engine hit 3 limits as the index grew: - Tail latency: Production p99 sat near 800 ms. The dataset exceeded RAM, so mlock was not an option. Cold reads triggered major page faults that stalled queries. - Merge spikes: During disk index fusion, p99 climbed to about 1.2 s for 10 to 15 minutes. - Slow recovery: Deploying and syncing an extra cluster could take more than a week. Recovery also raised the share of partial responses. Perplexity team concluded that building from scratch was simpler and cheaper than maintaining its fork. How Photon Works A load balancer routes each request to a Photon broker. The broker fans out to a shard group and watches for timeouts. Each shard runs retrieval, initial ranking, and second-stage ranking. The broker then merges candidates and fetches key document fields. - Adaptive posting lists: Short lists sit inline within a single page. Longer lists split into blocks of fixed document ID ranges. Sparse blocks store sorted offset arrays and use galloping search. Dense blocks use bitmaps, so membership becomes a single bit lookup. - Budgeted traversal: A WAND-like algorithm splits lists into driving lists and probe lists. Cheap presence checks bound each candidate’s maximum score first. Exact term frequencies are read only when a candidate can clear the threshold. - Docblob records: Each document gets a compact record of frequencies, field masks, and positions. Terms use Elias-Fano encoding, so ranking decodes only the matched terms. Ranking a candidate needs just 1 lookup per document. - Batched async reads: Record offsets are known upfront, so disk reads go out in batches through io uring . The cache checks the whole batch first. Readers take no locks, and eviction uses CLOCK instead of a shared LRU list. - Separate build and serve: Indexers build versioned shard indexes from YTsaurus tables on dedicated nodes. A controller rotates serving groups one at a time and warms caches with replayed search-log queries. A full web index now builds in a single-digit number of hours. Interactive Explainer: Inside Photon Production Results - p99 retrieval and ranking latency fell from about 800 ms to about 65 ms. This covers Photon’s stages only. - Photon runs on about 20% fewer serving machines than the old content nodes. - It stores about 2.5x as much data per document, which Perplexity used to improve ranking quality. - Pinning the same dataset with mlock would need an estimated 4.6x the resident memory Photon uses today. - Index version switches no longer cause latency spikes. Fast Search: Speed and Cost for Agents Fast Search pairs Photon with lighter ranking tuned for agentic workflows. Perplexity tested it on 6 benchmarks: WideSearch, BrowseComp https://openai.com/index/browsecomp/ , DSQA, FRAMES, SEAL-0, and SEAL-Hard. Across 3,554 tasks, Fast scored 64.3% at $59.73 in estimated model-plus-search cost. The default preset scored 64.0% at $187.60, so Fast was about 68% cheaper. The trade-off shows up in broader search quality. On internal long-tail benchmarks, relevance DCG fell from 2.45 to 2.21. Answer availability dropped from 0.596 to 0.567, a loss of 2.9 percentage points. Perplexity recommends Fast for day-to-day agent loops and the default for hard, ambiguous queries. curl -X POST 'https://api.perplexity.ai/search' \ -H "Authorization: Bearer $PERPLEXITY API KEY" \ -H 'Content-Type: application/json' \ -d '{"query": "latest stable Rust release", "search type": "fast", "max results": 5}' On Python SDK 0.43.4 and 0.43.5, pass extra body={"search type": "fast"} per the docs https://docs.perplexity.ai/docs/search/fast-search . Fast Search vs Closest Competitors | Feature | Perplexity Fast Search | Exa Instant | Parallel Search Turbo | Tavily ultra-fast | |---|---|---|---|---| | Request parameter | search type: "fast" | type: "instant" | mode: "turbo" | search depth: "ultra-fast" | | Vendor-reported latency | 160 ms p50, 230 ms p95 blog https://www.perplexity.ai/hub/blog/photon | ~250 ms typical docs https://exa.ai/docs/search/quickstart ; sub-200 ms at launch https://exa.ai/blog/exa-instant | ~200 ms docs https://docs.parallel.ai/search/modes | No figure published; lowest-latency depth docs https://docs.tavily.com/documentation/best-practices/best-practices-search | | List price per 1K requests | $1 pricing https://docs.perplexity.ai/docs/search/fast-search | $4 for up to 10 results pricing https://exa.ai/pricing | $1 docs https://docs.parallel.ai/search/modes | 1 credit: $8 pay-as-you-go, $5 to $7.50 on plans credits https://docs.tavily.com/documentation/api-credits | | Results per request | 1 to 20 | 10 in base price, $1 per 1K per extra result | Not specified | Not specified | | Known limits | Lower relevance than default preset | Extra results billed separately | English and Japanese queries only | Lower relevance than other depths | | Launched | Sep 24, 2026 | Feb 12, 2026 | Jul 13, 2026 blog https://parallel.ai/blog/parallel-search-turbo | Jan 5, 2026 blog https://www.tavily.com/blog/how-we-built-the-fastest-web-search-in-the-world | All latency figures are vendor-reported under different setups, so they are not like-for-like. Key Takeaways - Photon is Perplexity’s Rust retrieval and ranking engine, now serving all production traffic. - Production p99 latency dropped from about 800 ms to about 65 ms. - Fast Search reports 160 ms p50 and 230 ms p95 at $1 per 1,000 requests. - Fast cut estimated agent task cost by about 68% at comparable task quality. - It trades some retrieval relevance, so keep the default preset for hard queries. Check out the technical details https://www.perplexity.ai/hub/blog/photon and Fast Search docs https://docs.perplexity.ai/docs/search/fast-search . All credit goes to the researcher of this project. Also, feel free to follow us on Twitter https://x.com/intent/follow?screen name=marktechpost and don’t forget to join our 150k+ML SubReddit https://www.reddit.com/r/machinelearningnews/ and Subscribe to our Newsletter https://magic.beehiiv.com/v1/f5e63dd4-5653-4f09-83e2-321a8b1ba526?email={{email}} . Wait are you on telegram? now you can join us on telegram as well. https://t.me/machinelearningresearchnews Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? Connect with us https://forms.gle/MJjjVDPS7whH8Ngs6 Asif Razzaq is the CEO of Marktechpost AI Media Inc.. As a visionary entrepreneur and engineer, Asif is committed to harnessing the potential of Artificial Intelligence for social good. His most recent endeavor is the launch of an Artificial Intelligence Media Platform, Marktechpost, which stands out for its in-depth coverage of machine learning and deep learning news that is both technically sound and easily understandable by a wide audience. The platform boasts of over 2 million monthly views, illustrating its popularity among audiences.