Perplexity has released Photon, an in-house retrieval and ranking engine written in Rust. It replaces an open-source engine Perplexity had forked for its AI-native search stack. Photon now handles retrieval and ranking for all production traffic. It also powers a new Fast Search mode in the Perplexity Search API. Perplexity reports single-call latency of 160 ms at p50 and 230 ms at p95.
Is it deployable? Yes, as a hosted API. Set search_type: “fast” on POST /search and pay $1 per 1,000 requests. Photon itself is not open source, so the engine cannot be self-hosted.
Why Perplexity Replaced its Old Engine
The old engine hit 3 limits as the index grew:
Tail latency: Production p99 sat near 800 ms. The dataset exceeded RAM, so mlock was not an option. Cold reads triggered major page faults that stalled queries.
Merge spikes: During disk index fusion, p99 climbed to about 1.2 s for 10 to 15 minutes.
Slow recovery: Deploying and syncing an extra cluster could take more than a week. Recovery also raised the share of partial responses.
Perplexity team concluded that building from scratch was simpler and cheaper than maintaining its fork.
How Photon Works
A load balancer routes each request to a Photon broker. The broker fans out to a shard group and watches for timeouts. Each shard runs retrieval, initial ranking, and second-stage ranking. The broker then merges candidates and fetches key document fields.
Adaptive posting lists: Short lists sit inline within a single page. Longer lists split into blocks of fixed document ID ranges. Sparse blocks store sorted offset arrays and use galloping search. Dense blocks use bitmaps, so membership becomes a single bit lookup.
Budgeted traversal: A WAND-like algorithm splits lists into driving lists and probe lists. Cheap presence checks bound each candidate’s maximum score first. Exact term frequencies are read only when a candidate can clear the threshold.
Docblob records: Each document gets a compact record of frequencies, field masks, and positions. Terms use Elias-Fano encoding, so ranking decodes only the matched terms. Ranking a candidate needs just 1 lookup per document.
Batched async reads: Record offsets are known upfront, so disk reads go out in batches through io_uring. The cache checks the whole batch first. Readers take no locks, and eviction uses CLOCK instead of a shared LRU list.
Separate build and serve: Indexers build versioned shard indexes from YTsaurus tables on dedicated nodes. A controller rotates serving groups one at a time and warms caches with replayed search-log queries.
A full web index now builds in a single-digit number of hours.
Interactive Explainer: Inside Photon
