Yandex Introduces Sona: A Single Generative Recommender That Replaces Entire Recommendation Cascade

Most production recommenders are cascades. Candidate generators feed a pre-ranker, which feeds a heavy ranker built on hundreds of engineered features. Yandex’s Sona Technical Report describes a different design. Sona is a generative AI model that brings candidate generation and ranking into a single system, replacing the multiple stages typically used in recommendation pipelines. Yandex tested the model in a seven-day live production experiment on its smart speakers. In an online A/B test, it replaced more than 15 candidate generators, the pre-ranking stage, and the ranking stage with one served transformer.

What Problem Does Sona Solve?

Cascades split one decision across separately trained models. Each stage optimizes its own objective, and the ranker only sees what upstream stages let through. Yandex’s previous stack on the Yandex Music surface consumed hundreds of features, including signals from Argus, Yandex’s earlier recommender transformer. Sona puts candidate generation and ranking around one shared user representation. The encoder reads the listener’s history once per request. A decoder generates candidates. A Ranking Module scores them against the same encoder states. No component uses hand-engineered features. Inputs are logged event fields (track ID, artist ID, duration, likes, played time, surface flags) and learned Semantic IDs.

On Yandex smart speakers, playback can begin without the user first selecting an artist, genre, or mood. The research team describes this as a pure-recommendation setting.

How Sona’s Architecture Works

1. Semantic tokenizer

Following the Semantic ID formulation of Rajput et al., every track becomes a tuple of 3 discrete codes. A frozen multimodal LLM reads the mel-spectrogram of the first 90 seconds along with title, artists, and tags. It runs in prefill-only mode. A 4-layer refinement transformer then aligns those features with listening behavior, using InfoNCE on collaborative track pairs. Residual K-means quantizes the result into 3 codebooks of 32,000 entries each. This beat a CLMR audio baseline: Recall@1000 rose from 0.8111 to 0.8524.

2. Encoder with History Compression

Sona attends to 8,192 past events. Full attention over that length is expensive, so the encoder spends depth unevenly. The recent 2,048 events get a 7-layer self-attention stack. Older events pass through cross-attention and 1 full-history layer only. The paper reports this keeps most of the quality of full attention at about half the inference cost.

3. Decoder and Ranking Module

A 2-layer decoder emits Semantic ID tuples through constrained beam search with width 1,024. A catalog trie blocks invalid prefixes. Each tuple expands to every track sharing it. The Ranking Module, which consists of four cross-attention layers, then scores those tracks against the shared encoder memory.

Training: A Teacher That Never Ships

The Ranking Module learns from a frozen Teacher Ranker. The teacher is a 0.6B-parameter transformer, also without hand-engineered features. It is trained on a year of engagement events in 2 stages: next-item-prediction pre-training, then multi-head ranking fine-tuning. Removing pre-training dropped weighted pair accuracy from 0.6215 to 0.6153.

The team calls its distillation method Rollout Distillation. During training, the current decoder generates beam candidates. The teacher scores them, together with logged impressions. The Ranking Module regresses onto those scores with mean absolute error. The joint loss is L = L_NTP + L_rollout + L_impression. Both losses update the shared encoder. At serving time, the teacher is removed.

Training stays online. Events aggregate into sessions over a 15-minute window, feed a GPU trainer, and new weights reach serving every 10 minutes. End-to-end latency is 45 minutes at the median and 60 minutes at p99. Serving runs on NVIDIA Triton Inference Server with CUDA graphs and reaches 41% model FLOPs utilization.

Results: Online A/B Test on Live Traffic

The final experiment ran for 7 days on 15% of randomly selected users per split. Every change below is statistically significant and relative to the production control:

Active Users (primary metric): +4.53%

Total Listening Time: +6.30%

Likes: +11.42%

“Repeat” Commands: +17.99%

Deeply Engaged Users: +7.37%

These gains stack on top of improvements retained from earlier deployments. On Active Users, Sona’s uplift is 2.35x the +1.93% increment Argus previously delivered on this surface.

Sona vs OneRec vs HSTU: Feature Comparison

Sona is not the first end-to-end generative recommender in production. Kuaishou’s OneRec already serves a single encoder-decoder model. Meta’s HSTU Generative Recommenders reframed recommendation as sequential transduction over user actions in 2024. What Sona combines is a full cascade replacement, no hand-engineered features, and a distilled ranker, validated online.

FeatureSona (Yandex)OneRec (Kuaishou)HSTU GR (Meta)DomainMusic streamingShort videoLarge internet platform, multiple surfacesOne served model replaces the cascadeYes, in A/B testYes, about 25% of total QPSNo, reported as a new architecture for recommendation modelsUser inputsLogged event fields only, no hand-engineered features“Multi-scale feature engineering” pathways, including uid, age, genderUser action sequences (sequential transduction)Item output3-level Semantic IDs, 3 x 32,0003-level Semantic IDs via RQ-KmeansItem IDsRanking signalProduced from frozen 0.6B Teacher RankerRL with P-Score reward model (ECPO)HSTU ranking modelReinforcement learningNo, fully supervisedYes (ECPO)Not reportedScale reported8,192-event history, 0.6B teacher10x FLOPs of prior ranking model1.5 trillion parametersReported online gain+4.53% Active Users, +6.30% listening time, +11.42% likes+0.54% and +1.24% App Stay Time+12.4% in online A/B testsPublic code or weightsNoNot in the reportYes, GitHub

Sources: Sona, OneRec, HSTU. Online gains come from different platforms and metrics, so they are not directly comparable.

Key Takeaways

Sona replaced 15+ generators, pre-ranking, and ranking with 1 served model.

No hand-engineered features: only logged events and learned Semantic IDs.

A 0.6B teacher trains the ranker, then stays out of serving.

A/B test: +4.53% Active Users, 2.35x Argus’s prior gain.

Not deployable externally: no code or weights, not on full traffic.

FAQ

What is Yandex Sona?Sona is a generative AI model that combines candidate generation and ranking within a single system. Yandex tested it in a seven-day live production experiment on its smart speakers, where it replaced the existing multi-stage recommendation pipeline for the test group.

How is Sona different from OneRec?OneRec uses engineered user feature pathways and RL with a reward model. Sona uses only logged event fields and distills ranking from a frozen teacher.

Check out the Paper for full details.
The post Yandex Introduces Sona: A Single Generative Recommender That Replaces Entire Recommendation Cascade appeared first on MarkTechPost.