Sona: One Transformer Replaces Yandex's 15-Stage Music Recommender

October 5, 2026 • news

Yandex has replaced its multi-stage music recommendation cascade with Sona, a single generative transformer that unifies candidate generation, pre-ranking, and final ranking around one shared user representation. According to Yandex's technical report, covered by MarkTechPost, Sona runs as one served model, replacing more than 15 candidate generators and the subsequent ranking stages. The previous stack consumed hundreds of engineered features, including signals from Argus, Yandex's earlier recommender transformer; Sona instead uses only logged event fields and learned Semantic IDs. This pipeline consolidation matters because cascades split one decision across separately trained models, and each downstream ranker can see only what upstream stages let through.

History Compression and Semantic Tokenization

Sona eliminates hand-engineered features entirely. Item representations begin with a frozen multimodal large language model running in prefill-only mode, which ingests the first 90 seconds of a track's mel-spectrogram alongside its title, artists, and metadata tags. A four-layer refinement transformer aligns those features with listening behavior using an InfoNCE loss on collaborative track pairs. Residual K-means quantization maps the result to a tuple of three discrete codes from three codebooks of 32,000 entries each. Yandex reports that this approach improved Recall@1000 to 0.8524 from a 0.8111 CLMR audio baseline.

Sona's encoder attends to 8,192 historical events but spends depth asymmetrically. The most recent 2,048 events pass through a seven-layer self-attention stack; older events route through cross-attention into one full-history layer only. Yandex reports this preserves most of the quality of full attention at about half the inference cost.

Rollout Distillation and Continuous Training

A two-layer decoder emits Semantic ID tuples through constrained beam search with width 1,024, while a catalog trie blocks invalid prefixes. Each valid tuple expands to every track sharing it, and a four-layer cross-attention Ranking Module scores those tracks against the encoder's shared memory.

Ranking accuracy is distilled from a frozen 0.6B-parameter Teacher Ranker trained on one year of engagement data without engineered features. It uses next-item-prediction pre-training followed by multi-head ranking fine-tuning; removing pre-training drops weighted pair accuracy from 0.6215 to 0.6153. In Rollout Distillation, the decoder generates beam candidates, the teacher scores them alongside logged impressions, and the Ranking Module regresses toward those scores with mean absolute error. The joint loss—next-token prediction, rollout, and impression losses—updates the shared encoder. The teacher is removed before serving.

Training remains online: events aggregate into sessions over a 15-minute window, a GPU trainer processes them, and updated weights reach serving every 10 minutes. Yandex reports end-to-end latency of 45 minutes at the median and 60 minutes at p99. Serving uses NVIDIA Triton Inference Server with CUDA graphs and reaches 41% model FLOPs utilization.

Live Production Validation

Because playback on Yandex smart speakers can begin without a listener selecting an artist, genre, or mood, Yandex describes this as a pure-recommendation setting. In the final seven-day A/B test, 15% of randomly selected users per split received Sona instead of the production control. Every reported change is statistically significant relative to that control: active users rose 4.53%, total listening time 6.30%, likes 11.42%, deeply engaged users 7.37%, and repeat commands 17.99%. The active-user gain is 2.35 times the +1.93% increment Argus previously delivered on the same surface.

AI Mastery analysis

Sona is strong evidence that the rigid recommendation cascade is not a permanent architecture. By folding generation and ranking into one transformer, it removes the failure mode where a ranker cannot optimize candidates that upstream stages never surface. Continuous 10-minute training intervals are a distinct operational lever, but they demand streaming infrastructure that can sustain session aggregation and model rollout at this cadence.

Feature Sona (Yandex) OneRec (Kuaishou) HSTU GR (Meta)
Domain Music streaming Short video Large internet platform, multiple surfaces
Replaces Cascade Yes, in A/B test Yes, about 25% of total QPS No, reported as new model architecture
User Inputs Logged event fields only, no hand-engineered features "Multi-scale feature engineering" pathways (uid, age, gender) User action sequences (sequential transduction)
Item Output 3-level Semantic IDs, 3 x 32,000 3-level Semantic IDs via RQ-Kmeans Item IDs
Ranking Signal Produced from frozen 0.6B Teacher Ranker RL with P-Score reward model (ECPO) HSTU ranking model
Scale 8,192-event history, 0.6B teacher 10x FLOPs of prior ranking model 1.5 trillion parameters
Reported Online Gain +4.53% Active Users, +6.30% listening time, +11.42% likes +0.54% and +1.24% App Stay Time +12.4% in online A/B tests

Measured against Kuaishou's OneRec and Meta's HSTU Generative Recommenders, Sona's distinction is not end-to-end generation alone but the combination: a live cascade replacement with no hand-engineered features and supervised distillation from a teacher that never ships. OneRec retains demographic feature pathways and uses reinforcement learning; HSTU reframed recommendation as sequential transduction but was not reported as replacing a live cascade. The immediate cost is external reproducibility: Yandex publishes no code or weights, and the design assumes a platform generating massive continuous streams of pure-recommendation events.

Sources

Frequently asked questions

What is Yandex Sona?

Sona is a single generative transformer that replaces Yandex Music's 15-plus candidate generators, pre-ranking, and ranking stages with one served model. It was tested for seven days on Yandex smart speakers using logged event fields and learned Semantic IDs rather than hand-engineered features.

Does Sona use a separate ranking model?

At training time, Sona distills ranking from a frozen 0.6B-parameter Teacher Ranker, but the teacher is removed before serving. The served decoder generates Semantic ID candidates through constrained beam search with width 1,024, and a four-layer cross-attention Ranking Module scores them against shared encoder states.

How much did Sona improve engagement in Yandex's A/B test?

Yandex reports statistically significant gains over the production control: +4.53% active users, +11.42% likes, +6.30% total listening time, and +17.99% repeat commands. The active-user uplift was 2.35 times the +1.93% gain from Argus.

How often does Sona update its model?

Sona trains continuously: events aggregate into sessions over a 15-minute window, and updated weights reach serving every 10 minutes. Yandex reports end-to-end latency of 45 minutes at the median and 60 minutes at p99.

Related Reading