Perplexity's 9B Contextual Embedder Beats voyage-context-4 by 14.4 Points

October 1, 2026 • news
RAGEmbeddingsOpen WeightsBenchmarks

Perplexity Research and turbopuffer have released pplx-embed-v2-context-9b-preview, a 9-billion-parameter contextual embedding model designed for document-chunk indexing in RAG pipelines. Unlike standard dense retrievers that embed each chunk in isolation, the model encodes an entire document's chunks together in a single forward pass, so each chunk's vector carries positional and semantic information from the rest of the document. The practical result: a sentence like "Monthly rent is $3,450" can be disambiguated by lease address and term dates without duplicating that context into every chunk at index time.

Why conventional RAG training fails

The technical contribution is in the training signal, not only the architecture. Perplexity identifies four structural failures in conventional RAG training: binary relevance labels are too coarse; gold-passage annotation cost scales linearly with corpus size; labels are tied to a fixed chunking strategy; and every non-gold chunk in the positive document becomes a hard negative, even when it supplies verifiable supporting evidence.

The model's training apparatus addresses all four simultaneously. The teacher is Perplexity's own query-aware context-compression model, which reads query and document together and scores every token. Per-chunk relevance is computed as the mean of the top-n token scores inside that chunk; a temperature-scaled softmax converts those per-chunk scores into soft target distributions rather than one-hot labels. The student minimizes forward KL divergence against that distribution. A second document-level loss — InfoNCE scoring each document by its highest-scoring chunk, inspired by ColBERT's MaxSim operator — runs in parallel. Each training batch samples a random chunking strategy, so the soft token scores re-aggregate to new boundaries without re-annotation. The teacher runs only during training; it adds no latency or storage overhead at inference.

The model starts from an in-house 9B ColBERT retrieval model. A linear projection produces 2,048-dimension embeddings, with Matryoshka training also supporting 1,024 dimensions. Quantization-aware training enables native int8 output. Perplexity reports the released weights are a model soup across several checkpoints, trained on roughly 430 datasets spanning more than 50 languages, with no ConTEB data included.

context-bench results

Evaluation used context-bench, a benchmark privately held by turbopuffer to limit training contamination. It comprises 2,099 queries, 38,894 documents, and 2,458,072 sentence-level chunks across 21 domains, with a median target document length of roughly 6,100 tokens. Twelve contextual capabilities are tested, from pronoun resolution to table structure parsing. Every model is ranked exhaustively against all chunks, eliminating index-configuration effects.

Perplexity reports the model achieves 45.5% Answer Recall@10, 40.6% Evidence Recall@10, and 31.1% All-Evidence Recall@10 on context-bench, with Document Recall of 15.2% at K=1 and 61.6% at K=10. Against voyage-context-4, Perplexity reports a 14.4-point advantage on answer recall and a 5.0-point advantage on evidence recall at K=10. On ConTEB, Perplexity reports the highest average nDCG@10 among models evaluated, though the smaller pplx-embed-context-v1-4B wins on NarrativeQA and Nemotron-3-Embed-8B wins on COVID-QA. On general query-to-chunk retrieval, Perplexity reports best average performance; the model is slightly behind voyage-context-4 on query-to-document tasks. Average nDCG@10 across 74 MTEB tasks moves from 81.0% at 64-token chunks to 79.9% at 512-token chunks — a 1.1-point range.

Feature pplx-embed-v2-context-9b-preview voyage-context-4 pplx-embed-context-v1-4B Nemotron-3-Embed-8B
Chunk embedding approach Contextual (joint document pass) Contextual Contextual Independent per chunk
Parameters 9B (Perplexity); 8B (Hugging Face safetensors count) Not disclosed (MoE backbone) 4B ~8B
Dimensions 2048, 1024 (MRL) 2048, 1024, 512, 256 2560 (Matryoshka) 4096, sliceable
Native quantization int8 int8, uint8, binary, ubinary int8, binary Float
Max context evaluated 32,768 tokens 32K (120K with auto-chunking) 32K 32,768 tokens
Auto-chunking No Yes No No
Access Open weights, MIT license Hosted API — $0.12/1M tokens; first 200M free Open weights, MIT; Perplexity API Open weights, OpenMDW-1.1
Perplexity API Not yet available N/A Available Self-host

AI Mastery analysis

The shift from one-hot gold labels to soft teacher distributions is structurally significant: the model learns a continuous relevance gradient across chunks rather than a binary pass/fail boundary. This directly addresses a class of pipeline architecture problems — rather than model capability gaps — that increasingly explain RAG quality differences. The random chunking strategy per batch matters equally: the trained model is not brittle to chunk-boundary choices at deployment, which is one of the most persistent tuning burdens in production RAG.

The storage arithmetic is worth examining for teams currently on voyage-context-4. Perplexity reports that 1024-dim int8 vectors (1 KB per chunk) slightly exceed voyage-context-4 at 2048-dim float32 (8 KB per chunk) on its chunk-retrieval suite — an 8× storage reduction per vector with no reported quality penalty.

Two adoption constraints are real. First, the preview warning is unambiguous: embeddings produced now will not be compatible with future releases, requiring full re-indexing when a stable version ships. Second, the model requires trust_remote_code=True and transformers>=5.4.0, and uses separate encode_queries and encode methods; conflating them silently degrades retrieval quality. The absence of Perplexity API access means self-hosting a 9B model is the only current path, which limits who can operationalize this today versus simply benchmark it.

The competitive frontier for RAG retrieval has moved from chunk size and overlap tuning to training-objective design — specifically, how well a model captures inter-chunk evidence chains rather than isolated answer passages. The Evidence Recall and All-Evidence metrics on context-bench operationalize this directly, and turbopuffer's decision to hold the benchmark privately to limit contamination suggests this evaluation framing will itself become a template. Whether Perplexity's reported 14.4-point answer-recall gap over voyage-context-4 holds on domain-specific corpora outside context-bench's 21 domains remains the open question for practitioners.

Sources

Frequently asked questions

What benchmark scores does pplx-embed-v2-context-9b-preview achieve on context-bench?

Perplexity reports 45.5% Answer Recall@10, 40.6% Evidence Recall@10, and 31.1% All-Evidence Recall@10 on context-bench. Document Recall is 15.2% at K=1 and 61.6% at K=10.

How does pplx-embed-v2-context-9b-preview compare to voyage-context-4?

Perplexity reports a 14.4-point advantage on Answer Recall@10 and a 5.0-point advantage on Evidence Recall@10 against voyage-context-4. The model is slightly behind voyage-context-4 on query-to-document tasks.

Can pplx-embed-v2-context-9b-preview run via the Perplexity API?

No. At launch, the model is available only as open weights under the MIT license on Hugging Face. Perplexity API access has not yet been made available.

What are the storage requirements compared to voyage-context-4?

Perplexity reports that 1024-dim int8 vectors use 1 KB per chunk, versus 8 KB per chunk for voyage-context-4 at 2048-dim float32 — an 8× reduction with no reported quality penalty on chunk-retrieval tasks.

What are the technical requirements to load pplx-embed-v2-context-9b-preview?

The model requires transformers>=5.4.0 and must be loaded with trust_remote_code=True. Queries must use the encode_queries method; using encode instead silently degrades retrieval quality. Because this is a preview release, embeddings produced now will not be compatible with future versions.

Free interactive tools for the decisions this piece raises.

Related Reading