pplx-embed-v2-late: 0.6B Queries 9B Index at 63.5% ViDoRe
In this article
Perplexity has released pplx-embed-v2-late, a pair of MIT-licensed ColBERT-style multimodal embedding models: a 0.6B edge encoder and a 9B high-quality indexer that share one embedding space. Perplexity reports 92.4% on MADQA for the 9B model and 61.2% on ViDoRe v3 Markdown for the 0.6B model, giving engineering teams a concrete choice between retrieval accuracy and strict latency limits.
Architecture and shared embedding space
Unlike dense models that compress a document into one vector, pplx-embed-v2-late keeps one 128-dimensional vector per token. Query-document matching uses MaxSim scoring: each query token finds its strongest document token match, and those maxima are summed. Pages are encoded directly as images, so the models bypass external OCR pipelines.
Both sizes are built on Qwen3.5 with bidirectional attention. The 0.6B model is based on Qwen3.5-0.8B pruned to 12 text layers and has about 240 million active parameters for text and 340 million for images, out of 594,321,600 F32 parameters on its Hugging Face card. The 9B model has 7.4 billion active parameters, though MarkTechPost notes Hugging Face classifies it as an 8B model.
Perplexity distilled both checkpoints from an internal 18-billion-parameter ColBERT teacher using a token-level LEAF-style objective trained on pair and triplet data. The 0.6B model was fully fine-tuned; for the 9B model, only the final eight transformer layers were fully fine-tuned, while the remaining transformer layers and vision encoder were adapted with LoRA.
Benchmark results
All benchmark figures below are Perplexity's self-reported results. On the MADQA agentic PDF QA benchmark, the 9B model scored 92.4%, beating the Mixedbread retriever at 88.9% but trailing Mixedbread Agentic Search at 93.4%; the 0.6B model scored 90.1%.
On 72 domain-specific text tasks, the 9B model leads tested baselines with an average nDCG@10 of 81.3%, while the 0.6B model's 78.0% sits 0.3 percentage points behind Google's gemini-embedding-2. For image retrieval, the 0.6B model scored 62.3% on ViDoRe v3, within 1.2 points of NVIDIA's nemotron-colembed-v2-8b, and the 9B model reached 65.2%. MarkTechPost highlights that Tencent's EVIE still leads ViDoRe v3 image retrieval, and Google's Gemini Embedding 2 beats the 9B on MIRACL-Vision and by 2 points on PPLX-Q2I.
On Q2D-Web Recall@1000, both models cleared the previous best of 69.3%, with the 9B scoring 74.8% and the 0.6B scoring 73.6%. On BrowseComp+, the 9B model's 64.0% accuracy was 4.9 points above the next ColBERT model and 8.7 points above the best dense model.
For asymmetric deployments, querying a 9B-encoded index with the 0.6B model scored 63.5% on ViDoRe v3 image retrieval, beating the 0.6B symmetric baseline of 62.3% without extra query compute. Perplexity reports this recovers roughly half the 9B text-quality gap at the same query cost.
Deployment and constraints
The models require sentence-transformers >= 6.0.0 and transformers >= 5.4.0. Text-only and image-only batches must be encoded separately; mixed text and image inputs are unsupported. Unlike PyLate, which places query and document markers at the second position, this architecture expects them at the first position.
MarkTechPost estimates the 0.6B model needs about 1.2 GB of bf16 weight memory, fitting on laptops and edge devices, while the 9B model needs 16 to 18 GB on datacenter or high-memory GPUs. A hosted Perplexity API endpoint is planned but not yet live.
| Feature | pplx-embed-v2-late | NVIDIA nemotron-colembed-vl-8b-v2 | TopK topk-embed-v1 | Google Gemini Embedding 2 |
|---|---|---|---|---|
| Size | 0.6B, 9B | ~8.8B | 0.8B, 2B open | Not disclosed |
| Vector width | 128 per token | 4,096 per token | 2,048 per token (small) | 128 to 3,072, 1 vector |
| Inputs | Text, images, page renders | Text queries, page images | Text, page images | Text, image, video, audio, PDF |
| Shared space | Yes | Not stated | Not stated | Not applicable |
| License | MIT | CC-BY-NC-4.0 | Apache 2.0 (small) | Proprietary |
AI Mastery analysis
The release shows that shifting to constrained inference topologies is an effective optimization when local query latency matters more than cloud scale. By forcing the 0.6B and 9B models into the same embedding space through LEAF-style distillation, Perplexity gives systems engineers a way out of the ColBERT compute trap: run heavy 9B indexing offline, then execute only the 0.6B query encoding live against the remote index.
The main tradeoff is storage. Because the model outputs a 128-dimensional vector per token, index sizes scale linearly with document length. Although 128 dimensions are far narrower than the 2,048 or 4,096 dimensions emitted by TopK or NVIDIA ColBERT equivalents, retaining a vector for every token still changes database provisioning compared with single-vector dense retrieval. Implementations will need aggressive vector database optimization so MaxSim scoring does not erase the latency advantages of the 0.6B query encoder.
Decoupling query-encoder size from document-indexer size creates a practical lever for production RAG systems. As multimodal retrieval moves beyond single-vector bottlenecks, asymmetric model pairings are becoming the main way to balance index accuracy against live-query constraints.
Sources
Frequently asked questions
How much VRAM does pplx-embed-v2-late require?
MarkTechPost estimates about 1.2 GB of bf16 memory for the 0.6B model's weights alone, while the 9B model needs roughly 16 to 18 GB. Published checkpoints are stored in F32, which doubles download size.
Does pplx-embed-v2-late run on a single GPU?
The 0.6B model is designed for laptops, edge devices or small GPUs, while the 9B model is built for datacenter or high-memory GPUs. Both model cards show CUDA GPU usage.
What is the best benchmark score for pplx-embed-v2-late?
In Perplexity's self-reported results, the 9B model scores 92.4% accuracy on MADQA, the best result in that comparison. The 0.6B model's weakest listed score is 61.2% on ViDoRe v3 Markdown.
Can the 0.6B model query an index built with the 9B model?
Yes. Perplexity reports a 9B-encoded index queried by the 0.6B model scores 63.5% on ViDoRe v3 image retrieval, beating the 0.6B symmetric baseline of 62.3% at the same query cost and recovering roughly half the text-quality gap.
Related Reading
Mistral Claims 82% Cyber Fix Score for 1T Le Chonk
Mistral's 1T-parameter Mistral Large 4 preview posts an 82% cyber fix score, with open weights promised by the end of October 2026.
Perplexity's 9B Contextual Embedder Beats voyage-context-4 by 14.4 Points
Perplexity releases pplx-embed-v2-context-9b-preview, an open-weights 9B RAG embedder that scores 45.5% Answer Recall@10 and cuts storage 8ร vs float32.
Cohere's 218B MoE Scores 83.6 on WMT26, Beats DeepL and Google Translate
North Small Translate activates 25B of 218B parameters per token, scores 83.6 on WMT26 across 50 languages, and runs on a single B200 in 4-bit mode.