lupAI
novos-desenvolvimentos

Perplexity unveils new contextual embedding model for enhanced retrieval in rag systems

PerplexitySource: MarkTechPost01/10/2026, 02:44
Perplexity and turbopuffer have introduced pplx-embed-v2-context-9b-preview, a contextual embedding model designed to improve retrieval in Retrieval-Augmented Generation (RAG) systems. The model embeds entire documents and learns to retrieve answers along with supporting evidence, rather than relying on a single 'gold passage.' It is available as a self-hosted preview, with weights on Hugging Face under the MIT license, though it is not yet accessible via the Perplexity API. Training involves a teacher model that scores tokens based on queries and documents, enabling flexible chunk boundaries without re-annotation. The model, derived from a 9B ColBERT retrieval model, supports 1024 and 2048 dimensions, with quantization-aware training for int8 embeddings. Perplexity reports it outperforms Voyage-context-4 in chunk-retrieval tasks, with sensitivity measured by mean nDCG@10 across 74 MTEB tasks. The model addresses challenges in traditional RAG systems, where chunks often depend on external entities or definitions. By using late chunking and mean-pooling, it reduces reliance on fixed chunking strategies. Perplexity highlights issues with binary labels and linear annotation costs, proposing a soft target approach using token scores. The release includes an interactive walkthrough explaining how teacher distillation replaces single gold passage labeling, with results from context-bench showing improved performance metrics.
Perplexity unveils new contextual embedding model for enhanced retrieval in rag systems — lupAI