In this article
- Prerequisites
- Step 1: Installing Dependencies and Preparing Source Data
- Step 2: Generating Dense and Sparse Embeddings with BGE-M3
- Step 3: Defining a Dual-Vector Collection Schema
- Step 4: Creating Dedicated Indices for Both Modalities
- Step 5: Executing Hybrid Search with Reciprocal Rank Fusion
- Strategy Comparison Matrix
- What to Watch Out For
- Where to Go Next
Retrieval-Augmented Generation applications eventually hit a wall when they rely only on dense vector search. Dense embeddings are strong at capturing conceptual meaning, but they are notoriously unreliable for exact keyword matching. If your users search for specific acronyms, proprietary error codes, or niche proper nouns, dense vectors often collapse those precise tokens into a general semantic region and return documents that are related in meaning but factually useless.
A common fix is a hybrid retrieval architecture that combines dense vector similarity with sparse lexical search. This guide walks through a production-oriented hybrid pipeline in Milvus. It generates sparse and dense representations from the same BGE-M3 model, stores them in one collection schema, and fuses both result sets with Reciprocal Rank Fusion. The approach preserves semantic recall while restoring exact token precision for jargon-heavy corpora.
This guide is adapted from the Milvus Bootcamp's sparse_dense_embeddings_tutorial.ipynb, available in the Milvus Bootcamp repository under the Apache-2.0 license.
Prerequisites
You need a running local Milvus instance bound to port 19530, which is the connection target used throughout the source notebook. The tutorial assumes a current Milvus 2.x build with native support for sparse vectors and hybrid search. You also need a Python environment that can install the official Milvus client and the machine learning utilities required to load BGE-M3 locally.
Because pymilvus[model] loads the BGE-M3 tokenizer and weights into your Python process, ensure that the machine has enough available memory for the model plus the Milvus segments that col.load() will map into RAM. The example corpus is tiny, but the same ingestion path applies to much larger collections.
Step 1: Installing Dependencies and Preparing Source Data
The first code block installs the Milvus client, the optional model extra, and scikit-learn. The pymilvus[model] extra pulls in the model-loading utilities used to run BGE-M3 directly through the client rather than calling a separate embedding microservice.
!pip install -U pymilvus
!pip install -U 'pymilvus[model]'
!pip install -U scikit-learn
The source data is deliberately small so the retrieval behavior is easy to inspect. In a real application, this is the point where chunking strategy becomes critical. Chunks that are too long dilute the term frequencies that sparse retrieval depends on, while chunks that are too short lose the semantic context dense vectors need. When working with long-context embeddings such as Perplexity pplx-embed-v2, the chunk boundaries should still be chosen so that exact domain terms remain, even if the model can accept a large input window.
docs = [
"Artificial intelligence was founded as an academic discipline in 1956.",
"Alan Turing was the first person to conduct substantial research in AI.",
"Born in Maida Vale, London, Turing was raised in southern England.",
]
query = "Who started AI research?"
Step 2: Generating Dense and Sparse Embeddings with BGE-M3
Hybrid search requires each document to be embedded twice: once as a dense vector and once as a sparse vector. BGE-M3 is a single model architecture that can expose both representations from one inference call. The import block also includes TfidfVectorizer, utility, random, string, and numpy because the original notebook uses those tools while inspecting sparse vocabulary behavior before switching to the BGE-M3 path.
import random
import string
import numpy as np
from sklearn.feature_extraction.text import TfidfVectorizer
from pymilvus import (
utility,
FieldSchema, CollectionSchema, DataType,
Collection, AnnSearchRequest, RRFRanker, connections,
)
from pymilvus.model.hybrid import BGEM3EmbeddingFunction
The BGEM3EmbeddingFunction is instantiated with CPU execution and FP32 precision. The device and use_fp16 arguments matter in production: running on a CUDA device with FP16 enabled normally accelerates batch embedding and reduces memory bandwidth pressure. For this small example, CPU and FP32 keep the setup simple and deterministic.
ef = BGEM3EmbeddingFunction(use_fp16=False, device="cpu")
dense_dim = ef.dim["dense"]
Calling the function against the document list and query returns a dictionary with "dense" and "sparse" keys. The dense dimensionality is read from the model rather than hard-coded because Milvus requires an exact dimension for FLOAT_VECTOR fields before insertion.
docs_embeddings = ef(docs)
query_embeddings = ef([query])
Step 3: Defining a Dual-Vector Collection Schema
Milvus is strongly typed and relies on a rigid schema to optimize memory mapping and hardware acceleration. The first step is connecting to the local server on port 19530.
connections.connect("default", host="localhost", port="19530")
The schema declares an auto-generated primary key, a text field, a sparse vector field, and a dense vector field. DataType.SPARSE_FLOAT_VECTOR does not take a dimension because sparse vectors are stored as variable-length pairs of active token IDs and floating-point weights. In contrast, DataType.FLOAT_VECTOR must be locked to the exact dense_dim reported by BGE-M3.
fields = [
# Use auto generated id as primary key
FieldSchema(name="pk", dtype=DataType.VARCHAR,
is_primary=True, auto_id=True, max_length=100),
FieldSchema(name="text", dtype=DataType.VARCHAR, max_length=512),
FieldSchema(name="sparse_vector", dtype=DataType.SPARSE_FLOAT_VECTOR),
FieldSchema(name="dense_vector", dtype=DataType.FLOAT_VECTOR,
dim=dense_dim)
]
schema = CollectionSchema(fields, "")
col = Collection("sparse_dense_demo", schema)
Step 4: Creating Dedicated Indices for Both Modalities
Dense and sparse vectors have different mathematical properties, so they need separate index types and similarity metrics. The sparse field uses an inverted index with inner product scoring, which effectively multiplies the weights of overlapping terms. The dense field uses a FLAT index with cosine distance, which performs an exhaustive search over every dense vector. That guarantees exact retrieval for the current small dataset, but it becomes expensive at production scale.
sparse_index = {"index_type": "SPARSE_INVERTED_INDEX", "metric_type": "IP"}
dense_index = {"index_type": "FLAT", "metric_type": "COSINE"}
col.create_index("sparse_vector", sparse_index)
col.create_index("dense_vector", dense_index)
After the indices exist, entities are inserted as three aligned lists: original text, sparse embeddings, and dense embeddings. The col.flush() call seals the buffered data into queryable segments. The separate col.load() call then maps those segments into RAM; queries will fail until the collection has been loaded.
entities = [docs, docs_embeddings["sparse"], docs_embeddings["dense"]]
col.insert(entities)
col.flush()
col.load()
Step 5: Executing Hybrid Search with Reciprocal Rank Fusion
The actual hybrid query is built from two AnnSearchRequest objects. Each request launches an independent nearest-neighbor search against one vector field. The sparse request uses IP and the dense request uses COSINE, matching the index metrics defined earlier.
sparse_search_params = {"metric_type": "IP"}
sparse_req = AnnSearchRequest(query_embeddings["sparse"],
"sparse_vector", sparse_search_params, limit=2)
dense_search_params = {"metric_type": "COSINE"}
dense_req = AnnSearchRequest(query_embeddings["dense"],
"dense_vector", dense_search_params, limit=2)
res = col.hybrid_search([sparse_req, dense_req], rerank=RRFRanker(),
limit=2, output_fields=["text"])
The important detail is RRFRanker(). Reciprocal Rank Fusion compares the ordinal rank of each document in the sparse and dense result lists instead of comparing raw similarity scores. Since inner product values from sparse retrieval and cosine values from dense retrieval sit on different numerical ranges, adding those raw scores directly would skew the final order. RRF avoids that problem by merging the two ranked lists and promoting documents that performed well in both passes.
Strategy Comparison Matrix
When choosing a retrieval architecture, the decision usually comes down to precision, recall, and infrastructure cost. Adding a sparse path changes the operational profile of the database in ways that dense-only teams often underestimate.
| Strategy | Strengths | Weaknesses | Memory & Compute Cost |
|---|---|---|---|
| Dense Only | Understands semantic intent and handles synonyms or paraphrases well. | Fails on exact keyword matches, serial numbers, and domain-specific acronyms. | Baseline: One fixed-size dense vector per document. |
| Sparse Only | Strong lexical precision and heavy weighting for rare keywords. | No semantic understanding; degrades with typos or query rephrasing. | Variable: Inverted index size grows with vocabulary breadth. |
| Hybrid + RRF | Combines semantic context with exact token precision for stronger overall recall. | More moving parts; requires alignment between the sparse and dense embedding paths. | High: Stores two vector representations, runs two search passes, and adds reranking overhead. |
What to Watch Out For
Hybrid search introduces failure modes that do not exist in dense-only systems.
Index mismatching silently breaks retrieval. The sparse vector maps heavily to specific token IDs. If you query a BGE-M3 sparse index using a manually generated TF-IDF vector, or even a different tokenizer version, the inverted index searches for mismatched token identifiers and may return poor results without an obvious error. Keep the query embedding path identical to the ingestion path.
RRF scores are opaque. Reciprocal Rank Fusion produces a rank-based merge score, not a true geometric distance. A document with a fused score of 0.6 in one result set is not directly comparable to a 0.6 in a differently sized result list. Absolute score thresholds therefore need recalibration whenever the candidate list size or retriever weighting changes.
Watch query-node saturation. Each hybrid_search call runs two nearest-neighbor searches before the fusion step. If a cluster is provisioned for one search per request, enabling hybrid search changes the query-node workload materially. Production deployments need enough headroom for the combined sparse, dense, and reranking paths.
Prepare to replace FLAT for larger collections. The dense FLAT index used here guarantees exact cosine retrieval, but brute-force scanning becomes expensive as the corpus grows. At scale, you would drop and recreate the dense index with an approximate algorithm such as HNSW to keep query latency manageable while accepting a small trade-off in recall.
Where to Go Next
Once the local hybrid pipeline works, the next step is to make it scale. Move BGEM3EmbeddingFunction to a CUDA device with FP16 enabled, replace FLAT with HNSW for the dense field, and test whether unweighted RRF is appropriate for your domain. Corpora with very precise terminology often benefit from applying a custom weight to the sparse results so exact lexical matches are promoted more aggressively.
Frequently asked questions
What is hybrid search in Milvus?
Hybrid search combines dense semantic embeddings with sparse lexical embeddings in one query. This guide uses BGEM3EmbeddingFunction to generate both representations, stores them in SPARSE_FLOAT_VECTOR and FLOAT_VECTOR fields, and fuses the ranked results with RRFRanker. It is useful when a corpus contains exact acronyms, error codes, or proper nouns that dense vectors alone often miss.
Why does Milvus use Reciprocal Rank Fusion for hybrid search?
Sparse inner product scores and dense cosine similarity scores have different numerical ranges, so directly summing them can distort the final ranking. Reciprocal Rank Fusion compares each result's ordinal position from the sparse and dense passes, then merges those ranks. This lets documents that appear near the top of both lists rise without requiring score normalization.
What Milvus index types are used for sparse and dense vectors?
Sparse vectors use SPARSE_INVERTED_INDEX with an IP metric, which matches overlapping terms by inner product. Dense vectors in this tutorial use FLAT with a COSINE metric for exact brute-force retrieval. At production scale, FLAT is typically replaced with HNSW or another approximate index to reduce query latency.
How do I query a Milvus collection with sparse and dense vectors?
Create two AnnSearchRequest objects: one for the sparse_vector field with IP and one for the dense_vector field with COSINE. Pass both to Collection.hybrid_search along with RRFRanker and the desired limit. The response contains the top-k documents ranked by fused position, and output_fields can include stored values such as the original text.
Can I use BGE-M3 for both dense and sparse embeddings?
Yes, BGEM3EmbeddingFunction in pymilvus.model.hybrid returns both dense and sparse vectors from a single model. This keeps tokenization and model weights matched across representation types, avoiding mismatched sparse token IDs. Using one model also simplifies ingestion because both embeddings are produced in the same inference pass.
Related Guides

Build a Self-Correcting RAG Agent with LangGraph and Milvus
Wire adaptive routing, corrective RAG, and self-RAG into a stateful LangGraph agent backed by Milvus Lite — no OpenAI key required.
Designing a Production-Grade RAG Architecture
Large Language Models are powerful—but infamously unreliable when forced to guess. Learn how to build a production-grade RAG architecture that eliminates hallucinations through hybrid search, reranking, and structured ingestion.

Build Semantic Search over Meeting Audio with Whisper and Pinecone
Transcribe meeting audio with Whisper Large v3, embed chunks with Pinecone's integrated llama-text-embed-v2 model, and query transcripts semantically.