Build Semantic Search over Meeting Audio with Whisper and Pinecone
In this article
- Prerequisites
- Step 1: Installing the required libraries
- Step 2: Configuring the environment
- Step 3: Initialising the transcription pipeline
- Step 4: Structuring records and provisioning the vector index
- Step 5: Upserting data in batches
- Step 6: Querying the semantic index
- Cleanup
- Architecture Tradeoffs: Local vs Integrated Embedding
- What to watch out for
- Where to go next
Audio data is notoriously difficult to index and search. Traditional lexical search relies on exact keyword matching, which fails when a colleague asks about "the pricing model" but the meeting transcript actually says "cost structure." By coupling a powerful Automatic Speech Recognition (ASR) model with a vector database, you can build a semantic search engine over your recorded meetings, allowing users to query hours of spoken content using natural language.
Engineers are frequently tasked with building these pipelines for sales teams searching calls for feature requests, engineering managers tracking technical decisions across syncs, and legal teams reviewing compliance. Building this infrastructure yourself, rather than relying on consumer-facing SaaS transcription applications, ensures sensitive internal meeting recordings never leave your controlled cloud environment.
This stack has two main compute stages. The ASR stage uses Hugging Face Transformers to run OpenAI's Whisper Large v3 locally. The retrieval stage uses a Pinecone vector index with integrated embedding, so Pinecone generates embeddings from the transcript text during upsert. The source notebook does not specify minimum GPU requirements; size your inference environment according to the checkpoint you choose.
This guide is adapted from Pinecone's meeting_transcription_semantic_search.ipynb, available in the Pinecone Examples repository under the MIT licence.
Prerequisites
Before assembling the pipeline, ensure you have the following in place:
- A Hugging Face account to download the ASR models and sample datasets.
- A Pinecone account and API key to provision the vector database.
- A Python 3.x environment.
- Sufficient local or cloud GPU compute to run the inference pipeline without memory exhaustion.
Step 1: Installing the required libraries
Our architecture relies on three primary libraries. We use the Hugging Face transformers library for pipeline orchestration and model execution, the datasets library to pull sample audio if you do not have a local file, and the pinecone client to interface with the vector index.
!pip install datasets transformers pinecone
Step 2: Configuring the environment
You must securely pass your Pinecone API key to the environment. In production, you would inject it through a secret manager or orchestration tool. Here, the code retrieves it from your local environment variables. An empty audio_path string tells the script to fall back to an open-source speech dataset; set it to a local file path to transcribe your own recording.
# Grab your desired audio file compatible with Hugging Face Pipelines and put it here
from getpass import getpass
import os
audio_path = ""
transcription_result = []
api_key = os.environ.get('PINECONE_API_KEY')
Step 3: Initialising the transcription pipeline
To convert spoken audio into searchable text, we instantiate the openai/whisper-large-v3 model through the Hugging Face pipeline. Whisper is an encoder-decoder Transformer that is robust to background noise and varying accents.
The key parameter here is return_timestamps=True. Without it, the model would return one continuous string of text. With timestamps, the pipeline segments the transcript into logical chunks based on natural pauses in the speech. These chunks are the foundational units we embed and store; semantic search is most effective when the text blocks are short and focused on a single topic.
Whisper Large v3 is heavy but accurate for batch processing offline files. If your requirements later shift to live transcription during a meeting, you need a streaming architecture rather than this batch pipeline. See the low-latency techniques in /news/mai-transcribe-2-streaming-2-5-wer.
from datasets import load_dataset
from transformers import pipeline
pipeline = pipeline(
task="automatic-speech-recognition",
model="openai/whisper-large-v3",
)
if audio_path == "":
# use Hugging Face Sample Code instead, located here https://huggingface.co/learn/audio-course/en/chapter7/transcribe-meeting
concatenated_librispeech = load_dataset(
"sanchit-gandhi/concatenated_librispeech", split="train")
transcription_result = pipeline(concatenated_librispeech[0]["audio"]["array"], return_timestamps=True)
transcription_result
else:
# Use your own audio file, check out this for details: https://huggingface.co/openai/whisper-large-v3
transcription_result = pipeline(audio_path, return_timestamps=True)
Step 4: Structuring records and provisioning the vector index
With transcription complete, inspect the timestamped chunks before building database records:
print(transcription_result["chunks"])
Now reformat the chunk data into dictionaries that Pinecone can accept. Each record needs a unique _id, here derived from the chunk index, and the transcript text mapped to the sentence key.
Next, initialise the Pinecone client and create a dense index with integrated embedding via create_index_for_model. The embed argument selects llama-text-embed-v2, and the field_map tells Pinecone that the sentence key in each record contains the text to embed. This keeps embedding generation inside Pinecone during upsert instead of requiring a separate local embedding model and manual vector upload.
## use sentences as chunks, and transform into records for upsertion
# Turn into records
records = [
{
"_id": str(idx),
"sentence": chunk["text"],
# add any other desired metadata here
}
for idx, chunk in enumerate(transcription_result["chunks"])
]
# Import the Pinecone library
from pinecone import Pinecone
# Initialize a Pinecone client with your API key
pc = Pinecone(api_key=api_key)
namespace = "meeting-1"
# Create a dense index with integrated embedding
index_name = "meeting-transcription-index"
if not pc.has_index(index_name):
pc.create_index_for_model(
name=index_name,
cloud="aws",
region="us-east-1",
embed={
"model":"llama-text-embed-v2",
"field_map":{"text": "sentence"}
}
)
index = pc.Index(index_name)
# query.
Step 5: Upserting data in batches
Pushing many text chunks in a single request can cause oversized payloads or timeouts on longer audio files. The batch_upsert function slices the records into batches of 96 items before sending them to the index. Because integrated embedding is enabled, these calls transmit raw JSON text rather than high-dimensional vectors.
The namespace parameter isolates this transcript under "meeting-1", which is useful when multiple meetings share the same index.
# upsert into pinecone
def batch_upsert(records, batch_size=96, namespace=namespace):
# Great for longer audio files and batches of sentences
for i in range(0, len(records), batch_size):
batch = records[i:i+batch_size]
index.upsert_records(namespace=namespace, records=batch)
batch_upsert(records)
Step 6: Querying the semantic index
The index.search method accepts a natural language string because the integrated embedding pipeline transforms the query into a vector using the same llama-text-embed-v2 model configured earlier. The top_k: 5 parameter returns the five most semantically similar transcript chunks from the "meeting-1" namespace.
The time.sleep(10) call accounts for the brief period it can take for newly upserted records to finish embedding and propagate inside the index before they are searchable.
# Replace with your own query here if needed
import time
query = "Tell me about the king of France"
# Depending on the size of your dataset, it may take a few seconds for it to finish
# embedding and populating into the index.
time.sleep(10)
results = index.search(
namespace=namespace,
query={
"inputs": {"text": query},
"top_k": 5,
},
)
print(results)
Cleanup
After validating the pipeline, delete the index to avoid leaving test resources in place. The notebook includes this call as a comment so it is not run accidentally:
# Cleanup
#pc.delete_index(name=index_name)
Architecture Tradeoffs: Local vs Integrated Embedding
Choosing whether to embed text locally or rely on Pinecone's integrated inference changes the performance profile of the application. Here is how the approaches compare for transcription indexing:
| Metric | Local Embedding (Traditional) | Integrated Inference (Pinecone) |
|---|---|---|
| Compute Load | High. Requires running an embedding model alongside the heavy ASR model in local VRAM. | Low. Local GPU is entirely dedicated to the Whisper ASR pipeline. |
| Network Traffic | Heavy. Requires uploading dense arrays of thousands of floats over HTTP. | Light. Uploads raw JSON text strings; arrays stay on the server. |
| Code Complexity | High. Developer must manage model loading, pooling strategies, and tensor conversion. | Low. Handled natively via the database client's mapping configuration. |
| Vendor Lock-in | Low. You own the vectors and can migrate them to any database. | Higher. The vectors are generated by the database provider's hosted model infrastructure. |
What to watch out for
While this pipeline is effective for basic search, production use exposes several failure modes that require defensive engineering.
Lack of Speaker Diarization:
The Whisper model transcribes audio and segments it by pauses, but it has no native concept of who is speaking. In a real meeting, context often depends on speaker identity, such as whether a budget authorisation came from a senior executive or a junior contributor. To make search more useful for enterprise teams, run a diarization model such as Pyannote alongside Whisper, map speaker tags to the text, and inject them as metadata into the Pinecone records.
Whisper Hallucinations on Silence:
Whisper can over-generate repetitive text during long silences or background static. If a microphone stays live during a break, those hallucinations can be embedded as if they were real meeting content. Use Voice Activity Detection (VAD) to remove silent regions before transcription, or the hallucinated text will degrade the search index.
Chunking Context Loss:
The return_timestamps=True parameter splits text based on audio pauses. If a speaker pauses mid-sentence, one idea can be split across two records. A query may then miss the right result because half of the semantic meaning is in one chunk and half is in another. In production, use overlapping windows over the chunks before upserting them to preserve context boundaries.
Asynchronous Consistency:
The time.sleep(10) call is a symptom of eventual consistency after upsert. If you build a user-facing application where users expect newly uploaded audio to be searchable immediately, replace the fixed sleep with a polling mechanism or webhook that confirms the records are indexed before notifying the frontend.
Where to go next
Now that the transcript chunks are embedded and searchable, connect this retrieval engine to a large language model. Pass the top five retrieved chunks into the context window of a model such as GPT-4 or Claude to build an audio Retrieval-Augmented Generation (RAG) system. That upgrades the application from returning transcript snippets to summarising meeting decisions and answering questions based on the spoken history of the organisation.
Frequently asked questions
What speech recognition model does this pipeline use?
The pipeline uses OpenAI's Whisper Large v3 through the Hugging Face Transformers pipeline with task set to automatic-speech-recognition. The inference call passes return_timestamps=True so the audio is split into timestamped text chunks rather than one long transcript.
How does Pinecone embed the transcript chunks?
The index is created with create_index_for_model using llama-text-embed-v2. The field_map argument maps the sentence field in each record to the text input expected by the embedding model, so Pinecone generates embeddings during upsert and no local embedding model is required.
What batch size is used when upserting transcripts?
The example uses batch_upsert with batch_size=96 and upsert_records to send records to the meeting-1 namespace. Batching avoids oversized HTTP payloads when processing long audio files with many timestamped chunks.
Why does the query code include time.sleep(10)?
The sleep gives the index time to finish embedding and populating newly upserted records before the search query runs. The search then requests top_k=5 results from the meeting-1 namespace for a natural language text input.
Can I use my own audio file instead of the sample dataset?
Yes. Set audio_path to a local .mp3 or .wav file that is compatible with Hugging Face pipelines. If audio_path is left as an empty string, the code loads the sanchit-gandhi/concatenated_librispeech dataset from Hugging Face.
Related Guides

Generate Synthetic RAG Test Sets With Ragas
Build, save and reuse a Ragas knowledge graph to generate single-hop and multi-hop evaluation queries from Markdown documents.

Build a ReAct Agent with Guaranteed JSON Output Using Outlines
Use Outlines' constrained generation to make a ReAct agent that cannot emit malformed JSONβno fine-tuning, no GPU required.
DigitalOcean Managed Agents: MicroVM Runtime and 16,000-Tool Gateway
DigitalOcean launches Managed Agents in public preview β microVM isolation, session persistence, and 16,000+ governed tools via a unified MCP endpoint.