Build Semantic Search over Meeting Audio with Whisper and Pinecone

October 5, 2026 β€’ guides

Audio data is notoriously difficult to index and search. Traditional lexical search relies on exact keyword matching, which fails when a colleague asks about "the pricing model" but the meeting transcript actually says "cost structure." By coupling a powerful Automatic Speech Recognition (ASR) model with a vector database, you can build a semantic search engine over your recorded meetings, allowing users to query hours of spoken content using natural language.

Engineers are frequently tasked with building these pipelines for sales teams searching calls for feature requests, engineering managers tracking technical decisions across syncs, and legal teams reviewing compliance. Building this infrastructure yourself, rather than relying on consumer-facing SaaS transcription applications, ensures sensitive internal meeting recordings never leave your controlled cloud environment.

This stack has two main compute stages. The ASR stage uses Hugging Face Transformers to run OpenAI's Whisper Large v3 locally. The retrieval stage uses a Pinecone vector index with integrated embedding, so Pinecone generates embeddings from the transcript text during upsert. The source notebook does not specify minimum GPU requirements; size your inference environment according to the checkpoint you choose.

This guide is adapted from Pinecone's meeting_transcription_semantic_search.ipynb, available in the Pinecone Examples repository under the MIT licence.

Prerequisites

Before assembling the pipeline, ensure you have the following in place:

  • A Hugging Face account to download the ASR models and sample datasets.
  • A Pinecone account and API key to provision the vector database.
  • A Python 3.x environment.
  • Sufficient local or cloud GPU compute to run the inference pipeline without memory exhaustion.

Step 1: Installing the required libraries

Our architecture relies on three primary libraries. We use the Hugging Face transformers library for pipeline orchestration and model execution, the datasets library to pull sample audio if you do not have a local file, and the pinecone client to interface with the vector index.

!pip install datasets transformers pinecone

Step 2: Configuring the environment

You must securely pass your Pinecone API key to the environment. In production, you would inject it through a secret manager or orchestration tool. Here, the code retrieves it from your local environment variables. An empty audio_path string tells the script to fall back to an open-source speech dataset; set it to a local file path to transcribe your own recording.

# Grab your desired audio file compatible with Hugging Face Pipelines and put it here
from getpass import getpass
import os 
audio_path = ""
transcription_result = []

api_key =  os.environ.get('PINECONE_API_KEY')

Step 3: Initialising the transcription pipeline

To convert spoken audio into searchable text, we instantiate the openai/whisper-large-v3 model through the Hugging Face pipeline. Whisper is an encoder-decoder Transformer that is robust to background noise and varying accents.

The key parameter here is return_timestamps=True. Without it, the model would return one continuous string of text. With timestamps, the pipeline segments the transcript into logical chunks based on natural pauses in the speech. These chunks are the foundational units we embed and store; semantic search is most effective when the text blocks are short and focused on a single topic.

Whisper Large v3 is heavy but accurate for batch processing offline files. If your requirements later shift to live transcription during a meeting, you need a streaming architecture rather than this batch pipeline. See the low-latency techniques in /news/mai-transcribe-2-streaming-2-5-wer.

from datasets import load_dataset
from transformers import pipeline

pipeline = pipeline(
    task="automatic-speech-recognition",
    model="openai/whisper-large-v3",
)


if audio_path == "":
    # use Hugging Face Sample Code instead, located here https://huggingface.co/learn/audio-course/en/chapter7/transcribe-meeting
    concatenated_librispeech = load_dataset(
    "sanchit-gandhi/concatenated_librispeech", split="train")
    transcription_result = pipeline(concatenated_librispeech[0]["audio"]["array"], return_timestamps=True)
    transcription_result
else:
    # Use your own audio file, check out this for details: https://huggingface.co/openai/whisper-large-v3
    transcription_result = pipeline(audio_path, return_timestamps=True)

Step 4: Structuring records and provisioning the vector index

With transcription complete, inspect the timestamped chunks before building database records:

print(transcription_result["chunks"])

Now reformat the chunk data into dictionaries that Pinecone can accept. Each record needs a unique _id, here derived from the chunk index, and the transcript text mapped to the sentence key.

Next, initialise the Pinecone client and create a dense index with integrated embedding via create_index_for_model. The embed argument selects llama-text-embed-v2, and the field_map tells Pinecone that the sentence key in each record contains the text to embed. This keeps embedding generation inside Pinecone during upsert instead of requiring a separate local embedding model and manual vector upload.

## use sentences as chunks, and transform into records for upsertion

# Turn into records
records = [
    {
        "_id": str(idx),
        "sentence": chunk["text"],
        # add any other desired metadata here
    }
    for idx, chunk in enumerate(transcription_result["chunks"])
]

# Import the Pinecone library
from pinecone import Pinecone

# Initialize a Pinecone client with your API key
pc = Pinecone(api_key=api_key)
namespace = "meeting-1"
# Create a dense index with integrated embedding
index_name = "meeting-transcription-index"
if not pc.has_index(index_name):
    pc.create_index_for_model(
        name=index_name,
        cloud="aws",
        region="us-east-1",
        embed={
            "model":"llama-text-embed-v2",
            "field_map":{"text": "sentence"}
        }
    )

index = pc.Index(index_name)
# query.

Step 5: Upserting data in batches

Pushing many text chunks in a single request can cause oversized payloads or timeouts on longer audio files. The batch_upsert function slices the records into batches of 96 items before sending them to the index. Because integrated embedding is enabled, these calls transmit raw JSON text rather than high-dimensional vectors.

The namespace parameter isolates this transcript under "meeting-1", which is useful when multiple meetings share the same index.

# upsert into pinecone
def batch_upsert(records, batch_size=96, namespace=namespace):
    # Great for longer audio files and batches of sentences
    for i in range(0, len(records), batch_size):
        batch = records[i:i+batch_size]
        index.upsert_records(namespace=namespace, records=batch)

batch_upsert(records)

Step 6: Querying the semantic index

The index.search method accepts a natural language string because the integrated embedding pipeline transforms the query into a vector using the same llama-text-embed-v2 model configured earlier. The top_k: 5 parameter returns the five most semantically similar transcript chunks from the "meeting-1" namespace.

The time.sleep(10) call accounts for the brief period it can take for newly upserted records to finish embedding and propagate inside the index before they are searchable.

# Replace with your own query here if needed
import time
query = "Tell me about the king of France"

# Depending on the size of your dataset, it may take a few seconds for it to finish
# embedding and populating into the index.
time.sleep(10)

results = index.search(
    namespace=namespace,
    query={
        "inputs": {"text": query},
        "top_k": 5,
    },
)

print(results)

Cleanup

After validating the pipeline, delete the index to avoid leaving test resources in place. The notebook includes this call as a comment so it is not run accidentally:

# Cleanup

#pc.delete_index(name=index_name)

Architecture Tradeoffs: Local vs Integrated Embedding

Choosing whether to embed text locally or rely on Pinecone's integrated inference changes the performance profile of the application. Here is how the approaches compare for transcription indexing:

Metric Local Embedding (Traditional) Integrated Inference (Pinecone)
Compute Load High. Requires running an embedding model alongside the heavy ASR model in local VRAM. Low. Local GPU is entirely dedicated to the Whisper ASR pipeline.
Network Traffic Heavy. Requires uploading dense arrays of thousands of floats over HTTP. Light. Uploads raw JSON text strings; arrays stay on the server.
Code Complexity High. Developer must manage model loading, pooling strategies, and tensor conversion. Low. Handled natively via the database client's mapping configuration.
Vendor Lock-in Low. You own the vectors and can migrate them to any database. Higher. The vectors are generated by the database provider's hosted model infrastructure.

What to watch out for

While this pipeline is effective for basic search, production use exposes several failure modes that require defensive engineering.

Lack of Speaker Diarization:
The Whisper model transcribes audio and segments it by pauses, but it has no native concept of who is speaking. In a real meeting, context often depends on speaker identity, such as whether a budget authorisation came from a senior executive or a junior contributor. To make search more useful for enterprise teams, run a diarization model such as Pyannote alongside Whisper, map speaker tags to the text, and inject them as metadata into the Pinecone records.

Whisper Hallucinations on Silence:
Whisper can over-generate repetitive text during long silences or background static. If a microphone stays live during a break, those hallucinations can be embedded as if they were real meeting content. Use Voice Activity Detection (VAD) to remove silent regions before transcription, or the hallucinated text will degrade the search index.

Chunking Context Loss:
The return_timestamps=True parameter splits text based on audio pauses. If a speaker pauses mid-sentence, one idea can be split across two records. A query may then miss the right result because half of the semantic meaning is in one chunk and half is in another. In production, use overlapping windows over the chunks before upserting them to preserve context boundaries.

Asynchronous Consistency:
The time.sleep(10) call is a symptom of eventual consistency after upsert. If you build a user-facing application where users expect newly uploaded audio to be searchable immediately, replace the fixed sleep with a polling mechanism or webhook that confirms the records are indexed before notifying the frontend.

Where to go next

Now that the transcript chunks are embedded and searchable, connect this retrieval engine to a large language model. Pass the top five retrieved chunks into the context window of a model such as GPT-4 or Claude to build an audio Retrieval-Augmented Generation (RAG) system. That upgrades the application from returning transcript snippets to summarising meeting decisions and answering questions based on the spoken history of the organisation.

Frequently asked questions

What speech recognition model does this pipeline use?

The pipeline uses OpenAI's Whisper Large v3 through the Hugging Face Transformers pipeline with task set to automatic-speech-recognition. The inference call passes return_timestamps=True so the audio is split into timestamped text chunks rather than one long transcript.

How does Pinecone embed the transcript chunks?

The index is created with create_index_for_model using llama-text-embed-v2. The field_map argument maps the sentence field in each record to the text input expected by the embedding model, so Pinecone generates embeddings during upsert and no local embedding model is required.

What batch size is used when upserting transcripts?

The example uses batch_upsert with batch_size=96 and upsert_records to send records to the meeting-1 namespace. Batching avoids oversized HTTP payloads when processing long audio files with many timestamped chunks.

Why does the query code include time.sleep(10)?

The sleep gives the index time to finish embedding and populating newly upserted records before the search query runs. The search then requests top_k=5 results from the meeting-1 namespace for a natural language text input.

Can I use my own audio file instead of the sample dataset?

Yes. Set audio_path to a local .mp3 or .wav file that is compatible with Hugging Face pipelines. If audio_path is left as an empty string, the code loads the sanchit-gandhi/concatenated_librispeech dataset from Hugging Face.

Free interactive tools for the decisions this piece raises.

Related Guides