Build Multimodal Image Search in Elasticsearch

October 7, 2026 • guides

Modern search applications are moving beyond exact keyword matching. Multimodal similarity search maps text and images into a shared vector space so a natural-language query can retrieve visually similar catalog items without manual tagging. This guide walks through deploying a CLIP model directly in Elasticsearch, defining a dense_vector index, bulk-indexing image embeddings, and running k-nearest neighbor search.

This guide is adapted from Elasticsearch Labs's image-similarity.ipynb, available at the Elasticsearch Labs GitHub repository, and is reproduced here under the Apache-2.0 licence.

Prerequisites

  • An Elastic Cloud deployment running Elasticsearch 8.x with machine learning inference enabled.
  • Autoscaling configured to provision at least one ML node.
  • An Elastic Cloud ID and an API key with cluster management permissions.
  • A Python environment with the required dependencies:
pip install sentence-transformers==2.7.0 eland elasticsearch transformers torch tqdm Pillow streamlit

Step 1: Authenticating and connecting to the cluster

Start by importing the necessary operational libraries:

from elasticsearch import Elasticsearch
from elasticsearch.helpers import parallel_bulk
import requests
import os
import sys

import zipfile
from tqdm.auto import tqdm
import pandas as pd
from PIL import Image
from sentence_transformers import SentenceTransformer
import urllib.request

# import urllib.error
import json
from getpass import getpass

Next, capture your cluster credentials. Hardcoding credentials in deployment scripts is a security risk, so the prompt approach is used here:

# https://www.elastic.co/search-labs/tutorials/install-elasticsearch/elastic-cloud#finding-your-cloud-id
ELASTIC_CLOUD_ID = getpass("Elastic Cloud ID: ")

# https://www.elastic.co/search-labs/tutorials/install-elasticsearch/elastic-cloud#creating-an-api-key
ELASTIC_API_KEY = getpass("Elastic Api Key: ")

Finally, instantiate the client. The request_timeout=600 parameter is deliberately high because vector indexing and HNSW graph creation can be computationally expensive.

es = Elasticsearch(
    cloud_id=ELASTIC_CLOUD_ID,
    # basic_auth=(ELASTIC_CLOUD_USER, ELASTIC_CLOUD_PASSWORD),
    api_key=ELASTIC_API_KEY,
    request_timeout=600,
)

es.info()  # should return cluster info

Step 2: Deploying the text embedding model

Multimodal search requires a model that understands both text and images in the same context. Here we use a pre-trained Contrastive Language-Image Pre-training (CLIP) model imported directly into Elasticsearch with Eland.

eland_import_hub_model --cloud-id "$ELASTIC_CLOUD_ID" --hub-model-id sentence-transformers/clip-ViT-B-32-multilingual-v1 --task-type text_embedding --es-api-key "$ELASTIC_API_KEY" --start --clear-previous

The --task-type text_embedding value may seem counterintuitive for an image search guide, but it controls query-time behaviour. The image vectors are stored statically in the index; the inference that happens at query time is the transformation of the user's text search string into a vector.

If you run the command from a Jupyter notebook, prefix it with ! and use curly-brace interpolation for the Python variables:

!eland_import_hub_model --cloud-id {ELASTIC_CLOUD_ID} --hub-model-id sentence-transformers/clip-ViT-B-32-multilingual-v1 --task-type text_embedding --es-api-key {ELASTIC_API_KEY} --start --clear-previous

Step 3: Defining the index schema and vector mappings

Vector indices need a dense_vector field and explicit distance rules. The dims value must match the model output shape; the similarity metric determines how distances are computed in the latent space.

# Destination Index name
INDEX_NAME = "images"

# flag to check if index has to be deleted before creating
SHOULD_DELETE_INDEX = True

INDEX_MAPPING = {
    "properties": {
        "image_embedding": {
            "type": "dense_vector",
            "dims": 512,
            "index": True,
            "similarity": "cosine",
        },
        "photo_id": {"type": "keyword"},
        "photo_image_url": {"type": "keyword"},
        "ai_description": {"type": "text"},
        "photo_description": {"type": "text"},
        "photo_url": {"type": "keyword"},
        "photographer_first_name": {"type": "keyword"},
        "photographer_last_name": {"type": "keyword"},
        "photographer_username": {"type": "keyword"},
        "exif_camera_make": {"type": "keyword"},
        "exif_camera_model": {"type": "keyword"},
        "exif_iso": {"type": "integer"},
    }
}

# Index settings
INDEX_SETTINGS = {
    "index": {
        "number_of_replicas": "1",
        "number_of_shards": "1",
        "refresh_interval": "5s",
    }
}

# check if we want to delete index before creating the index
if SHOULD_DELETE_INDEX:
    if es.indices.exists(index=INDEX_NAME):
        print("Deleting existing %s" % INDEX_NAME)
        es.indices.delete(index=INDEX_NAME, ignore=[400, 404])

print("Creating index %s" % INDEX_NAME)
es.indices.create(
    index=INDEX_NAME, mappings=INDEX_MAPPING, settings=INDEX_SETTINGS, ignore=[400, 404]
)

The dims parameter is hardcoded to 512, which matches the output shape of the selected CLIP architecture. If a document includes a vector with any other dimension, Elasticsearch rejects the payload. Setting "index": True indexes the vectors into the HNSW graph rather than only storing them for exact score recalculation.

The similarity metric changes how the internal algorithm measures distance:

Similarity Metric Behaviour Best Used For
cosine Evaluates the angle between two vectors, ignoring their magnitude. Text and image embeddings where semantic direction matters more than term frequency or length.
dot_product Evaluates both the angle and magnitude simultaneously. Optimised, normalized embedding spaces. Offers faster computation at scale if vectors are rigorously pre-normalized.
l2_norm Measures the straight-line Euclidean distance between two data points. Computer vision applications or sensor data that do not rely on contrastive learning.

Step 4: Downloading and extracting the dataset

Download the Unsplash Lite dataset and the precomputed image embeddings:

curl -L https://unsplash.com/data/lite/1.2.0 -o unsplash-research-dataset-lite-1.2.0.zip
curl -L https://raw.githubusercontent.com/radoondas/flask-elastic-nlp/main/embeddings/images/image-embeddings.json.zip -o image-embeddings.json.zip

Then extract both archives:

# Unzip downloaded files
UNSPLASH_ZIP_FILE = "unsplash-research-dataset-lite-1.2.0.zip"
EMBEDDINGS_ZIP_FILE = "image-embeddings.json.zip"

with zipfile.ZipFile(UNSPLASH_ZIP_FILE, "r") as zip_ref:
    print("Extracting file ", UNSPLASH_ZIP_FILE, ".")
    zip_ref.extractall("data/unsplash/")

with zipfile.ZipFile(EMBEDDINGS_ZIP_FILE, "r") as zip_ref:
    print("Extracting file ", EMBEDDINGS_ZIP_FILE, ".")
    zip_ref.extractall("data/embeddings/")

Step 5: Joining and indexing the multimodal dataset

Real-world metadata is messy. Before pushing records to the cluster, merge the human-readable metadata with the raw embeddings, cleaning null values to prevent document parsing errors.

First, define a generator for parallel_bulk actions:

def gen_rows(df):
    for row in df.itertuples(index=False):
        yield {
            "_index": INDEX_NAME,
            "_id": row.photo_id,
            "_source": {
                "image_embedding": row.image_embedding,
                "photo_id": row.photo_id,
                "photo_image_url": row.photo_image_url,
                "ai_description": row.ai_description,
                "photo_description": row.photo_description,
                "photo_url": row.photo_url,
                "photographer_first_name": row.photographer_first_name,
                "photographer_last_name": row.photographer_last_name,
                "photographer_username": row.photographer_username,
                "exif_camera_make": row.exif_camera_make,
                "exif_camera_model": row.exif_camera_model,
                "exif_iso": row.exif_iso,
            },
        }

Now load the metadata, clean it, merge with the embeddings, and bulk-index the result:

df_unsplash = pd.read_csv("data/unsplash/" + "photos.tsv000", sep="\t", header=0)

# follwing 8 lines are fix for inconsistent/incorrect data
df_unsplash["photo_description"].fillna("", inplace=True)
df_unsplash["ai_description"].fillna("", inplace=True)
df_unsplash["photographer_first_name"].fillna("", inplace=True)
df_unsplash["photographer_last_name"].fillna("", inplace=True)
df_unsplash["photographer_username"].fillna("", inplace=True)
df_unsplash["exif_camera_make"].fillna("", inplace=True)
df_unsplash["exif_camera_model"].fillna("", inplace=True)
df_unsplash["exif_iso"].fillna(0, inplace=True)
## end of fix

# read subset of columns from the original/downloaded dataset
df_unsplash_subset = df_unsplash[
    [
        "photo_id",
        "photo_url",
        "photo_image_url",
        "photo_description",
        "ai_description",
        "photographer_first_name",
        "photographer_last_name",
        "photographer_username",
        "exif_camera_make",
        "exif_camera_model",
        "exif_iso",
    ]
]

# read all pregenerated embeddings
df_embeddings = pd.read_json("data/embeddings/" + "image-embeddings.json", lines=True)

df_merged = pd.merge(df_unsplash_subset, df_embeddings, on="photo_id", how="inner")

count = 0
for success, info in parallel_bulk(
    client=es,
    actions=gen_rows(df_merged),
    thread_count=5,
    chunk_size=1000,
    index=INDEX_NAME,
):
    if success:
        count += 1
        if count % 1000 == 0:
            print("Indexed %s documents" % str(count), flush=True)
            sys.stdout.flush()
    else:
        print("Doc failed", info)

print("Indexed %s image embeddings documents" % str(count), flush=True)
sys.stdout.flush()

The parallel_bulk helper parallelises network I/O across five threads, so the cluster's indexing throughput is usually the bottleneck rather than the local upload speed. The explicit dataframe sanitisation replaces NaN values with empty strings or zeroes because Elasticsearch mappings are strictly typed.

With the graph built, issue a natural-language query using the knn block. The query_vector_builder tells Elasticsearch to generate the query vector at search time from the uploaded model.

# Search queary
WHAT_ARE_YOU_LOOKING_FOR = "Valentine day flowers"

source_fields = [
    "photo_description",
    "ai_description",
    "photo_url",
    "photo_image_url",
    "photographer_first_name",
    "photographer_username",
    "photographer_last_name",
    "photo_id",
]
query = {
    "field": "image_embedding",
    "k": 5,
    "num_candidates": 100,
    "query_vector_builder": {
        "text_embedding": {
            "model_id": "sentence-transformers__clip-vit-b-32-multilingual-v1",
            "model_text": WHAT_ARE_YOU_LOOKING_FOR,
        }
    },
}

response = es.search(index=INDEX_NAME, fields=source_fields, knn=query, source=False)

print(response.body)

# the code writes the response into a file for the streamlit UI used in the optional step.
with open("json_data.json", "w") as outfile:
    json.dump(response.body["hits"]["hits"], outfile)

# Use the `loads()` method to load the JSON data
dfr = json.loads(json.dumps(response.body["hits"]["hits"]))
# Pass the generated JSON data into a pandas dataframe
dfr = pd.DataFrame(dfr)
# Print the data frame
dfr

results = pd.json_normalize(json.loads(json.dumps(response.body["hits"]["hits"])))
# results
results[
    [
        "_id",
        "_score",
        "fields.photo_id",
        "fields.photo_image_url",
        "fields.photo_description",
        "fields.photographer_first_name",
        "fields.photographer_last_name",
        "fields.ai_description",
        "fields.photo_url",
    ]
]

The execution has two discrete internal phases. First, the ML node passes the string "Valentine day flowers" through the CLIP text encoder and generates a 512-dimensional query vector. Second, the data nodes execute an approximate nearest neighbor search to find the closest image embeddings in the index.

k controls how many documents are returned to the user; num_candidates controls how many candidate nodes are tracked and evaluated per shard during graph traversal. Increasing num_candidates can improve recall by exploring more of the graph, but it also increases computation and query latency.

What to watch out for

Memory pressure on ML nodes. Uploading a Transformer model into Elasticsearch means the model weights must reside on the ML node. Under-provisioned nodes, or multiple heavy models loaded at once, can cause garbage collection pauses or out-of-memory crashes that interrupt search traffic.

Dimension mismatch errors. Model architectures evolve rapidly. If you later move to a model that outputs 768 or 1024 dimensions, your indexing pipeline will break because Elasticsearch dense_vector mappings are immutable once created. You will need to reindex into a new index with the updated schema.

Latency at scale. Vector search combines ML inference with HNSW graph traversal. If inference overhead slows down your pipeline, you are not alone; engineering teams frequently overhaul search architectures to meet latency targets, as in the Uber Eats search latency 50 percent pipeline rewrite.

Storage footprint. Dense vectors consume substantially more disk and memory than traditional text fields. Size data tiers appropriately because HNSW indexing structures require both disk space and significant RAM to remain performant.

Where to go next

After validating retrieval quality, consider a hybrid search architecture. Combining this multimodal vector search with traditional BM25 keyword matching via Reciprocal Rank Fusion (RRF) lets you retain high precision for exact product numbers or specific categories while preserving the semantic flexibility of image embeddings.

Frequently asked questions

How do I run semantic image search in Elasticsearch?

Deploy a CLIP model such as sentence-transformers/clip-ViT-B-32-multilingual-v1 with Eland, map an image_embedding field as a 512-dimensional dense_vector, and use an Elasticsearch knn query with query_vector_builder to convert a text query into a vector at search time.

What model does Elasticsearch use for multimodal image similarity?

This guide uses the Sentence Transformers CLIP model sentence-transformers/clip-ViT-B-32-multilingual-v1. It is imported into Elasticsearch with Eland using the text_embedding task type, which lets the cluster encode the user's text query while the image embeddings remain stored in the index.

Why does the dense_vector mapping require dims set to 512?

The dims value must match the output shape of the chosen embedding model. The CLIP ViT-B-32 multilingual model produces 512-dimensional vectors, so any embedded document with a different dimension would be rejected by Elasticsearch.

How do I bulk index image embeddings into Elasticsearch?

Use the elasticsearch.helpers.parallel_bulk helper with a generator that yields _index, _id, and _source actions. The guide joins a metadata DataFrame with precomputed image embeddings on photo_id, sanitizes missing values, and indexes with 5 threads and a chunk size of 1000.

What is the difference between k and num_candidates in an Elasticsearch knn query?

k is the number of nearest neighbors returned to the user. num_candidates controls how many candidate nodes are tracked and evaluated per shard during graph traversal; a higher value can improve recall at the cost of additional computation and query latency.

Free interactive tools for the decisions this piece raises.

Related Guides