Build an Agentic RAG System with smolagents and ChromaDB

September 29, 2026 • guides
RAGLangChainAI Agents

Adapted from the Hugging Face smolagents repository — rag_using_chromadb.py, licensed under Apache-2.0. Attribution: Hugging Face smolagents (https://github.com/huggingface/smolagents).


Retrieval-augmented generation is well understood at this point — you embed a knowledge base, pull the top-k chunks at query time, and stuff them into a prompt. What changes here is the reasoning layer sitting on top of that retrieval: instead of a single retrieve-then-answer pass, a CodeAgent from smolagents writes and executes actual Python code to plan its own search strategy, interpret what comes back, and decide whether to search again. That loop is what makes this an agentic RAG system. The agent treats your retriever as a callable tool it can invoke with different query strings across multiple steps, which matters enormously for questions that span several concepts or that require cross-referencing disparate chunks.

Who should care? Engineers building internal documentation assistants, support-ticket deflection systems, or any product where users ask compound questions that a single vector lookup can't reliably answer. If your questions tend to be narrow and self-contained, a standard RAG pipeline is faster and cheaper. But if users regularly ask things like "what are the differences between approach A and approach B, and when should I choose each," multi-step agentic retrieval recovers accuracy that a single-shot approach leaves on the table. The recent wave of capable reasoning agents — see our coverage of agentic coding benchmarks — makes this architecture increasingly practical to deploy.

Prerequisites

  • Python 3.10 or later
  • Packages: smolagents, langchain, langchain-chroma, langchain-huggingface, datasets, transformers, tqdm, sentence-transformers
  • A Groq API key set as GROQ_API_KEY in your environment (or an Anthropic/OpenAI key if you swap the model — the comments in the source make this straightforward)
  • Enough local disk space for the ChromaDB persistence directory (./chroma_db) — expect a few hundred MB for the HuggingFace documentation corpus

No GPU is required. The embedding model (all-MiniLM-L6-v2) and the tokenizer (thenlper/gte-small) both run on CPU without complaint, just more slowly.

Step 1: Load and split the knowledge base

The first task is turning raw documents into consistently sized chunks. The source uses the HuggingFace documentation dataset from the Hub, but the commented-out PDF loader block shows how to substitute your own files — the rest of the pipeline is identical either way. Crucially, the text splitter is initialised from the same tokenizer family as the embedding model so that chunk sizes are measured in tokens rather than characters, which prevents the embedder from silently truncating long chunks.

The deduplication loop is easy to overlook but important: the same paragraph often appears across multiple documentation pages, and duplicate embeddings waste storage and dilute retrieval precision.

import os

import datasets
from langchain.docstore.document import Document
from langchain.text_splitter import RecursiveCharacterTextSplitter
from langchain_chroma import Chroma

# from langchain_community.document_loaders import PyPDFLoader
from langchain_huggingface import HuggingFaceEmbeddings
from tqdm import tqdm
from transformers import AutoTokenizer

# from langchain_openai import OpenAIEmbeddings
from smolagents import LiteLLMModel, Tool
from smolagents.agents import CodeAgent


# from smolagents.agents import ToolCallingAgent


knowledge_base = datasets.load_dataset("m-ric/huggingface_doc", split="train")

source_docs = [
    Document(page_content=doc["text"], metadata={"source": doc["source"].split("/")[1]}) for doc in knowledge_base
]

## For your own PDFs, you can use the following code to load them into source_docs
# pdf_directory = "pdfs"
# pdf_files = [
#     os.path.join(pdf_directory, f)
#     for f in os.listdir(pdf_directory)
#     if f.endswith(".pdf")
# ]
# source_docs = []

# for file_path in pdf_files:
#     loader = PyPDFLoader(file_path)
#     docs.extend(loader.load())

text_splitter = RecursiveCharacterTextSplitter.from_huggingface_tokenizer(
    AutoTokenizer.from_pretrained("thenlper/gte-small"),
    chunk_size=200,
    chunk_overlap=20,
    add_start_index=True,
    strip_whitespace=True,
    separators=["\n\n", "\n", ".", " ", ""],
)

# Split docs and keep only unique ones
print("Splitting documents...")
docs_processed = []
unique_texts = {}
for doc in tqdm(source_docs):
    new_docs = text_splitter.split_documents([doc])
    for new_doc in new_docs:
        if new_doc.page_content not in unique_texts:
            unique_texts[new_doc.page_content] = True
            docs_processed.append(new_doc)

The separator list matters: the splitter tries \n\n first, falling back through progressively finer boundaries until it can fit the chunk within 200 tokens. The 20-token overlap gives the embedding model enough shared context at chunk boundaries to link ideas that straddle a split point.

Step 2: Embed and persist the vector store

With clean chunks in hand, you embed them and write the resulting vectors to a local ChromaDB directory. Using persist_directory means you pay the embedding cost exactly once — subsequent runs can reload the store directly without re-embedding.

print("Embedding documents... This should take a few minutes (5 minutes on MacBook with M1 Pro)")
# Initialize embeddings and ChromaDB vector store
embeddings = HuggingFaceEmbeddings(model_name="sentence-transformers/all-MiniLM-L6-v2")


# embeddings = OpenAIEmbeddings(model="text-embedding-3-small")

vector_store = Chroma.from_documents(docs_processed, embeddings, persist_directory="./chroma_db")

The commented-out OpenAIEmbeddings line is worth noting: if you need higher retrieval quality and are willing to pay per embedding call, swapping to text-embedding-3-small requires only changing that one line. The rest of the pipeline doesn't care which embedder produced the vectors, as long as you use the same model at query time.

Embedding option Model Cost Quality Latency (CPU)
Local (default) sentence-transformers/all-MiniLM-L6-v2 Free Good for English technical text ~5 min for full HF docs corpus
OpenAI (commented out) text-embedding-3-small $0.02 / 1M tokens Generally stronger, multilingual API-bound, depends on batch size
HuggingFace Inference API Any Hub model Free tier available Varies by model Network-bound

Step 3: Wrap the vector store as an agent tool

smolagents agents interact with the outside world exclusively through Tool subclasses. The RetrieverTool below is a thin wrapper that accepts a plain-English query string, fires off a similarity search, and returns the top three chunks formatted so the agent can parse them as distinct documents. The description explicitly tells the agent to phrase queries in the affirmative — "how to push a model to the Hub" rather than "how do I push a model to the Hub?" — because embedding models typically produce better similarity scores for declarative statements than for interrogative ones.

class RetrieverTool(Tool):
    name = "retriever"
    description = (
        "Uses semantic search to retrieve the parts of documentation that could be most relevant to answer your query."
    )
    inputs = {
        "query": {
            "type": "string",
            "description": "The query to perform. This should be semantically close to your target documents. Use the affirmative form rather than a question.",
        }
    }
    output_type = "string"

    def __init__(self, vector_store, **kwargs):
        super().__init__(**kwargs)
        self.vector_store = vector_store

    def forward(self, query: str) -> str:
        assert isinstance(query, str), "Your search query must be a string"
        docs = self.vector_store.similarity_search(query, k=3)
        return "\nRetrieved documents:\n" + "".join(
            [f"\n\n===== Document {str(i)} =====\n" + doc.page_content for i, doc in enumerate(docs)]
        )


retriever_tool = RetrieverTool(vector_store)

The k=3 parameter keeps context windows manageable. Increasing it to 5 or 10 will surface more relevant material for broad questions but also floods the agent's working context with noise, which can degrade answer quality and inflate token costs on every LLM call.

Step 4: Assemble and run the CodeAgent

The final step wires together the LLM, the retriever tool, and the agent loop. CodeAgent is the key choice here: rather than emitting tool calls as JSON (which ToolCallingAgent does), it generates executable Python code that the smolagents runtime actually runs. This lets the agent compose multiple retrieval calls, store intermediate results in variables, and apply conditional logic before producing its final answer — a qualitatively different capability from a single-turn retrieve-and-respond pattern.

# Choose which LLM engine to use!

# from smolagents import InferenceClientModel
# model = InferenceClientModel(model_id="Qwen/Qwen3-Next-80B-A3B-Thinking")

# from smolagents import TransformersModel
# model = TransformersModel(model_id="Qwen/Qwen3-4B-Instruct-2507")

# For anthropic: change model_id below to 'anthropic/claude-4-sonnet-latest' and also change 'os.environ.get("ANTHROPIC_API_KEY")'
model = LiteLLMModel(
    model_id="groq/openai/gpt-oss-120b",
    api_key=os.environ.get("GROQ_API_KEY"),
)

# # You can also use the ToolCallingAgent class
# agent = ToolCallingAgent(
#     tools=[retriever_tool],
#     model=model,
#     verbose=True,
# )

agent = CodeAgent(
    tools=[retriever_tool],
    model=model,
    max_steps=4,
    verbosity_level=2,
    stream_outputs=True,
)

agent_output = agent.run("How can I push a model to the Hub?")


print("Final output:")
print(agent_output)

max_steps=4 is the budget the agent has to call tools and reason before it must produce an answer. Four steps is generous for most documentation questions; if you find the agent consistently exhausting all four on simple queries, reducing to two or three cuts latency and cost. verbosity_level=2 prints each code block the agent generates and each tool result it receives — invaluable during development, noisy in production, so set it to 0 or 1 once things are working. stream_outputs=True prints tokens as they arrive so you see progress rather than staring at a blank terminal during LLM inference.

What to watch out for

Stale embeddings after document updates. ChromaDB's persistence directory holds the vectors from the first run. If you update your source documents and re-run the script, Chroma.from_documents creates a new collection on top of the old one, leading to duplicates or stale entries depending on your version of langchain-chroma. The safe approach is to delete ./chroma_db before re-ingesting, or to load the existing store with Chroma(persist_directory="./chroma_db", embedding_function=embeddings) and manage additions explicitly.

Embedding model mismatch. If you embed with all-MiniLM-L6-v2 on ingestion and reload the store with a different embedding model, every similarity search will return nonsense without raising an error. ChromaDB does not store which model produced the vectors. Document your embedding model choice and treat it as a locked dependency for the lifetime of a given collection.

max_steps as a cost multiplier. Each agent step is a separate LLM call. With max_steps=4 and k=3 retrieved chunks per tool call, a complex query can accumulate several thousand input tokens across the full run. Set max_steps conservatively and benchmark your actual token usage before going to production.

The affirmative query instruction is not enforced. The tool description asks the agent to phrase queries in affirmative form, but nothing in the code enforces this. A weaker LLM may ignore the instruction and pass interrogative strings, which can measurably lower retrieval precision. If you observe poor retrieval quality, logging the actual query strings passed to forward() is the first diagnostic step.

CodeAgent executes real code. Because CodeAgent generates and runs Python, any tool it calls with user-controlled inputs is a potential code-execution surface. smolagents sandboxes this in the local Python process by default, so the agent cannot call arbitrary shell commands — but it is still worth reviewing what your tools expose, especially if query strings come from untrusted users. Agentic systems with code execution deserve the same scrutiny as any other code-execution endpoint.

Where to go next

The smolagents repository contains additional examples covering multi-agent orchestration and tool-calling variants. If you need to serve this system at scale, look at caching retrieval results for repeated queries, replacing the local ChromaDB instance with a managed vector database, and moving LLM calls behind a queue with rate-limit handling. The ToolCallingAgent alternative shown in the commented-out lines of the source is worth benchmarking against CodeAgent on your specific query distribution — for narrow, well-defined questions the JSON-based caller is often faster and cheaper, while the code-executing agent earns its overhead on compound or exploratory queries.

Frequently asked questions

What is the difference between CodeAgent and ToolCallingAgent in smolagents?

CodeAgent generates executable Python code that the smolagents runtime actually runs, so it can compose multiple tool calls, store intermediate results in variables, and apply conditional logic before answering. ToolCallingAgent emits tool calls as JSON structures instead, which is faster and cheaper for narrow, well-defined questions but cannot chain logic between calls. For compound or exploratory queries, CodeAgent's overhead is usually justified.

How do I update the ChromaDB vector store when my documents change?

Delete the ./chroma_db directory before re-running the ingestion script — Chroma.from_documents creates a new collection on top of any existing one, which causes duplicates or stale entries. Alternatively, load the existing store with Chroma(persist_directory='./chroma_db', embedding_function=embeddings) and manage additions explicitly using its add_documents method.

Can I use my own PDF files instead of the HuggingFace documentation dataset?

Yes. The source includes a commented-out block using PyPDFLoader from langchain_community that loads a directory of PDFs into the same source_docs list. Uncomment those lines, set pdf_directory to your folder, install langchain-community, and the rest of the pipeline — splitting, embedding, and the agent loop — is identical.

Do I need a GPU to run this pipeline?

No. The embedding model (sentence-transformers/all-MiniLM-L6-v2) and the tokenizer (thenlper/gte-small) both run on CPU. The source notes that embedding the full HuggingFace documentation corpus takes roughly five minutes on a MacBook M1 Pro; CPU-only cloud instances will be slower. The LLM inference is handled remotely via the Groq API, so no local GPU is needed for that either.

What does max_steps control and how does it affect cost?

max_steps caps how many tool-call-and-reason iterations the CodeAgent can perform before it must produce a final answer. Each step is a separate LLM API call, so a setting of 4 with k=3 retrieved chunks per call can accumulate several thousand input tokens across a single query. Set it conservatively — two or three steps handles most documentation questions — and benchmark actual token usage before going to production.

Why does the RetrieverTool description tell the agent to use affirmative phrasing?

Sentence-transformer embedding models tend to produce higher cosine similarity scores when comparing two declarative statements than when comparing a question to a statement. The instruction nudges the agent to rephrase 'how do I push a model to the Hub?' as 'pushing a model to the Hub', which typically improves retrieval precision. Nothing in the code enforces this, so weaker models may ignore it — log the strings passed to forward() if retrieval quality seems poor.

Free interactive tools for the decisions this piece raises.

Related Guides