In this article
This guide is adapted from rag_testset_generation.md in the Ragas project, which is distributed under the Apache-2.0 license. It shows how to load Markdown documents, construct and persist a knowledge graph, and generate a synthetic evaluation set containing single-hop and multi-hop queries.
The resulting test set can be reviewed and reused when comparing changes to a retrieval pipeline. For example, you can hold the test set constant while comparing retrieval-oriented embedding models such as Perplexity PPLX Embed v2.
Prerequisites
You need Python, Ragas, langchain-community, langchain-openai, an OpenAI API key, and a directory of Markdown files. The model setup below is the OpenAI option from Ragas's own documentation snippet, choose_generator_llm.md, reproduced as published. It defines the two variables every later step depends on: generator_llm, which writes the questions and enriches the knowledge graph, and generator_embeddings, which Ragas uses to relate chunks of your documents to one another.
Install Ragas and the OpenAI integration:
pip install ragas langchain-openai
Set the API key without writing it into your shell history:
read -s OPENAI_API_KEY
export OPENAI_API_KEY
Initialise the LLM and embeddings. The LLM is a LangChain chat model wrapped in LangchainLLMWrapper so Ragas can call it; the embeddings come from Ragas's own OpenAIEmbeddings class, built on the official openai client:
from ragas.llms import LangchainLLMWrapper
from langchain_openai import ChatOpenAI
from ragas.embeddings import OpenAIEmbeddings
import openai
generator_llm = LangchainLLMWrapper(ChatOpenAI(model="gpt-4o"))
openai_client = openai.OpenAI()
generator_embeddings = OpenAIEmbeddings(client=openai_client)
The documentation uses gpt-4o here. Question quality depends heavily on this model, since it decides what is worth asking and writes the reference answers, so treat a cheaper substitute as something to evaluate rather than a free saving. The same snippet has AWS, Azure and other provider options; whichever you choose, assign the results to the same two variable names and the rest of this guide runs unchanged.
1. Fetch the sample documents
Clone the sample Markdown collection from Hugging Face:
git clone https://huggingface.co/datasets/vibrantlabsai/Sample_Docs_Markdown
You can replace this directory with your own Markdown collection after confirming the workflow.
2. Load the Markdown files
Install the LangChain community package used by the source workflow:
pip install langchain-community
Load every Markdown file beneath the cloned directory. Each file becomes a LangChain Document containing page_content and metadata.
from langchain_community.document_loaders import DirectoryLoader
path = "Sample_Docs_Markdown/"
loader = DirectoryLoader(path, glob="**/*.md")
docs = loader.load()
3. Generate a test set directly from the documents
Create a TestsetGenerator with the model wrappers initialised in the prerequisites. testset_size=10 requests ten generated samples.
from ragas.testset import TestsetGenerator
generator = TestsetGenerator(llm=generator_llm, embedding_model=generator_embeddings)
dataset = generator.generate_with_langchain_docs(docs, testset_size=10)
Convert the result to a pandas DataFrame for inspection:
dataset.to_pandas()
If your input consists of LlamaIndex documents rather than LangChain documents, use generate_with_llama_index_docs instead of generate_with_langchain_docs.
How the generator works
The one-call method in step 3 hides two separate operations, and Ragas's documentation describes the pipeline in exactly those terms:
- Knowledge graph creation. Your documents become nodes in a
KnowledgeGraph, and a series of transformations enriches it: splitting documents into smaller pieces, extracting information from them, and adding relationships between related pieces using the LLM and the embedding model. - Test set generation. The enriched graph is used to build scenarios, each describing what a question should draw on, and the scenarios are turned into the final test samples.
Running the two stages yourself buys you two things. You can save the enriched graph and generate from it again without repeating the enrichment, which is where much of the LLM work happens. And you can see, and change, both the transformations and the mix of question types instead of accepting the defaults blind. Steps 4 to 6 do exactly that.
4. Build the knowledge graph explicitly
The direct generation method manages graph construction internally. To save and reuse the graph, create a KnowledgeGraph yourself:
from ragas.testset.graph import KnowledgeGraph
kg = KnowledgeGraph()
Before documents are added, the graph is empty:
KnowledgeGraph(nodes: 0, relationships: 0)
Append one DOCUMENT node for each loaded document. The node stores both the page content and the original document metadata.
from ragas.testset.graph import Node, NodeType
for doc in docs:
kg.nodes.append(
Node(
type=NodeType.DOCUMENT,
properties={"page_content": doc.page_content, "document_metadata": doc.metadata}
)
)
For the sample collection, the graph now contains ten document nodes and no relationships:
KnowledgeGraph(nodes: 10, relationships: 0)
5. Apply the default transformations
Use the same LLM and embedding model that generated the initial test set. default_transforms creates the transformation sequence, and apply_transforms applies it to the graph.
from ragas.testset.transforms import default_transforms, apply_transforms
# define your LLM and Embedding Model
# here we are using the same LLM and Embedding Model that we used to generate the testset
transformer_llm = generator_llm
embedding_model = generator_embeddings
trans = default_transforms(documents=docs, llm=transformer_llm, embedding_model=embedding_model)
apply_transforms(kg, trans)
Save the transformed graph as JSON, reload it, and display the result:
kg.save("knowledge_graph.json")
loaded_kg = KnowledgeGraph.load("knowledge_graph.json")
loaded_kg
The source workflow produced the following graph for the sample collection:
KnowledgeGraph(nodes: 48, relationships: 605)
These counts are outputs from the source example, not fixed requirements for other document collections. The jump from 10 nodes and no relationships to 48 nodes and 605 relationships is the transformations at work: documents were split into additional nodes, and the relationships record which pieces are related to which. Those relationships are what make multi-hop questions possible in the next step, because a multi-hop synthesizer needs connected pieces of information to combine. Because knowledge_graph.json holds the result of all that enrichment, keep it next to your test set: regenerating questions later from the saved file skips the transformation work entirely.
6. Generate queries from the saved graph
Create another TestsetGenerator, this time passing the reloaded graph through knowledge_graph:
from ragas.testset import TestsetGenerator
generator = TestsetGenerator(llm=generator_llm, embedding_model=embedding_model, knowledge_graph=loaded_kg)
Build the default query distribution:
from ragas.testset.synthesizers import default_query_distribution
query_distribution = default_query_distribution(generator_llm)
The source displays this distribution:
[
(SingleHopSpecificQuerySynthesizer(llm=llm), 0.5),
(MultiHopAbstractQuerySynthesizer(llm=llm), 0.25),
(MultiHopSpecificQuerySynthesizer(llm=llm), 0.25),
]
Generate ten samples with that distribution and convert the result to a DataFrame:
testset = generator.generate(testset_size=10, query_distribution=query_distribution)
testset.to_pandas()
default_query_distribution returns a plain list of (synthesizer, weight) pairs, and the weights set what share of the generated samples each one produces. With testset_size=10, the default mix aims for about half single-hop questions and a quarter each of the two multi-hop kinds. Because it is an ordinary list, you can pass your own to generator.generate to shift the balance, for example toward multi-hop questions if your retriever's weak point is combining information spread across documents. The three synthesizers serve different query patterns:
| Query synthesizer | Weight shown in the source | Query pattern |
|---|---|---|
SingleHopSpecificQuerySynthesizer |
0.5 | Targets specific information available through a single-hop query. |
MultiHopAbstractQuerySynthesizer |
0.25 | Produces abstract questions that combine information through multiple hops. |
MultiHopSpecificQuerySynthesizer |
0.25 | Produces specific questions that combine information through multiple hops. |
Review and retain the generated test set
Inspect the DataFrame before using generated samples in an evaluation. Confirm that each question is answerable from the intended source material and that its reference information represents the behaviour you want to test.
Keep knowledge_graph.json when you need to generate additional samples from the same transformed graph. Save the reviewed test set separately as part of your evaluation assets so that retrieval configurations can be compared against the same questions.
Expect to discard some samples. Synthetic questions can be ambiguous, can lean on details that only make sense to someone holding the source document, or can carry a reference answer that overstates what the text says. A short manual pass over ten samples is cheap; an evaluation built on bad questions quietly measures the wrong thing.
Generation alone does not score a retriever. Run the retriever against the reviewed questions, collect the retrieved contexts, and then apply the Ragas evaluation metrics appropriate to your pipeline.
Frequently asked questions
How do I generate a synthetic RAG test set with Ragas?
Load your source files as LangChain documents, initialise a Ragas TestsetGenerator with an LLM and embedding model, and call generate_with_langchain_docs. Set testset_size to the number of generated samples you want, then use to_pandas to inspect the result.
How do I save and reload a Ragas knowledge graph?
Call kg.save("knowledge_graph.json") after applying the graph transformations. Reload it with KnowledgeGraph.load("knowledge_graph.json") and pass the resulting graph to TestsetGenerator through the knowledge_graph argument.
What query types does the default Ragas distribution generate?
The distribution shown here assigns 0.5 to SingleHopSpecificQuerySynthesizer, 0.25 to MultiHopAbstractQuerySynthesizer and 0.25 to MultiHopSpecificQuerySynthesizer. These cover direct questions about specific information and questions that combine information from multiple parts of the graph.
Can Ragas generate test sets from LlamaIndex documents?
Yes. When the input consists of LlamaIndex documents, use generate_with_llama_index_docs instead of generate_with_langchain_docs and supply the corresponding Ragas model wrappers.
Why should I inspect a generated Ragas test set?
Synthetic queries should be checked against the underlying source material before they become regression cases. Export the test set with to_pandas, review its questions and reference information, and retain only samples suitable for your evaluation.
Related Guides
Mastering Advanced RAG Evaluation: From Basic Metrics to LLM-as-a-Judge
Learn how to systematically evaluate Retrieval-Augmented Generation (RAG) pipelines using the RAG Triad, Ragas, and LLM-as-a-Judge techniques.
Getting Started with LangChain in 2026
A comprehensive tutorial on building your first RAG application using the latest LangChain updates.

Build an Agentic RAG System with smolagents and ChromaDB
Use smolagents CodeAgent to run multi-step retrieval over a ChromaDB vector store, letting the agent plan its own search strategy.