Automate Prompt Optimization with DSPy GEPA

September 28, 2026 • guides
LLMsPython

This guide is adapted from the DSPy GEPA optimization tutorial, available under the MIT licence.

Prompt engineering by hand is a losing game at scale. You tune a set of instructions for one model, and the next provider update breaks your assumptions. You move to a smaller, cheaper model and spend days rewriting prompts that used to work. GEPA, DSPy's reflective prompt optimizer, automates that cycle. It uses a separate "reflection" language model to study your scored examples, hypothesise better instructions, and iteratively replace the prompt in your DSPy program. You supply a metric; the optimizer does the rest.

Any team running LLM tasks in production where latency, cost, or consistency matter will find a use for this. The technique is particularly effective when you want to move a task from an expensive frontier model to a smaller, faster one without accepting a quality regression. DSPy's own tutorial cites Shopify moving a GPT-5 task to a small Qwen model optimized with GEPA. Small open-weight models keep getting more capable (see our report on ThinkingCap-Qwen3.8-27B cutting thinking tokens), and GEPA lets you adopt one without hand-tuning prompts for it.

GEPA itself runs on CPU; the only GPU you might need is for a locally-hosted student LM. With auto="light", optimization evaluates around six candidate prompts. You pay for API calls to both the student LM and the reflection LM, but the reflection LM is called only a handful of times per run rather than once per example, so using a large frontier model there stays affordable.


Prerequisites

  • Python 3.10 or later
  • dspy installed (pip install dspy) — GEPA is built in, no separate package required
  • spacy installed if you use the full haiku metric (pip install spacy)
  • haiku_metric.py downloaded from the public gist and placed next to your script
  • An OpenAI API key set as OPENAI_API_KEY, or equivalent credentials for your chosen provider
  • A DSPy program already written — this guide uses a haiku generator as the running example
  • A split dataset: train, val, and ideally a held-out test split

Step 1: Design a metric that gives feedback

Ordinary DSPy metrics return a scalar score. GEPA unlocks an extra channel: text feedback that the reflection LM reads when writing the next candidate prompt. Returning a dspy.Prediction with both a score and a feedback string is how you activate that channel.

The haiku task used throughout this guide penalises any output that names the season directly — a classical haiku should evoke, not announce:

def haiku_score_gepa(example, prediction, trace=None, pred_name=None, pred_trace=None):
    """
    Penalize verbatim use of the input season string.
    A haiku should evoke the season through imagery, not name it
    directly.
    """
    text = prediction.haiku.lower()
    if example.season.strip().lower() in text:
        return dspy.Prediction(
            score=0.0,
            feedback="Don't reference the input season verbatim."
        )
    return dspy.Prediction(score=1.0, feedback=None)

This pattern generalises far beyond haiku. A code-review metric can explain why a generated diff is unsafe. An extraction metric can flag which field was missing. Any note a human labeller would write in the margin of an annotation can become feedback here. The reflection LM aggregates those notes across failing examples and uses them to rewrite the prompt, so the richer your feedback strings, the more targeted the next candidate instruction will be.

For the full haiku metric — which also checks line count, 5-7-5 syllable structure, part-of-speech balance, article density, and tense — download haiku_metric.py from the public gist and place it next to your notebook. Import it as shown in Step 2.


Step 2: Configure the GEPA optimizer

With a metric in hand, instantiate the optimizer. Three decisions matter: which metric to use, which model acts as the reflection LM, and what budget to allow.

from haiku_metric import haiku_metric

reflection_lm = dspy.LM("openai/gpt-5.4")

optimizer = dspy.GEPA(
    metric=haiku_metric,
    reflection_lm=reflection_lm,
    auto="light",
    num_threads=2,
)

The reflection_lm is deliberately separate from the student LM your program uses at inference time. When optimising a small, cheap student model, you want a larger, more capable reasoner writing its instructions — that asymmetry is the point. The reflection LM is invoked only a handful of times per optimization run, so even a premium frontier model costs little in this role.

auto="light" sets the exploration budget. The three built-in levels come from AUTO_RUN_SETTINGS in DSPy's dspy/teleprompt/gepa/gepa.py. DSPy converts the candidate target into a number of metric calls using your train and validation sizes, so wall-clock time depends on your models and data rather than on the level alone:

auto value Candidate target (source) Relative cost Best for
"light" 6 Lowest Fast iteration, early-stage experiments
"medium" 12 Higher Production tuning once the metric is stable
"heavy" 18 Highest High-stakes tasks, final model selection

num_threads=2 controls how many examples are scored in parallel. Hosted inference providers impose rate limits, so keep this conservative until you know your account's quota.


Step 3: Compile the optimized program

With the optimizer configured, pass your DSPy program along with the training and validation splits:

optimized_haiku_bot = optimizer.compile(haiku_bot, trainset=train, valset=val)

Internally, compile runs an iterative loop. It executes your program on every training example with the current instructions and scores each result using your metric. It then passes those examples — with their scores and any feedback strings — to the reflection LM, which proposes revised instructions. The program is re-run with the revised instructions, scored again against the validation set, and whichever candidate instructions scored best are retained. This loop repeats until the auto budget is exhausted.

The two dataset splits serve distinct purposes. The trainset feeds the reflection LM examples from which to learn; it influences which instructions get written. The valset is used to rank candidate instructions against each other; it determines which instructions get kept. Keep a completely separate test set that touches neither split — that is the only honest measure of how the optimised program will behave in production.

DSPy's tutorial is explicit that there is no universal split ratio: keep as much data as possible in trainset, and make valset the smallest sample that still ranks candidates reliably, because every candidate is scored against all of it. If you omit valset entirely, GEPA reuses trainset for selection — acceptable for inference-time search where overfitting those examples is tolerable, but not for evaluating generalisation.


Step 4: Save and inspect the optimized program

After compile returns, persist the result immediately. DSPy serialises the full program state — including all optimised instructions for every predictor module — into a single JSON file:

optimized_haiku_bot.save("react_gpt_nano_haiku_optimized.json")

Opening that file is worth your time. Every Predict and ChainOfThought module in your program gets its own optimised instruction block, and reading them shows exactly what the reflection LM learned. The haiku program's synthesis step began with a bare docstring instruction:

Write a classical haiku given the provided inputs.

After compilation, GEPA replaced it with a structured multi-paragraph specification:

Write a classical haiku from three inputs:

Inputs:
- location
- season
- mood

Output requirements:
- Return only the haiku itself.
- Exactly 3 lines.
- No title, no labels, no explanation, no reasoning, no quotation marks.

Primary success criteria, in order:
1. Exact 5-7-5 syllable counts, one line per count.
2. Exactly 3 lines.
3. A concrete seasonal image or cue appropriate to the given season.
4. Do not repeat the input season or mood words verbatim.
5. Keep diction sparse and image-heavy, with strong noun/verb focus.

Haiku style requirements:
- Use a classical haiku approach: brief, image-centered, present tense, emotionally restrained.
- Evoke the location, season, and mood indirectly through concrete imagery rather than naming them outright.
- Build the poem around one small, observable moment tied to the location.
- Prefer concrete nouns and active present-tense verbs.
- Favor lexical density: most words should carry imagery or action.
- Keep adjectives very sparse; avoid piling on descriptors.
- Avoid abstraction, explanation, commentary, and explicit emotional naming.
- Do not use first-person pronouns.
- Keep article use minimal.

Location handling:
- Anchor the poem clearly in the given location with at least one concrete object, surface, sound, or visual detail from that place.
- If the location is unusual or man-made, pair one specific man-made image from the setting with one seasonal sign.

No human wrote that. The reflection LM inferred those rules directly from the failures and feedback strings in the training examples. Notably, the same program compiled against a different student LM will produce different instructions — the reflection LM tailors its output to what the student model responds to, which means the JSON is not transferable across student models.


What to watch out for

Metric leakage into the instructions. GEPA reads your metric's feedback strings and folds them into candidate prompts. If your feedback is too mechanical — for instance, always repeating the exact check that failed — the optimised prompt can end up listing your evaluation criteria verbatim. That overfits the metric rather than the underlying task, making the program brittle against metric changes.

Rate limits compound with num_threads. The student LM is called once per example per candidate prompt, every iteration. With 50 training examples, six candidates, and three iterations, that is 900 calls to the student endpoint. On free-tier or low-quota accounts this will hit rate limits. Start with num_threads=2 and a small dataset, then scale up.

A weak metric produces a confidently wrong prompt. GEPA is only as good as what you measure. If your metric does not capture a failure mode — say, it checks syllable count but not whether the output is actually in English — the optimised prompt will score well while still exhibiting that failure. Invest in the metric before you invest in the optimizer budget.

valset size affects exploration depth. Every candidate prompt is scored against every validation example. A large valset is expensive per candidate, which means fewer candidates are evaluated within the same wall-clock budget. If your valset runs to hundreds of examples, consider whether auto="light" will explore enough candidates to escape the local optimum your initial prompt occupies.

Do not edit the saved JSON by hand. The optimised instructions are embedded inside a structured program state. Manually tweaking the instruction text is fragile and will not reflect what the optimizer has learned about the full module configuration. If you want to iterate on the instructions, adjust your metric or feedback strings and rerun compile.


Where to go next

The optimizer mechanics described here — Pareto sampling, per-predictor feedback routing, auto budget translation, and the detailed_results audit trail — are covered in GEPA's in-depth documentation at the DSPy repository. DSPy's docs also include a choosing-an-optimizer guide that maps each built-in optimizer (MIPROv2, BootstrapFewShot, GEPA, and others) to the regime where it performs best.

Once you have a saved optimized program, the natural next question is when to rerun the optimizer — particularly when a new model is released. Because GEPA is fully automated, model evaluation becomes a single compile call rather than a multi-day prompt-engineering effort. That makes routine benchmarking of newly released models against your current solution practical, rather than a project in itself. The automation story fits naturally alongside the infrastructure improvements discussed in our AWS AgentCore v2 coverage, where cold-start elimination makes it feasible to trigger optimizer runs on demand in production pipelines.

Frequently asked questions

What is GEPA in DSPy and how does it differ from MIPROv2?

GEPA is DSPy's reflective prompt optimizer. It uses a separate reflection language model to read scored training examples — including text feedback strings — and write improved instructions for your DSPy program. MIPROv2 uses Bayesian search over instruction and few-shot combinations. GEPA's feedback channel lets domain-specific failure explanations directly shape the next candidate prompt, making it well-suited for tasks where you can articulate why an output failed.

How do I install DSPy GEPA?

GEPA is built into DSPy — install it with `pip install dspy`. There is no separate package. If you use the full haiku metric from the linked gist, download `haiku_metric.py` and place it next to your script, then install its dependency with `pip install spacy`.

What should I use as the reflection LM in DSPy GEPA?

The reflection LM should be more capable than the student model your program uses at inference time — that asymmetry is the point. It is called only a handful of times per optimization run, so a frontier model like `gpt-5.4` is affordable in this role even when your student is a small, cheap model.

How many examples do I need to run GEPA optimization?

DSPy's own GEPA tutorial says there is no universal split ratio. Keep as much data as possible in the trainset, which supplies examples for reflective prompt updates, and make the valset the smallest sample that still ranks candidate prompts reliably, since GEPA scores every candidate on it. If you omit valset, GEPA reuses the trainset for selection. Keep a separate test set as the honest measure of production performance.

Can I reuse optimized DSPy GEPA prompts across different student models?

No. The reflection LM tailors the optimised instructions to what the specific student model responds to. Saving a JSON compiled against `gpt-5.4-nano` and loading it with a different student will likely underperform because the instructions were written for that model's behaviour. Run a fresh `compile` call for each student model you want to evaluate.

What does the GEPA `auto` budget setting control?

`auto` sets the optimization budget. In DSPy's source (AUTO_RUN_SETTINGS in dspy/teleprompt/gepa/gepa.py) `"light"` targets 6 candidate prompts, `"medium"` 12 and `"heavy"` 18, and DSPy translates that into a number of metric calls from your dataset sizes. Wall-clock time depends on your models and data, so start with `"light"` for early experiments and raise it only for final model selection.

Related Guides