GEPA: Optimizing prompts, code, and systems with LLM reflection and evolutionary search
Optimize prompts, code, and more with AI-powered Reflective Optimization
At a glance
- What is it?
- GEPA is an open-source Python framework for optimizing text artifacts such as prompts, code, agent architectures, and configurations through LLM powered reflection on execution traces and Pareto efficient evolutionary search. It requires 35x fewer evaluations than reinforcement learning methods.
- Who is it for?
- Adopt GEPA if you need to optimize prompts or code and want to reduce evaluation cost compared to reinforcement learning or manual iteration. The framework works best when you can measure success on a training set and provide diagnostic feedback (error messages, profiling data) to the optimization loop.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 2 days ago.
- What is it written in?
- Mainly Jupyter Notebook, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 28, 2026, and from our analysis. They are not legal advice.
Editorial analysis
Optimization through LLM reflection on execution traces
GEPA differs from traditional optimizers by analyzing why candidates fail, not just that they fail. Instead of collapsing execution results into a scalar reward, GEPA uses an LLM to read full execution traces such as error messages, profiling data, reasoning logs, and other diagnostic output. The LLM diagnoses the root cause and proposes targeted fixes.
The README shows the key results: GPT-4.1 Mini improves from 46.6% to 56.6% accuracy on AIME 2025 math problems with just 150 evaluations. Databricks reports achieving results that match Claude Opus 4.1 with open-source models, costing 90x less. On ARC-AGI, agent architecture discovery through GEPA improved accuracy from 32% to 89%.
Pareto-frontier selection and multi-objective optimization
GEPA maintains a Pareto frontier of candidates that excel on different task subsets, rather than optimizing for a single metric. This allows the framework to explore different trade-offs: one candidate might solve reasoning problems well but struggle with math; another might do the opposite. The Pareto-frontier approach discovers candidates that are optimal across different dimensions simultaneously, avoiding the trap of optimizing a single aggregated metric that might mask important failure modes. On each iteration, GEPA selects a frontier candidate, evaluates it on a minibatch, reflects on failures, mutates it, and accepts the mutation if it improves any part of the Pareto front. This can include system-aware merging, which combines strengths of two Pareto-optimal candidates excelling on different tasks.
This approach is 35x faster than reinforcement learning methods like GRPO, which require 5,000 to 25,000 evaluations. GEPA achieves similar or better results with 100 to 500 evaluations. The efficiency gain stems from reading full execution traces and diagnosing root causes, rather than treating optimization as a black-box reward signal.
Actionable side information: Diagnostic feedback from evaluators
The key concept enabling GEPA's efficiency is Actionable Side Information (ASI): structured diagnostic output from your evaluator that explains why a candidate failed. When you evaluate a prompt against a test case, you provide not only the pass/fail result but also error messages, profiling statistics, or reasoning traces.
The README example shows this in the optimize_anything API:
def evaluate(candidate: str) -> float:
result = run_my_system(candidate)
oa.log(f"Output: {result.output}") # Actionable Side Information
oa.log(f"Error: {result.error}") # feeds back into reflection
return result.scoreThe LLM reads these logs and uses them to inform the next mutation. Without ASI, optimization degrades to trial and error.
Installation and prompt optimization tutorial
Install GEPA from PyPI:
pip install gepaOptimize a system prompt in a few lines:
import gepa
trainset, valset, _ = gepa.examples.aime.init_dataset()
seed_prompt = {
"system_prompt": "You are a helpful assistant. Answer the question. "
"Put your final answer in the format '### <answer>'"
}
result = gepa.optimize(
seed_candidate=seed_prompt,
trainset=trainset,
valset=valset,
task_lm="openai/gpt-4.1-mini",
max_metric_calls=150,
reflection_lm="openai/gpt-5",
)
print("Optimized prompt:", result.best_candidate['system_prompt'])The optimize function takes a seed prompt, training and validation sets, the LLM to optimize (task_lm), and the LLM for reflection (reflection_lm). It returns the best candidate found. The framework includes examples for AIME math problems, ARC-AGI tasks, and blackbox optimization. The repository provides working tutorials and downloadable examples that demonstrate optimization on real datasets. You can also install from the main branch for the latest development version: `pip install git+https://github.com/gepa-ai/gepa.git`.
DSPy integration for AI pipeline optimization
The most powerful use of GEPA is within DSPy, which provides dspy.GEPA as an optimizer. You define a DSPy program with prompts you want to optimize, then compile it with GEPA:
import dspy
optimizer = dspy.GEPA(
metric=your_metric,
max_metric_calls=150,
reflection_lm="openai/gpt-5",
)
optimized_program = optimizer.compile(student=MyProgram(), trainset=trainset, valset=valset)DSPy handles metric computation and DSPy-specific optimizations, while GEPA drives the search. The integration allows you to optimize DSPy Predict modules, ChainOfThought modules, and custom programs without rewriting your evaluation logic. DSPy tutorials for GEPA are available on dspy.ai/tutorials, including executable notebooks for the AIME benchmark and general AI programs. This approach scales well for teams building AI systems that need continuous improvement on benchmark tasks.
optimize_anything: Beyond prompts to code and configurations
The optimize_anything API extends GEPA beyond prompts. You can optimize code snippets, agent architectures, scheduling policies, or any text artifact you can measure. You provide an evaluator function and GEPA handles the search.
The README mentions production examples: cloud scheduling policies that beat expert heuristics by 40.2%, coding agent resolve rates that improved from 55% to 82% on Jinja, and ARC-AGI agent architecture discovery. All of these use optimize_anything to evolve text artifacts guided by evaluation metrics and diagnostic feedback.
Agent skill integration for Claude Code and other agents
GEPA ships as an Agent Skill, so coding agents that read .claude/skills/ (Claude Code, Cursor, VS Code/Copilot, Codex, Gemini CLI) auto-discover it. You can also install it via plugin:
/plugin marketplace add gepa-ai/gepa
/plugin install gepa-optimize-anything@gepaThis allows agents to drive optimization for you without leaving your development environment. The agent can recognize when a prompt or piece of code needs optimization and invoke GEPA to improve it. The Agent Skill guide at https://gepa-ai.github.io/gepa/guides/agent-skill/ provides detailed instructions for setup and usage.
Production usage and framework requirements
GEPA has 50+ production uses across Shopify, Databricks, Dropbox, OpenAI, Pydantic, MLflow, and Comet ML, according to the README. The framework requires Python 3.10 to 3.14, with Python 3.14+ requiring alternate dependencies like mlflow-skinny instead of full MLflow. Core dependencies include litellm (capped at <1.92 due to compatibility constraints), tqdm for progress tracking, and cloudpickle for serialization. Optional integrations support datasets, MLflow, and wandb for experiment tracking. Install via `pip install gepa` or from the git repository for the latest development version at https://github.com/gepa-ai/gepa.git.
Because GEPA runs many evaluations, it is most practical for optimizations that can tolerate hours or days of compute time. It is not a real-time tool; it is a batch optimization framework for high-stakes decisions where the cost of the evaluation is worth the improvement in performance. The last push was 2026-09-28, with recent releases v0.1.4 (2026-07-15), v0.1.3 (2026-07-14), and v0.1.2 (2026-07-14). The MIT license makes it suitable for commercial and research use. Documentation is available at https://gepa-ai.github.io/gepa/ with guides, tutorials, blog posts, and a Discord community at https://discord.gg/WXFSeVGdbW.
Editorial conclusion
Adopt GEPA if you need to optimize prompts or code and want to reduce evaluation cost compared to reinforcement learning or manual iteration. The framework works best when you can measure success on a training set and provide diagnostic feedback (error messages, profiling data) to the optimization loop. Avoid it if you need sub-second iteration or have only a single evaluation criterion with no diagnostic data. Before starting, verify that your evaluation function can provide actionable side information (logs, errors, traces) that an LLM can diagnose.
Frequently asked questions
What does GEPA stand for?
GEPA stands for Genetic-Pareto. It combines genetic algorithms (iterative mutation and selection) with Pareto frontier optimization (maintaining candidates that excel on different task subsets).
How does GEPA prompt optimization work?
GEPA selects candidates from the Pareto frontier, evaluates them on test cases, uses an LLM to read error messages and diagnostic output and diagnose why they failed, mutates the prompt based on those diagnoses, and accepts the mutation if it improves the frontier.
What is Actionable Side Information in GEPA?
Actionable Side Information (ASI) is diagnostic feedback like error messages, profiling data, or reasoning logs that your evaluator returns alongside the pass/fail result. GEPA uses ASI to inform mutations and accelerate optimization.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/gepa-ai-gepa)