# prompt-ops: Meta's CLI for turning a working prompt into a Llama-optimized one

> prompt-ops is a Python package and CLI from meta-llama that rewrites an existing system prompt against your own query-response pairs. Here is how the migrate workflow is wired, what the YAML config controls, and where the tool stops being the right choice.

**meta-llama/prompt-ops** — An open-source tool for LLM prompt optimization.

- Repository: https://github.com/meta-llama/prompt-ops
- Stars: 1,036 · Forks: 138
- Language: Python
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/meta-llama-prompt-ops

## The problem prompt-ops targets: prompts written for one model family, run on another

A system prompt that behaves well on one LLM often degrades when you move it to Llama. The usual fix is manual: edit a sentence, rerun a handful of queries, judge the answers by eye, repeat. That loop does not scale past a few dozen cases, and it leaves no record of which wording change produced which improvement.

prompt-ops is built for that specific gap. It takes three inputs, an existing system prompt, a JSON file of query-response pairs, and a YAML configuration file, and produces a rewritten prompt plus performance metrics. The README frames the audience plainly: people with a prompt that already works elsewhere and a dataset they can afford to evaluate against. The repository ships use-cases/ and configs/ directories with worked examples, including a facility-support analyzer and a logical reasoning case called web-of-lies-pdo.

It is not a prompt editor, a registry, or a monitoring product. Nothing in the repository layout suggests versioning of prompts or runtime serving. The output is a prompt and a set of numbers.

## How the migrate pipeline is put together

The README's diagram is the clearest description of the data flow available. Three things enter: the existing system prompt, a set of query-response pairs, and the YAML configuration. They meet at a single step named prompt-ops migrate, and one optimized prompt comes out.

The dependency list in pyproject.toml explains what sits underneath. dspy is pinned exactly at 2.6.13, so the optimization and evaluation machinery comes from DSPy rather than from code written in this repository. litellm is a loose lower bound at 1.63.0 and acts as the unified client, which is why the .env.example file lists OPENAI_API_KEY, OPENROUTER_API_KEY and CEREBRAS_API_KEY side by side: the same pipeline can point at different providers without changing the calling code. pandas, numpy and scipy are present, which fits a workflow that tabulates per-example scores and compares configurations.

The README also mentions a newer direction. A paper on PDO, Prompt Duel Optimizer, describes a label-free method using dueling bandits and Thompson sampling, with results reported on BIG-bench Hard and MS MARCO. That is separate from the migrate workflow documented in the quick start, and the README points to use-cases/web-of-lies-pdo/ as the place to see it. Treat PDO as an additional method, not as a replacement for the config-driven pipeline.

## Installing prompt-ops from source and running a first optimization

The README recommends installing from source rather than from PyPI, stating that package names are in transition. The virtual environment is created with conda and Python 3.10, which matches the requires-python setting of >=3.10 in pyproject.toml.

```bash
conda create -n prompt-ops python=3.10
conda activate prompt-ops
git clone https://github.com/meta-llama/prompt-ops.git
cd prompt-ops
pip install -e .
```

After the editable install, the console script declared under [project.scripts] as prompt_ops.interfaces.cli:cli becomes available as prompt-ops. The README's next command scaffolds a project directory containing a sample configuration and dataset.

```bash
prompt-ops create my-project
cd my-project
```

Before running anything that calls a model, the API key goes into the .env file. The quick start shows OPENROUTER_API_KEY; .env.example also lists OPENAI_API_KEY and CEREBRAS_API_KEY, so whichever provider your configuration names needs its key present.

```bash
OPENROUTER_API_KEY=your_key_here
```

From there the README's workflow is a single optimization command against the generated config, followed by reading the optimized prompt and its metrics. The README does not document the exact flags of that run command in the text available, so check the generated config.yaml and the CLI help before assuming an argument name. What you should expect to see is a new prompt and a set of evaluation numbers, not a diff of individual edits.

## Where prompt-ops breaks down

The dataset requirement is the first real constraint. The README asks for as few as 50 query-response pairs, which sounds modest until you try to assemble them. If your task is new, you have no responses to score against, and the optimizer has no signal. The label-free PDO method is the project's answer to that case, but the README presents it through a paper and a single use case rather than through the main quick start, so a reader without labels is on a less-travelled path.

The pinned dspy==2.6.13 is a second constraint worth naming. An exact pin keeps behaviour reproducible, and it also means you cannot casually move to a newer DSPy release without checking whether the optimizer still works. This is a deliberate trade-off, not an oversight, but it belongs in your upgrade planning.

Evaluation cost is the third. Every optimization run calls a model provider through LiteLLM, and the README's own framing is minutes rather than seconds. There is no cached offline mode described in the README, so budget for API spend proportional to dataset size and the number of candidates explored.

Finally, the packaging situation is unresolved. The README says the PyPI install may have naming transition issues and is still on version 0.0.7, while pyproject.toml declares version 0.0.9 under the name prompt-ops. If you need a stable published artifact with a predictable name, the source install is the documented route.

## Compared with hand-tuning or a general DSPy setup

The obvious alternative is to use DSPy directly. prompt-ops is built on top of it, pinned to a specific version, and adds a CLI, a YAML configuration format, a project scaffold command, and a migrate step that takes an existing prompt as the starting point. If you are already comfortable writing DSPy programs and defining your own metrics in Python, the extra layer buys you convenience rather than capability, and you inherit the pin.

The other alternative is doing nothing structured: keep editing the prompt by hand. That is genuinely cheaper for a one-off task with twenty examples you can eyeball. It stops being cheaper the moment you need to explain why a prompt changed, or reproduce a previous result. prompt-ops produces metrics alongside the prompt, which is the part manual iteration rarely leaves behind.

A third option is a hosted prompt-optimization service. The difference in approach is where your data goes. prompt-ops runs locally as a Python package and sends only model calls to your chosen provider, with keys read from a local .env file. That matters if your query-response pairs contain anything you would rather not upload to a third-party optimization platform.

## Maintenance, licensing and upgrade cost

The repository is not archived, and the last push was on 2026-04-21. That is roughly five months before the date of writing, so the project is not abandoned, but there is no recent release entry to point at either. Plan for source installs rather than a steady stream of versioned packages.

The licence is MIT, declared both in the repository's LICENSE file and in the license field of pyproject.toml. MIT is permissive: you can use, modify and redistribute the code, including in commercial products, provided the copyright notice and permission notice are retained. That is the standard reading of the text, not legal advice, and if you are embedding the package in a product with unusual distribution terms, have counsel read the actual file rather than this summary.

The upgrade cost concentrates in two places. The exact dspy==2.6.13 pin means any move to a newer DSPy is a deliberate migration with testing, not a routine bump. The LiteLLM lower bound is looser, so provider behaviour can shift under you without a version change on your side. Neither is documented as a supported upgrade path in the README.

## Conclusion

prompt-ops suits teams that already have a system prompt and a labelled query-response set and want to stop hand-tuning wording for Llama models. It is the wrong tool if you have no evaluation data at all, since the README asks for as few as 50 examples and the optimizer has nothing to score against without them. Before adopting it, verify the PyPI situation yourself: the README recommends installing from source while package names transition, and the project metadata in pyproject.toml still declares the name prompt-ops at version 0.0.9.

## FAQ

### What is prompt-ops and who is it for?

prompt-ops is a Python package from meta-llama that automatically optimizes prompts for Llama models. It is aimed at people who already have a system prompt that works with another LLM and want it adapted using their own query-response examples.

### How do I install prompt-ops?

The README recommends installing from source: create a conda environment with Python 3.10, clone the repository, and run pip install -e . in the project root. It notes that the PyPI route may have naming transition issues.

### How much data does prompt-ops need to optimize a prompt?

The README asks for a JSON file of query-response pairs and states that as few as 50 examples are enough for evaluation and optimization.

### Which model providers can prompt-ops call?

It uses LiteLLM as a unified API client. The .env.example file lists OPENAI_API_KEY, OPENROUTER_API_KEY and CEREBRAS_API_KEY, and the quick start shows an OPENROUTER_API_KEY entry in the .env file.

### Does prompt-ops work without labelled examples?

The README describes PDO, Prompt Duel Optimizer, as a label-free method using dueling bandits and Thompson sampling, with a use case under use-cases/web-of-lies-pdo/. That path is presented through the paper and that use case rather than through the main quick start.

## Sources

- [Issues](https://github.com/meta-llama/prompt-ops/issues)
- [License: MIT](https://github.com/meta-llama/prompt-ops/blob/main/LICENSE)
- [meta-llama/prompt-ops on GitHub](https://github.com/meta-llama/prompt-ops)
- [README](https://github.com/meta-llama/prompt-ops/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/meta-llama-prompt-ops
