# Bespoke Curator: synthetic data pipelines for post-training and structured extraction

> Bespoke Curator is a Python library from Bespoke Labs for generating and curating synthetic datasets with LLM inference at scale. It is aimed at teams building post-training data or structured extraction pipelines, and its main trade-off is that it pulls in a large dependency tree, including LiteLLM and instructor.

**bespokelabsai/curator** — Synthetic data curation for post-training and structured data extraction

- Repository: https://github.com/bespokelabsai/curator
- Website: https://docs.bespokelabs.ai/bespoke-curator
- Stars: 1,735 · Forks: 147
- Language: Python
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/bespokelabsai-curator

## What Bespoke Curator is for, and who should care

The README describes Bespoke Curator as a way to create synthetic data pipelines, for two jobs: training a model, and extracting structured data. The topics list on the repository adds synthetic-data, synthetic-dataset-generation, instruction-tuning, and fine-tuning, which matches the examples directory: examples/bespoke-stratos-data-generation, examples/persona-hub, examples/ungrounded-qa, and examples/reannotation all generate training data, while examples/function-calling and examples/litellm-recipe-generation lean toward structured extraction.

The intended user is a Python engineer who already has a dataset and an inference budget, and who needs more rows of a specific shape. The library is not a labeling UI or a data-versioning system. It is the generation layer. If you are training a small model on a domain where no public dataset exists, or you need to turn a pile of documents into typed records, that is the target case. If you just want to call an LLM once from a script, the dependency weight is not worth it.

## How the generation and curation loop actually works

The mechanism visible in the repository is a Python function that maps an input row to an output row, wrapped by a Curator class that handles inference. The pyproject.toml dependencies tell you what sits underneath: litellm for provider routing, instructor for structured outputs, datasets for the row container, pydantic for the schema, xxhash for hashing, tiktoken for token counting, and aiofiles plus nest-asyncio for asynchronous file and event-loop handling. The README states there is built-in optimization for asynchronous operations, caching, and fault recovery.

So the data flow is: you declare a prompt and an output schema, Curator fans the requests out asynchronously through LiteLLM or vLLM, caches results so a re-run does not re-pay for identical work, and writes the result into a Hugging Face dataset. The viewer, exposed as the curator-viewer console script in pyproject.toml, reads that output while generation is in progress. The examples/viewer directory exists for that. Fault recovery matters here because a batch of thousands of requests will hit rate limits and transient errors, and the README claims recovery at every scale, though it does not spell out the retry policy in the README text itself.

## Installing Bespoke Curator and running a first curation job

The README gives one install command. It requires Python 3.10 or newer, per pyproject.toml, and the package on PyPI is named bespokelabs-curator.

```bash
pip install bespokelabs-curator
```

The Makefile shows the development path instead, which installs the code_execution and vllm extras plus dev dependencies through Poetry. Use that only if you are working on the library itself.

```bash
poetry install --extras "code_execution vllm" --with dev
poetry run pre-commit install
```

Once installed, the viewer is a separate console entry point. The pyproject.toml maps curator-viewer to bespokelabs.curator.viewer.__main__:main, so after installation you can start it from the shell.

```bash
curator-viewer
```

What you should see is the viewer process starting; the README's CLI animation is the only illustration of its output, and the documentation at docs.bespokelabs.ai/bespoke-curator is where the getting-started, tutorials, and how-to guides live. The README itself does not include a complete runnable generation snippet, so the first real job should be built from the tutorials rather than from the README alone.

## Where the dependency tree and versioning become a problem

The honest limitation is weight and version drift. pyproject.toml pins litellm to exactly 1.83.7 and pandas to 2.2.2, and requires anthropic ^0.84.0, vertexai 1.71.1, and mistralai ^1.5.1. Those are not loose constraints, and a project that already pins a different LiteLLM or pandas will have to resolve the conflict. The optional extras make this worse in a good way: vllm is optional, and tinker is optional and requires Python 3.11 or newer, so the finetune extra is unavailable on the 3.10 floor the package otherwise supports.

There is also a release mismatch worth checking. The latest GitHub release listed is v0.1.27 from 2026-03-15, while pyproject.toml on main declares version 0.1.29. That gap is normal for a project that publishes from main, but it means the tag and the installed package may not correspond, and the release notes for 0.1.28 and 0.1.29 are not published in the repository's release list. For a 0.x library, expect API movement between minor versions.

Finally, the wrong-tool case: if your extraction task fits a single prompt and a few hundred rows, a direct provider SDK call is simpler. Curator earns its place when the row count is large enough that caching and fault recovery save real money, or when you need the same pipeline to run against several providers.

## Bespoke Curator compared with writing your own LiteLLM loop

The obvious alternative is a hand-rolled script on top of LiteLLM, which Curator already depends on. The difference is what Curator adds around it: the dataset container, the caching layer keyed by xxhash, the viewer, the structured-output wiring through instructor, and the batch API integrations the README lists for OpenAI, Anthropic, and Gemini. If you write the loop yourself, you own retry logic, cache invalidation, and progress display, and you will probably rebuild a worse version of the viewer.

A second alternative, for fine-tuning specifically, is to skip curation and use a provider's managed dataset tooling. The README shows the opposite direction: the Tinker integration and the FireworksTrainer both take Curator output and push it into managed supervised fine-tuning, with the Fireworks example described as uploading curated data, training a LoRA, and sampling from the deployed model behind the same trainer interface as Tinker. That is the real argument for Curator over a bespoke loop: the output format is already what the trainers expect.

## Maintenance, licence, and what upgrading costs

The repository is not archived and the last push was on 2026-09-02, so it is being touched recently. The release cadence in the repository's release list is uneven: v0.1.25 in May 2025, v0.1.26 in July 2025, then v0.1.27 in March 2026. The README's What's New section runs from the January 2025 launch through a June 2026 Fireworks AI integration, which suggests the project is still adding provider support even when tagged releases lag.

The licence is Apache-2.0, declared in pyproject.toml and included as LICENSE at the repository root. Apache-2.0 is permissive and includes an explicit patent grant, which matters if you are embedding the library in a commercial training pipeline. Note that the licence covers Curator, not the models you call through it or the datasets you produce; those carry their own terms from whichever provider or source you use. That is a factual boundary, not legal advice.

Upgrade cost is dominated by the pinned dependencies. Moving from one Curator version to the next may force a LiteLLM or pandas bump, and the optional extras gate features behind Python versions. Check the extras you rely on before upgrading.

## Conclusion

Adopt Bespoke Curator if your team already has an inference budget and needs a Python pipeline that generates and curates synthetic data with structured outputs, batch APIs, and a viewer. Do not adopt it if you want a stable API surface: the version is 0.1.29 in pyproject.toml while the latest GitHub release is v0.1.27, so check that the two agree before pinning. Verify first that your Python version is 3.10 or newer, that your provider is one LiteLLM or vLLM supports, and that the optional vllm or finetune extras you need install cleanly on your platform.

## FAQ

### How do I install Bespoke Curator?

The README gives a single command, pip install bespokelabs-curator, and pyproject.toml requires Python 3.10 or newer. For development on the library itself, the Makefile uses poetry install with the code_execution and vllm extras.

### What is Bespoke Curator used for?

The README describes it as a way to create synthetic data pipelines, for training a model or extracting structured data. The repository topics list synthetic-data, instruction-tuning, and fine-tuning, and the examples cover data generation, function calling, multimodal input, and code execution.

### Does Bespoke Curator support batch APIs?

Yes. The README lists batch processing support for OpenAI, Anthropic, and other compatible APIs, and adds Gemini batch support. It states that batch processing cuts token costs in half.

### What is the licence for Bespoke Curator?

It is Apache-2.0, declared in pyproject.toml and shipped as LICENSE at the repository root. That covers the library, not the models you call or the datasets you generate.

### Is there a viewer for monitoring generation?

Yes. The README lists a viewer to monitor data while it is being generated, and pyproject.toml defines the curator-viewer console script pointing at bespokelabs.curator.viewer.__main__:main.

## Sources

- [bespokelabsai/curator on GitHub](https://github.com/bespokelabsai/curator)
- [License: Apache-2.0](https://github.com/bespokelabsai/curator/blob/main/LICENSE)
- [Project website](https://docs.bespokelabs.ai/bespoke-curator)
- [README](https://github.com/bespokelabsai/curator/blob/main/README.md)
- [Releases](https://github.com/bespokelabsai/curator/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/bespokelabsai-curator
