# PrefectHQ/marvin: structured LLM output and agentic tasks in Python

> Marvin is an Apache-2.0 Python framework that turns unstructured input into typed Python objects and lets you chain agents into tasks and threads. It is a good fit when you already use Pydantic and want validated results, and a poor fit when you want a hosted platform or a non-Python runtime.

**PrefectHQ/marvin** — an ambient intelligence library

- Repository: https://github.com/PrefectHQ/marvin
- Website: https://marvin.mintlify.app
- Stars: 6,198 · Forks: 415
- Language: Python
- License: Apache-2.0
- Published: 2026-09-09 · Updated: 2026-09-09 · Language: en
- Canonical page: https://hysenlabs.com/projects/prefecthq-marvin

## The problem Marvin solves, and who actually needs it

Most LLM calls return a string. Application code then needs a number, an enum, a list of records, or an object with named fields. The usual workaround is prompt engineering plus a parser plus retry logic, and that parser is where projects quietly rot, because the model's output shape drifts and the parser was written for last month's prompt.

Marvin's answer is to make the target type the interface. You pass a Python type, and the library handles the prompt construction and the parsing. The README shows four top-level utilities for this: marvin.extract pulls native types out of unstructured text, marvin.cast converts text into a structured type, marvin.classify maps text onto a predefined label set, and marvin.generate produces a requested number of objects from a description.

The second problem is orchestration. A single call is easy; a workflow with a research step, a classification step, and a summarization step is not. Marvin 3.0, described in the README as ported from ControlFlow, adds Agent, Task, and thread abstractions on top of the same typed-result foundation.

This is aimed at Python engineers building internal tools, data pipelines, and support automation, not at people who want a chat UI. If you are writing TypeScript or Go, there is nothing here for you; the package is Python 3.10 or newer per pyproject.toml.

## How the typed-output path works under the hood

The dependency list in pyproject.toml is the clearest statement of the architecture. pydantic-ai is a direct dependency, so Marvin is a layer over that library rather than its own model client. pydantic and pydantic-settings handle validation and configuration, partial-json-parser suggests streaming or incomplete JSON is parsed rather than discarded, and jinja2 handles prompt templating.

There is also a persistence layer: aiosqlite, sqlalchemy[asyncio], and alembic, plus a migrations directory and alembic.ini at the repository root. That means tasks and their results are stored in a local SQLite database by default, which is what makes the observable task history in the README's shell example possible. The console output showing an agent ID, a tool call, and a status is rendered by rich, another dependency.

So the data flow is: your type and instructions become a prompt, pydantic-ai calls the model, the response is parsed into your type, validated, and the run is recorded through SQLAlchemy into SQLite. The README notes that Marvin uses OpenAI by default but natively supports all Pydantic AI models, which follows directly from pydantic-ai being the underlying client.

One consequence worth naming: because alembic is in the dependency set, schema migrations for the local database are part of the package, and the wheel build config includes migrations and alembic.ini. Upgrading Marvin can therefore involve a database migration, not just a library swap.

## Installing Marvin and getting a first typed result

The README gives one install command, using uv. The justfile confirms uv is the expected tool for development as well, with a check-uv recipe that fails with install instructions if uv is missing.

```bash
uv add marvin
```

The README then shows that you configure a provider with an environment variable. OpenAI is the default provider, so the minimum setup is an API key.

```bash
export OPENAI_API_KEY=your-api-key
```

With that in place, the smallest useful thing is casting free text into a TypedDict. The README uses a location example; note that the printed values are the README's own illustrative output, not a guarantee of the numbers you will get.

```python
from typing import TypedDict
import marvin

class Location(TypedDict):
    lat: float
    lon: float

result = marvin.cast("the place with the best bagels", Location)
print(result)
```

Classification is the same shape but returns one of your enum members, which is the pattern most teams reach for first when routing support tickets or tagging records.

```python
from enum import Enum
import marvin

class SupportDepartment(Enum):
    ACCOUNTING = "accounting"
    HR = "hr"
    IT = "it"
    SALES = "sales"

result = marvin.classify("shut up and take my money", SupportDepartment)
print(result)
```

The README states this prints SupportDepartment.SALES. If you want to see a runnable version of any of these before writing your own, the repository ships examples/hello_cast.py, examples/hello_classify.py, examples/hello_extract.py, and examples/hello_generate.py.

## Tasks, agents and threads: when the single call is not enough

marvin.run is the entry point the README presents for a bare task. It takes a string, and it optionally takes a result_type, so the same call can return prose or a typed value.

```python
import marvin
answer = marvin.run("the answer to the universe", result_type=int)
print(answer)
```

The README states this prints 42. That example is worth pausing on, because it shows the trade-off clearly: the type constrains the output format, not the correctness of the content.

When you need context or tools, you construct a marvin.Task directly. The README's example passes instructions, a result_type of IPvAnyAddress from pydantic, a list of tools, and a context dictionary, then calls .run(). The console output it shows includes a tool invocation and a callback named MarkTaskSuccessful, which is the mechanism by which the agent signals completion rather than the framework guessing.

Agents are the specialization layer. The README's example constructs an Agent with a name and instructions and calls .run() on it, which is how you keep a consistent persona or a domain constraint across many calls instead of restating it in every prompt. Threads, described in the README as a way to compose tasks into a customizable thread, are the orchestration layer above that.

The honest summary: marvin.run is a convenience wrapper, Task is the unit you should build on, and threads are where the design gets opinionated. The README describes threads but does not show a full thread example, so treat that layer as the one to prototype before you commit to it.

## The tool-calling warning is the real limitation

The README carries a warning directly above its shell example: the example produces type safe results but runs untrusted shell commands. That is not boilerplate. It is the framework telling you that a tool is a Python function the agent decides to call, with arguments the agent decides, and the type annotation on the return value does nothing to constrain the arguments.

In the example, the tool is a function that calls subprocess.check_output on a list of strings, and the agent chose ipconfig getifaddr en0 on macOS. In a container with no sensitive files, that is fine. In a process that holds database credentials or has network access to internal services, the same pattern is a remote code execution path with an LLM as the decision maker.

There is no allowlist, sandbox, or confirmation step visible in the README's tool example. If your threat model requires one, you are building it yourself, and you should assume the agent will occasionally pass arguments you did not anticipate.

The second limitation is API churn. The README states that Marvin 3.0 introduces a new way to work with AI and that the structured-output utilities are the ones from marvin 2.x, kept at the top level of the package. That is a compatibility bridge, not a stability promise. The release history shows v3.2.5 in January 2026, v3.2.6 later that month, and v3.2.7 in March 2026, with the last push to the repository on 2026-09-09. Code written against a 3.2.x minor version should be pinned.

## Alternatives, and how they differ in approach

The most direct comparison is the library Marvin is built on. Pydantic AI is a direct dependency in pyproject.toml, and the README says Marvin natively supports all Pydantic AI models. If your need is a single typed model call, using Pydantic AI directly removes a layer and a migration surface. What you give up is the task and thread orchestration, the local run history, and the four convenience utilities.

LangChain and its ecosystem take the opposite approach: a large set of integrations and abstractions for chains, retrievers, and agents, assembled from many packages. Marvin is narrower and more opinionated, with a fixed dependency set and a task-centric model rather than a graph of composable components. If you want to plug in ten vector stores and four tracing backends, Marvin is the wrong shape.

Outlines and similar constrained-decoding libraries solve the structured-output problem at the token level, masking logits so the model literally cannot emit an invalid token. Marvin solves it at the validation level, parsing and validating after generation. The constrained-decoding approach gives stronger format guarantees but ties you to specific model serving setups; Marvin's approach works across whatever pydantic-ai can reach.

The practical question is whether you want a framework or a client. Marvin is a framework, and frameworks cost you upgrade work.

## Licence, maintenance and upgrade cost

Marvin is Apache-2.0, with the licence referenced from pyproject.toml as a file rather than an SPDX string. Apache-2.0 is permissive and includes an explicit patent grant, which matters if you are shipping this inside a commercial product. It is not a copyleft licence, so it does not force you to publish your own code. That is a description of the licence text, not legal advice; your counsel decides what your obligations are.

On maintenance, the facts are these: the repository is not archived, and the last push was on 2026-09-09. The most recent tagged release listed is v3.2.7 from 2026-03-04. So there is a gap between the last release and the last push, which is normal for a project that merges to main between tags, but it also means the code you get from main is ahead of the code you get from PyPI.

The upgrade cost is higher than a typical library because of the persistence layer. alembic and the migrations directory ship inside the wheel, and the build config explicitly includes migrations and alembic.ini. Any upgrade that changes the schema requires a migration against your existing SQLite database. Pin the version, read the release notes for the version you are moving to, and back up the database file before running an upgrade in anything you care about.

Development setup, per the justfile, is uv sync for dependencies, uv run pre-commit run --all-files for checks, and cd docs && mintlify dev for the documentation site.

## Conclusion

Adopt Marvin if your application is already Python and Pydantic shaped, and you want typed results from an LLM without writing your own validation layer. Do not adopt it if you need a hosted service, a non-Python runtime, or a stable API surface for long-lived code, because the README itself notes that 3.0 reorganized the package and the release cadence shows ongoing churn. Before committing, verify two things in your own environment: that the model provider you intend to use is reachable through pydantic-ai, and that the tool-calling path in marvin.Task is acceptable for your threat model, since the README's own shell example carries a warning about running untrusted commands.

## FAQ

### How do I install PrefectHQ/marvin?

The README gives a single command, uv add marvin, and notes the package is on PyPI. You then configure a provider with an environment variable such as OPENAI_API_KEY, since OpenAI is the default.

### Does PrefectHQ/marvin work with models other than OpenAI?

Yes. The README states that Marvin uses OpenAI by default but natively supports all Pydantic AI models, and pydantic-ai is a direct dependency in pyproject.toml.

### What is the difference between marvin.run and marvin.Task?

marvin.run is the simplest way to run a task and accepts an optional result_type. A marvin.Task is defined explicitly with instructions, a result_type, tools and context, and is run by a default agent when you call .run().

### Is it safe to give a Marvin task a shell tool?

The README places a warning above its shell example stating that while the example produces type safe results, it runs untrusted shell commands. No allowlist or sandbox around tool arguments is shown.

## Sources

- [License: Apache-2.0](https://github.com/PrefectHQ/marvin/blob/main/LICENSE)
- [PrefectHQ/marvin on GitHub](https://github.com/PrefectHQ/marvin)
- [Project website](https://marvin.mintlify.app)
- [README](https://github.com/PrefectHQ/marvin/blob/main/README.md)
- [Releases](https://github.com/PrefectHQ/marvin/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/prefecthq-marvin
