Model or dataset
vstorm-co/pydantic-deepagents avatar
vstorm-co/pydantic-deepagents

Pydantic Deep Agents: A Terminal Assistant and Python Framework with Live Run Forking

Open-source, self-hosted Claude Code - a terminal AI assistant and the Python framework behind it. Tool-calling, sandboxed execution, multi-agent teams, skills, checkpoints, unlimited context - on Pydantic AI, any model.

1,073 stars134 forksPythonMIT

At a glance

What is it?
pydantic-deep is both a self-hosted Claude-Code-style terminal assistant and a Python framework for building agents, running on Pydantic AI with any model provider. Its defining feature is Live Run Forking, which splits an in-flight agent run into parallel branches and uses an AI judge to pick the best outcome.
Who is it for?
pydantic-deep is the right pick for Python teams that want a self-hosted coding assistant they can also embed in their own agent pipelines, with the ability to run parallel solution branches and select the winner by test results. The MIT license and model-agnostic design mean you are not tied to a single provider.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 15 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

Two Products in One Repository

pydantic-deep ships as two things from the same codebase. The first is a terminal assistant: a TUI that runs in your terminal, plans tasks, edits files, executes shell commands, searches the web, maintains memory across sessions, spawns sub-agents, and connects to MCP servers. The README describes it as a self-hosted, open-source alternative to Claude Code that runs on any model. The key difference from Claude Code is model flexibility: pydantic-deep works with Claude, GPT, Gemini, and local models through Pydantic AI's provider abstraction.

The second is a Python framework. The same harness that powers the terminal assistant is accessible through a single function call, create_deep_agent(), which gives any Python program a filesystem, shell, planning, memory, sub-agents, sandboxed execution, MCP support, and unlimited context. The README targets developers who want to build their own assistant, research agent, or coding tool without rebuilding the infrastructure each time.

Both products share one unusual capability that the README states no comparable tool has: Live Run Forking, covered in its own section below. The repository also includes example files covering basic usage, Docker sandboxes, MCP server integration, custom tools, streaming, thinking-mode operation, human-in-the-loop flows, sub-agents, web tools, and skills.

Installation: CLI and Framework

The terminal assistant has a one-line install that handles uv and the CLI automatically:

bash
curl -fsSL https://raw.githubusercontent.com/vstorm-co/pydantic-deep/main/install.sh | bash
pydantic-deep

For Windows or manual installs, the README gives an alternative:

bash
pip install "pydantic-deep[cli]"

For use as a Python framework only, the base package is sufficient:

bash
pip install pydantic-deep

The framework entry point is create_deep_agent(). The README shows the minimal form:

python
from pydantic_deep import create_deep_agent

agent = create_deep_agent(model="anthropic:claude-sonnet-4-6")
result = await agent.run("Build a REST API for auth")

The package is on PyPI as pydantic-deep and requires Python 3.10 or later. The current release as of the last push on 2026-09-15 is version 0.3.43, based on the pyproject.toml.

Live Run Forking: The Core Differentiator

Live Run Forking lets an in-flight agent run split into parallel branches at a decision point. Each branch runs on a copy-on-write filesystem overlay: reads fall through to the parent state, and writes stay local to the branch. Each branch also has its own steering message and its own budget_usd cap.

The coordinator resolves the fork using one of four acceptance modes: manual (user picks), auto (AI judge picks automatically), auto_with_fallback (the default, which falls back to manual if the judge is uncertain), and vote. To opt into forking in the framework, pass one flag to create_deep_agent:

python
agent = create_deep_agent(
    model="anthropic:claude-sonnet-4-6",
    forking=True,
)

For test-driven selection, the README shows a LiveForkCapability that runs a real test command against every branch:

python
from pydantic_deep import LiveForkCapability

agent = create_deep_agent(
    forking=LiveForkCapability(test_command="pytest -q", test_timeout_s=120),
)

The confidence score is a weighted combination: quality spread at 0.4, test pass ratio at 0.4, and internal consistency at 0.2. In the CLI, /fork splits the current run, >>A and >>B steer individual branches, and /merge resolves the fork.

The README is direct that this feature adds cost. Each branch runs separately. Teams that want deterministic, single-path execution should set forking=False or omit the parameter entirely.

Architecture: Pydantic AI, Multi-Agent, and MCP

The framework is built on Pydantic AI, which provides type-safe structured output and model-agnostic tool calling. The pyproject.toml lists pydantic-ai-slim[web-fetch] as a core dependency, along with specialized packages for sub-agents, memory summarization, and capability shields (pydantic-ai-todo, pydantic-ai-backend, summarization-pydantic-ai, subagents-pydantic-ai, and pydantic-ai-shields).

The comparison table in the README positions pydantic-deep against Claude Code, Aider, LangGraph, and CrewAI across nine dimensions including sandboxed Docker execution, persistent memory, MCP servers, and type-safe structured output. According to the README's table, pydantic-deep is the only entry with all nine capabilities marked as first-class. The table is dated to 2026-06 and notes that corrections are welcome via PR.

Sandboxed Docker execution is available as an optional extra listed in pyproject.toml under the sandbox group, which depends on pydantic-ai-backend[docker]. YAML support for skills requires pyyaml, also an optional extra listed under the yaml group. The CLI extras bundle Textual (for the TUI), Typer, Rich, and prompt-toolkit. These optional groups let teams install only what they need and avoid pulling in Docker dependencies when the sandbox is not required.

The multi-agent swarm and message bus support means you can spawn sub-agents from within a running agent and have them communicate through a shared bus. The examples/subagents.py file in the repository demonstrates this pattern.

Limitations and When to Look Elsewhere

The package is classified as Beta in its PyPI classifiers. The CHANGELOG.md tracks changes from v0.3.24 onward; the frequency of minor version releases (three releases in the first week of August 2026 alone) suggests active iteration with possible breaking changes between minor versions. The pyproject.toml requires Python 3.10 or later, so teams on Python 3.9 cannot use this package without upgrading.

Live Run Forking is the feature most likely to create unexpected cost. The README's per-branch budget_usd cap provides a ceiling, but multiple parallel branches with long trajectories can multiply inference spend quickly. The AI judge adds one more call on top of that. Teams on tight inference budgets should test with a single branch before enabling forking. The confidence formula weights quality spread and test pass ratio equally at 0.4 each, with internal consistency at 0.2; a project with no automated tests cannot rely on the test_pass_ratio component.

The framework depends on a set of pydantic-ai auxiliary packages (pydantic-ai-todo, pydantic-ai-backend, summarization-pydantic-ai, subagents-pydantic-ai, pydantic-ai-shields) that are not from the main pydantic-ai repository. These are separate packages that may change independently of pydantic-deep's own version, which adds dependency management complexity compared to a framework with fewer third-party sub-package dependencies.

Aider is the most direct alternative for the terminal coding assistant use case. Aider covers file editing, git integration, and multi-model support without forking or the broader agent harness. It is a lighter dependency if the main need is interactive code editing rather than full agent orchestration. LangGraph is the closest alternative for the Python framework side: it offers a graph-based orchestration model that is well-documented but requires more upfront configuration than a single create_deep_agent() call.

Maintenance Status and License

The last push to the repository was on 2026-09-15. Recent releases include 0.3.43 on 2026-08-05, 0.3.42 on 2026-08-01, and 0.3.41 on 2026-08-01, indicating frequent patch releases. The repository has an active CI configuration at .github/workflows/ci.yml, a pre-commit setup, a GOVERNANCE.md, and a SECURITY.md. A Makefile documents the standard development workflow including install, format, lint, typecheck targets using uv.

The project is MIT-licensed, which allows commercial use, modification, and distribution with attribution. The installation script at install.sh is a curl-pipe-bash pattern. Users who prefer to review the installer before running it can fetch the script from the URL given in the README and inspect it before executing. The pyproject.toml marks the package as Development Status :: 4 - Beta, supports Python 3.10 through 3.13, and operates on Linux, macOS, and Windows as an OS-independent package.

The repository includes a CODE_QUALITY_REPORT.md and a SYSTEM_REVIEW.md in the root, which are uncommon files that may reflect the project's internal audit or quality tracking process. The examples/ directory covers a wide range of patterns including composite backends, filesystem backends, security gates, interactive chat, and streaming, giving developers concrete starting points beyond the minimal create_deep_agent() call.

Editorial conclusion

pydantic-deep is the right pick for Python teams that want a self-hosted coding assistant they can also embed in their own agent pipelines, with the ability to run parallel solution branches and select the winner by test results. The MIT license and model-agnostic design mean you are not tied to a single provider. The caveat is that Live Run Forking adds real cost: each branch runs a separate model session with its own budget_usd cap, and the AI judge adds another call. Teams that have no need for parallel exploration and want a lighter dependency should compare against Aider, which covers most editing workflows without the branching machinery. Check the CHANGELOG.md for the most recent stable release before adopting; the package is in Beta status per its PyPI classifiers, and the auxiliary pydantic-ai sub-packages introduce version coupling risk that is worth auditing before a production deployment.

Frequently asked questions

What is Pydantic Deep Agents?

Pydantic Deep Agents is both a self-hosted terminal AI assistant and a Python framework for building agents. The terminal assistant is a TUI that plans, edits files, runs commands, maintains memory across sessions, and connects to MCP servers on any model provider. The Python framework exposes the same harness through create_deep_agent() for embedding in your own code.

What are Pydantic AI Skills in this framework?

The README mentions skills as reusable, persistent work methods that the agent distills from repeated successful task patterns. YAML support for skills is available as an optional extra (pip install "pydantic-deep[yaml]"). The examples/skills/ directory and examples/skills_usage.py in the repository demonstrate how to define and use them.

How does Live Run Forking differ from standard agent runs?

A standard agent run follows one path. With forking enabled, the agent can split into parallel branches at a decision point, each with an isolated copy-on-write filesystem and its own steering message. An AI judge or test results then select the winning branch, whose history becomes the parent run's continuation. Each branch has its own budget_usd cap to control cost.

Does pydantic-deep work with models other than Claude?

Yes. The README states it works with any model supported by Pydantic AI, including Claude, GPT, Gemini, and locally hosted models. The model parameter in create_deep_agent() accepts any Pydantic AI compatible model identifier.

Official sources

  1. License: MIT
  2. Project website
  3. README
  4. Releases
  5. vstorm-co/pydantic-deepagents on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/vstorm-co-pydantic-deepagents.svg)](https://hysenlabs.com/projects/vstorm-co-pydantic-deepagents)