# flyto-core: a Python execution kernel that makes AI agent runs replayable

> flyto-core turns browser and API work into traced YAML recipes you can replay from a failed step. Here is what the package contains, how to install it, and where its boundaries lie.

**flytohub/flyto-core** — Flyto2 Core is the open-source execution kernel for automation and AI-agent workflows: 452 registry-backed modules, MCP-native transport, YAML recipes, evidence capture, replay, triggers, queue, versioning, and metering.

- Repository: https://github.com/flytohub/flyto-core
- Website: https://flyto2.com
- Stars: 481 · Forks: 84
- Language: Python
- License: Apache-2.0
- Published: 2026-08-08 · Updated: 2026-08-18 · Language: en
- Canonical page: https://hysenlabs.com/projects/flytohub-flyto-core

## The problem flyto-core addresses: agent work with no proof and no restart point

An AI agent that drives a browser or calls an API typically produces a final answer and very little else. When the answer is wrong, there is no per-step record of what happened, and no way to re-run only the part that broke. The README frames this bluntly: with a shell script you re-run the whole thing. flyto-core is the execution layer that sits under that work and records it.

The intended audience is narrow enough to state plainly. It is for engineers who already have a workflow that consists of concrete steps (open a page, extract values, call an endpoint, write a file) and who want those steps executed by a reviewed module rather than by code an agent generated at runtime. The README describes the boundary as MCP-compatible clients calling reviewed flyto-core modules through schemas, not arbitrary generated production code. That is the design bet: constrain the agent to a registry of known operations so the run can be traced and repeated.

It is not a general-purpose orchestration platform. The README states that flyto-core does not own intent or provider governance, procedure learning or scoring, or hosted product and account logic. Those belong to two sibling packages, flyto-ai and flyto-blueprint. flyto-core is standalone: it validates schemas, executes, replays, and emits evidence without either of the others installed.

## How execution works: registry modules, YAML recipes, trace and replay

The unit of work is a recipe, a YAML document with a name and an ordered list of steps. Each step names a module and passes params. The README's competitive-pricing example shows the shape: a browser.launch step, a browser.goto step whose params reference a url variable, a browser.evaluate step carrying a JavaScript snippet, then browser.screenshot and browser.viewport steps. Steps are addressed by id, which is what makes selective replay possible.

Modules come from a registry. The README says the current public inventory is 480 registry-backed modules across 88 catalog categories, and that docs/TOOL_CATALOG.md is generated from ModuleRegistry rather than hand-counted. That distinction matters when you evaluate the project: the catalog is a build artifact of the registry, so the number moves with the code rather than with someone's editing. Categories named in the README include triggers, queue modules, workflow versioning, metering hooks, browser automation, API calls, data transforms, verification, files, and crypto.

The replay mechanism is the part worth understanding before you adopt. After a run, flyto replay --from-step 8 re-executes step 8 and later against preserved context from the earlier steps. The README claims steps 1 through 7 are instant on replay and that full context is preserved. Whether that holds for your workflow depends on whether your early steps have side effects, since the preserved context is not the same thing as a rolled-back world. The README does not document rollback of side effects, so treat replay as a restart optimization, not a transaction.

Evidence and verification are first-class in the module list. The README names verification.* modules and describes evidence capture as part of the package's job, alongside trace and per-step timing. The sample run output shows each step with a status mark and a duration in milliseconds, plus notes such as saved intel-desktop.png or Web Vitals captured.

## Installing flyto-core and running a first recipe

The README gives a three-command install. The base package brings the core engine, the CLI, and an MCP server; the browser extra adds Playwright, and Chromium is installed separately as a one-time setup step.

```bash
pip install flyto-core            # Core engine + CLI + MCP server
pip install flyto-core[browser]   # + browser automation (Playwright)
playwright install chromium        # one-time browser setup
```

After that, the README's 30-second path is a single recipe invocation. The competitor-intel recipe takes a url parameter:

```bash
flyto recipe competitor-intel --url https://github.com/pricing
```

According to the README's sample output, you should see a numbered step list with a status mark and a duration per step, ending with screenshots written to disk (intel-desktop.png and intel-mobile.png in the example), Web Vitals captured, and a JSON report saved. The run shown covers twelve steps, including two viewport changes and two screenshots.

Two more recipes are listed for a first pass. full-audit takes a url and covers SEO, accessibility, and performance. scrape-to-csv takes a url and a selector and exports to CSV:

```bash
flyto recipe full-audit --url https://your-site.com
flyto recipe scrape-to-csv --url https://news.ycombinator.com --selector ".titleline a"
```

The README points to 41 built-in recipes in docs/RECIPES.md. If a run fails partway, the documented recovery is the replay command with a step number:

```bash
flyto replay --from-step 8
```

There is also a Python-level entry point in the repository root, demo.py, and a run_demo.sh script, though the README does not document their arguments.

## The package is classified as beta, and the version floor moved for a reason

The first thing to check before adopting is the classifier in pyproject.toml: Development Status :: 4 - Beta. That is the project's own label. It sits alongside a version number in the 2.x line, which can read as more settled than beta implies. Take the classifier at face value.

The Python floor is also worth reading. A comment in pyproject.toml records that the floor previously said 3.9 while the base dependency set could not resolve there at all, because aiohttp>=3.14.3 requires Python 3.10 or later. pip would accept the package on 3.9 and then fail to find a distribution. The comment notes that CI only ran 3.11, so nothing caught it. The floor now matches what the package actually needs, and the classifiers list 3.10 through 3.13. This is a small, honest piece of project history, and it tells you something about the project's testing surface: the gap existed because the CI matrix was narrow.

The module count is the other constraint. 480 modules across 88 categories is a large surface to learn, and the README does not claim they are uniform in maturity. Nothing in the README says every module is covered by tests, and nothing gives a per-module stability rating. The generated catalog tells you what exists, not what is safe to depend on.

Finally, the JavaScript side. The repository carries a package.json named flyto2-core-test-runtime, marked private, whose only stated purpose is declaring a JavaScript runtime for browser-contract tests, with jsdom as a dev dependency and Node engine constraints. That is test infrastructure, not a Node API for the project. If you arrived expecting a JavaScript client, the README does not describe one.

## When flyto-core is the wrong tool

If your workflow is a pure in-process data transformation with no browser, no network call, and no external side effect, the trace and replay machinery buys you little. You are paying for a registry lookup, a YAML layer, and evidence emission around code that a dozen lines of Python would do directly. The README's own framing supports this reading: the pitch is aimed at jobs that are annoying because they touch a page or an endpoint, not at computation.

Replay is the second place to be careful. The README documents replay from a step number and says context is preserved, but it does not document rollback. If step 4 sent an email or created a record, replaying from step 8 does not undo it, and nothing in the README suggests the engine tracks that. Design your recipes so that side-effecting steps come after the steps you expect to fail, or accept that you will clean up manually.

The third boundary is governance. If what you actually need is routing across model providers, or scoring and reuse of learned procedures, flyto-core explicitly does not own that. The README assigns provider governance to flyto-ai and procedure scoring to flyto-blueprint. Installing flyto-core alone will not give you either, and the README does not describe how the three packages interoperate beyond the table of responsibilities.

## How flyto-core differs from writing Playwright scripts directly

The README makes the comparison itself, and it is the most useful one available. The left column of its table is an 85-line Python program using playwright.async_api: launch a browser, open a page, run page.evaluate to pull pricing text out of elements matching a class substring, take a desktop screenshot, and so on. The right column is the same work as a YAML recipe with a dozen steps.

The difference is not the amount of code. It is what the code leaves behind. A hand-written Playwright script produces the artifacts you remembered to write. The README's characterization of that column is: no trace, no replay, no timing, and if step 5 fails you re-run everything. The recipe version, by the project's account, produces a full trace, per-step timing, and the ability to replay from any step.

That trade is real in both directions. You give up the full expressiveness of Python inside the workflow, since steps are module calls with params rather than arbitrary code. In exchange you get a structure the engine can address step by step, which is what makes selective replay and evidence capture possible at all. If your workflow needs control flow or data handling that the module registry does not cover, the recipe format will feel like a cage, and the README does not describe an escape hatch to inline Python.

A second alternative worth naming is a general workflow orchestrator such as Airflow or Prefect. Those schedule and retry tasks across a cluster; they do not ship a browser module registry or emit per-step screenshots and Web Vitals. The overlap is orchestration vocabulary, not capability.

## Maintenance, licensing and the cost of upgrading

The repository is not archived, and the last push was on 2026-08-14, the same day as the v2.28.1 release. The two preceding releases, v2.28.0 and v2.27.0, landed on 2026-08-13 and 2026-08-08. The release cadence visible in the release list is roughly weekly in that window, with patch and minor versions rather than a long-stable major line. The pyproject.toml in the repository declares version 2.32.0, ahead of the newest release listed, which suggests the file tracks development rather than the published artifact.

Upgrade cost is driven by the module surface. A minor release can change behavior in a specific area without touching the rest: v2.28.0 is described as IPv6 SSRF hardening and extension management, v2.27.0 as boundary coverage enforcement. Those are narrow, security-relevant changes. If your recipes depend on network modules, the SSRF hardening in particular is the kind of change that can alter which URLs a step will accept, and the release notes are the place to check before upgrading. The README does not document a deprecation policy or a compatibility window for recipes across minor versions, so pin your version and read CHANGELOG.md before moving.

Licensing is Apache-2.0, declared both in the README badge and in pyproject.toml under license and license-files, with LICENSE and NOTICE listed as the license files. Apache-2.0 is permissive and includes an explicit patent grant, which is usually what a company wants for a dependency embedded in internal tooling. The NOTICE file means there is attribution material to carry if you redistribute. That is the extent of what the repository supports; whether your specific redistribution triggers the NOTICE obligation is a question for your own counsel, not for this article.

## Conclusion

Adopt flyto-core if you need deterministic, traced execution of browser and API steps and you are comfortable with a beta-classification package whose module surface is large enough that you will spend real time in the catalog. Do not adopt it if you need a stable API contract today or if your work is pure in-process data transformation with no browser or network step to trace. Before committing, run the competitor-intel recipe against a page you control, inspect the emitted evidence, and confirm that replay from a mid-recipe step produces the same output as the original run.

## FAQ

### What is flyto-core used for?

It is an execution kernel for automation and AI-agent workflows: it validates schemas, runs steps from a registry of 480 modules, and emits trace, evidence and replay data. Typical uses named in the README are browser automation, API integration, web scraping, MCP server automation and Web Vitals checks.

### How do I install flyto-core?

Install the base package with pip install flyto-core for the engine, CLI and MCP server, or pip install flyto-core[browser] to add Playwright, followed by playwright install chromium as a one-time browser setup.

### Can flyto-core replay a workflow from a failed step?

Yes. The README documents flyto replay --from-step 8, which re-executes from that step onward with context from the earlier steps preserved. The README does not document rollback of side effects produced by earlier steps.

### Does flyto-core need flyto-ai or flyto-blueprint to run?

No. The README states that flyto-core is a standalone execution package that does not require the other two to validate and run a procedure or produce evidence. It also does not provide their intent, provider governance or procedure scoring features.

## Sources

- [Official documentation](https://flyto2.com)
- [Official README](https://github.com/flytohub/flyto-core#readme)
- [Project repository](https://github.com/flytohub/flyto-core)
- [Release notes](https://github.com/flytohub/flyto-core/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/flytohub-flyto-core
