Library / SDK
atong01/conditional-flow-matching avatar
atong01/conditional-flow-matching

TorchCFM: five flow matching objectives in one small package

TorchCFM: a Conditional Flow Matching library

2,598 stars226 forksPythonMIT

At a glance

What is it?
A PyTorch library for conditional flow matching, extracted from two preprints on optimal transport CFM and simulation-free Schrodinger bridges, with example notebooks for 2D toy problems, images, single-cell data and tabular data.
Who is it for?
TorchCFM is small by design. The installable surface is a handful of loss classes, the test suite is a tests/ directory wired to pytest with doctest modules on, and the example work is organized by data type rather than by algorithm.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 82 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 10, 2026, and from our analysis. They are not legal advice.

Editorial analysis

A library carved out of two papers, not an application

The repository records itself plainly: TorchCFM, a Conditional Flow Matching library, with the description line reading 'Conditional Flow Matching for Fast Continuous Normalizing Flow Training.' The package on PyPI is `torchcfm`, the license is MIT, and the language is Python with 2587 stars, 223 forks and 42 open issues. The repository was last pushed on 2026-07-20 and is not archived.

What sits in the tree explains what kind of project this is. There is `torchcfm/` for the library, `tests/` for the test suite, `examples/` for notebooks, `runner/` for experiment plumbing, and three root files that carry most of the engineering decisions: `setup.py`, `pyproject.toml` and `requirements.txt`. There is also an `assets/` directory holding the README images, including a grid of generated samples and an animated GIF of eight Gaussians turning into two moons.

That framing matters for expectations. This is not a tool you point at a dataset and get a trained model. It is the reusable layer underneath two specific research efforts, and the README says so: it exists to spread flow matching methods within the machine learning community. If you want a paper result reproduced, this is the right shape. If you want a finished generative model pipeline, the notebooks supply the missing pieces by example rather than by API.

Five loss classes, one abstraction over the coupling

The `torchcfm` package exists so the conditional distribution `q(z)` can be swapped without rewriting the training loop. Everything else follows from that one decision. The README lists five loss functions, and the differences between them are differences in how the pairing between source and target samples is drawn.

`ConditionalFlowMatcher` treats `z = (x_0, x_1)` with an independent product of marginals, `q(z) = q(x_0) q(x_1)`. `ExactOptimalTransportConditionalFlowMatcher` uses an exact optimal transport joint instead, and the README notes it appears as OT-CFM in Tong et al. 2023a and as Batch OT in Poolidan et al. 2023. `TargetConditionalFlowMatcher` drops the pair entirely, working from `z = x_1`, which follows Lipman et al. 2023 and learns a flow from a standard normal to the data through conditional flows that transport the Gaussian to each datapoint. The README is careful to add that this does not make the marginal flow an optimal transport map.

The last two widen the choice. `SchrodingerBridgeConditionalFlowMatcher` substitutes an entropically regularized OT plan, usually approximated by a minibatch OT plan in practice, which is the SB-CFM and SF2M family. `VariancePreservingConditionalFlowMatcher` keeps independent marginals but uses conditional Gaussian probability paths with a trigonometric interpolation that preserves variance over time, following Albergo et al. 2023a.

That is a real API surface rather than a naming exercise. Someone who has read the flow matching literature can map their preferred formulation onto a class name without reading the trainer.

What the two preprints claim, and where the README stops

The homepage field points at arXiv 2302.00482, and two preprint badges sit at the top of the README. The first is 'Improving and generalizing flow-based generative models with minibatch optimal transport', which introduces Optimal Transport Conditional Flow Matching. The README's summary is that OT-CFM approximates the dynamical formulation of optimal transport, using both the static optimal transport plan and the probability paths and vector fields derived from it.

The second preprint is 'Simulation-free Schrodinger bridges via score and flow matching', arXiv 2307.03672, which proposes Simulation-Free Score and Flow Matching. The README describes SF2M as combining OT-CFM with score-based methods to approximate Schrodinger bridges, a stochastic form of optimal transport.

Read the README's own description of the method for context on the pitch: conditional flow matching is presented as a simulation-free training objective for continuous normalizing flows, one that allows conditional generative modeling and speeds up training and inference, with performance that closes the gap between CNFs and diffusion models. That last clause is a claim about the method family, and the repository documents it as a claim rather than presenting a benchmark table of its own.

Where the README runs out is anything about hardware, expected runtimes or reproduction budgets. It ends in the citation instructions, with a details block for BibTeX that the captured README cuts off mid-element. For the experimental details, the two papers are the source, not this repository.

Examples are filed by data type, starting with the 2D toys

The `examples/` directory has four subdirectories: `2D_tutorials/`, `images/`, `single_cell/` and `tabular/`. That split tells you what the authors expected people to bring to the library, and it is worth reading the two notebook names the README calls out.

The first is `examples/2D_tutorials/model-comparison-plotting.ipynb`, which produced the animated eight-Gaussians-to-two-moons GIF. The README is unusually candid about what that notebook shows: density, vector field and trajectories for several simulation-free CNF training schemes. It also reports a concrete negative result, that action matching with a 3x64 MLP and SeLU activations underfits under ReLU, SiLU and Swish activations in an external implementation, while fitting better under the training setup used here. One forward pass on the README is given as the action matching baseline. That kind of detail is more informative than a results table.

The second is `examples/images/mnist_example.ipynb`, an unconditional MNIST generation example covering both deterministic and stochastic generation. Both notebooks carry Google Colab badges in the README, which tells you they are expected to run in a hosted environment rather than on a local cluster.

The `single_cell/` and `tabular/` directories carry the heavier dependencies. Nothing in the README describes their notebooks in detail, so treat them as pointers: the layout tells you the library was applied beyond toy problems, and the files themselves tell you how.

Install surface: eleven dependencies and an optional extra

The dependency list is short enough to read in one screen, which is unusual for a generative modelling library. It comes straight from `install_requires` in `setup.py`:

python
install_requires = [
    "torch>=1.11.0",
    "matplotlib",
    "torchdyn>=1.0.6",
    "pot",
    "torchdiffeq",
    "pandas>=2.2.2",
]

The full list also includes numpy, scipy, scikit-learn and absl-py. Three of these matter more than they look. `torchdyn` carries neural ODE machinery and is pinned at 1.0.6 or newer because `requirements.txt` carries an inline note that 1.0.4 is broken on PyPI. `pot` is the Python optimal transport library, which is what makes the exact and minibatch OT couplings practical. `torchdiffeq` supplies the differential equation solvers used at sampling time. `pandas` is pinned at 2.2.2 or newer, and release 1.0.6 says why: it removed numpy and pandas pinning after version 2.

The example environment in `requirements.txt` is heavier and worth reading separately, since it adds torchvision, lightning-bolts, scprep, scanpy and clean-fid. The scprep and scanpy entries line up with the `single_cell/` examples, and clean-fid lines up with the image examples. There is also an extra: `extras_require` defines a `forest-flow` group pulling in xgboost, scikit-learn and ForestDiffusion, matching the ForestFlow example added in release 1.0.5.

So the library installs light and the examples do not. A reader planning to reproduce the MNIST notebook needs more than `pip install torchcfm` gives them.

Project hygiene is better than the release history suggests

The publishing record is thin and worth reading honestly. Three releases are recorded: 1.0.5 on 2023-11-27, 1.0.6 on 2025-03-05 and 1.0.7 on 2025-03-11. The bodies are minimal, and 1.0.7 carries only a full changelog link. 1.0.5 is the informative one, adding an updated CIFAR-10 configuration for an FID of about 3.6, a ForestFlow example and tests for torchcfm itself.

The engineering setup is a cut above that, though. `pyproject.toml` configures pytest with `testpaths` pointing at `tests/`, enables doctest collection on modules, marks slow tests through a registered `slow` marker and filters DeprecationWarning and UserWarning. It also sets coverage exclusions for the usual boilerplate lines, and configures ruff with a line length of 99 and a selective rule set. The root of the tree holds `.pre-commit-config.yaml`, a `CODE_OF_CONDUCT.md` and a `.github/` directory, and the README carries a CI badge for tests and another for code quality, plus a codecov badge.

So the pattern is a research codebase with maintainer tooling and very few tagged releases. The repository was pushed on 2026-07-20, roughly three months before the snapshot of these pages, so work is landing on `main` without version bumps. With 42 open issues on a project of this size, the open question is less whether the code works and more whether a change lands without version ceremony. For a research dependency that gets installed from a branch or a commit hash, that distinction is minor.

Editorial conclusion

TorchCFM is small by design. The installable surface is a handful of loss classes, the test suite is a tests/ directory wired to pytest with doctest modules on, and the example work is organized by data type rather than by algorithm. For anyone training a continuous normalizing flow on 2D toy data, MNIST, single-cell trajectories or tabular sets, the five objective classes are the whole reason to use it, and picking between them is a decision about the coupling rather than about the solver. The harder half of the work, the training loop, the noise schedule and the architecture comparison, lives in the notebooks and in the two preprints the repository reproduces. Code that pushes further than that has to be written by the reader, and nothing in the package pretends otherwise.

Frequently asked questions

What is flow matching and how does it work?

Flow matching trains a continuous normalizing flow by regressing a network onto a target vector field along a chosen probability path, instead of simulating an ODE and differentiating through it. TorchCFM implements this as simulation-free loss classes in which the only real choice is how source and target samples are paired.

Is flow matching better than diffusion?

The README claims conditional flow matching is faster to train and faster at inference than diffusion, and that its performance closes the gap between continuous normalizing flows and diffusion models. The repository documents this as a property of the method family and leaves the evidence in the two cited preprints rather than in a local benchmark.

Can you provide some examples of flow matching?

The repository ships notebooks under examples/2D_tutorials/, examples/images/, examples/single_cell/ and examples/tabular/. Two are named in the README: model-comparison-plotting.ipynb, which maps eight Gaussians onto two moons, and mnist_example.ipynb, which generates MNIST digits deterministically and stochastically.

Official sources

  1. atong01/conditional-flow-matching on GitHub
  2. License: MIT
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/atong01-conditional-flow-matching.svg)](https://hysenlabs.com/projects/atong01-conditional-flow-matching)