Model or dataset
ByteDance-Seed/VeOmni avatar
ByteDance-Seed/VeOmni

VeOmni: a recipe zoo for training any modality, and a headline principle the paragraph under it does not keep

VeOmni: Scaling Any Modality Model Training with Model-Centric Distributed Recipe Zoo

2,225 stars280 forksPythonApache-2.0

At a glance

What is it?
VeOmni is ByteDance Seed's training framework for text, vision-language, mixture-of-experts, omni and video diffusion models, built on PyTorch's own distributed primitives with per-model recipe files and support for three accelerator vendors. Its most interesting claim is that you should not use a structured trainer, and the same paragraph that makes that claim also tells you it ships two of them.
Who is it for?
VeOmni is worth evaluating if you train more than one kind of model, because a recipe zoo is the right abstraction when the interesting decisions are per model rather than per training run, and the configuration layout mirrors the modality split so finding the right recipe is a directory listing rather than a documentation hunt. Three accelerators supported in one framework is also rarer than it sounds and matters if your hardware is not the default.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 3 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 28, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The headline says trainer-free, and the paragraph beneath it ships two trainers

The second of the framework's stated principles is the interesting one, and it is worth reading closely because the bold label and the sentence under it do not say the same thing. The label is trainer-free. The explanation is that the framework supports linear training scripts which avoid rigid, structured trainer classes, and it names the two it means, one general-purpose training library and the trainer class in the most popular model library. The stated reason is transparency: the scripts expose the entire training logic to the user for maximum control. The claim is real and it is the right thing to object to. A trainer class decides when a batch moves, where logging goes, what happens on an exception, and how a checkpoint is written, and for research code those are exactly the decisions you want to make yourself. But the same paragraph then says the framework also supports a basic trainer for text-only or vision-language or omni models, and a reinforcement learning trainer as a backend. So the honest reading is trainer-optional rather than trainer-free, and the bold label oversells the second half. That is not a criticism so much as a warning about where the flexibility lives. If you write a linear script you get the transparency the principle promises, and you also take on the loop, the checkpointing, the logging and the failure handling. If you use their trainer you get all of that written, and you are back inside an abstraction you chose to avoid. The presence of a reinforcement learning trainer as a pluggable backend is the more interesting signal, because that is the case where a fixed loop is much harder to do without, and the roadmap says reinforcement learning post-training for omni models is planned next.

Three accelerator vendors, and extras that the installer is told to treat as conflicting

The accelerator story is better than most frameworks manage. The feature list names support for training on one vendor's graphics processors, on a competing vendor's compute stack, and on a domestic accelerator from a third, and the dependency metadata turns that into something an installer can act on rather than a sentence in a readme. The hardware extras are declared mutually exclusive, with a declaration in the installer configuration that says so explicitly, and there is a comment showing the shape of the problem: one optional kernel library needs the graphics extra, so you install both and the tooling knows they are not alternatives. The graphics extra itself is where the real constraint lives, because it pins the framework, the vision library, the audio library and a video codec library to exact builds of a specific version with a specific compute toolkit, then adds a list of vendor-specific packages including a collective communications library and a pair of kernel-compilation runtimes. Two comments in that list are worth reading as engineering notes rather than as noise. One explains that a particular frontend is there for training a sparse attention variant from one model family, which tells you the dependency is a feature rather than an accident. The other warns against a different extra because it pins a version of the same kernel library that conflicts with the one you want. So the picture for an adopter is: the vendor-neutral path is genuinely supported,

toml
gpu = [
  "torch==2.11.0+cu130",
  "torchvision==0.26.0+cu130",
  "torchaudio==2.11.0+cu130",
]

and the vendor-specific path is pinned hard, with the conflicts documented in the file itself. If your hardware is the domestic accelerator, expect a different extra, expect it not to be the same set of packages, and expect the community extras to have been written for the default path.

A Python window two versions wide, and two installers with different pinning rules

Two facts in the package metadata will decide whether you can install this at all, and both are the kind of thing that only shows up when you try. The first is the interpreter requirement, which allows two minor versions and excludes the next one. A framework that excludes a current Python release is making a statement about what it has been tested against, and for a project on a fast-moving framework version that is a defensible constraint, but it does mean this package is the reason you cannot upgrade the interpreter, rather than a consequence of it.

toml
requires-python = ">=3.11, <3.13"

The second is a comment, and it is the most valuable line in the whole file. It explains that the model library is pinned by a named dependency group for users of one installer, and then says, in terms, that users of the other installer should pin that library themselves to a specific version. So the single most important version constraint in the dependency graph is enforced automatically on one path and left to you on the other. That is not a criticism of the project, since the comment is honest about it, but it is a real difference in behaviour between two commands that look equivalent, and it means the reproducibility of your environment depends on which installer you used. A third detail fits the same pattern. One data library is pinned to a narrow range spanning a single minor version, which is a strong statement that a wider range breaks something. And there is a plotting library among the runtime dependencies rather than the development ones, with a comment saying it is used by a monitoring utility that draws an expert-load heatmap, which is a nice piece of honesty about why a training framework depends on drawing pictures.

A recipe zoo organised by modality, and one model whose fast path is opt-in

The supported model table is the most useful page in the readme, and the value is not the list of names but the shape of it. The models span plain text models from under a billion parameters to seventy billion, mixture-of-experts variants with an active-parameter count in the name, vision-language models, an omni model that handles several modalities at once, a video generation model with a named image-to-video checkpoint, and a second video model. So the framework's claim about any modality is concrete rather than aspirational, and the configuration files back it up: the paths are organised into a text directory, a multimodal directory and a directory named for diffusion transformers, which is the modality taxonomy expressed as a directory listing. Each row gives you a configuration file, so adopting a model is copying a YAML file and changing the parts you need, and there is a documented guide and checklist for adding one that is not listed. Two rows deserve more than a mention. One is a hundred-billion-parameter open model whose configuration file name encodes both low-rank adaptation and expert parallelism at degree four, which tells you parameter-efficient fine-tuning and distributed expert routing are expected to be combined rather than chosen between. The other is the newest architecture in the table, listed as checkpoint dependent, where two of its components default to the unfused path and the faster kernel backends are optional and tied to a particular generation of accelerator hardware, with a design document linked specifically about kernel selection. So for the newest model you get correctness by default and speed if you opt in with the right hardware, which is the right default and worth knowing before you benchmark anything.

patchgen: generated patches for model support, with a preview mode and a check

The build tooling contains something unusual and worth understanding, because it tells you how the project expects to grow. There is a command called patchgen with two invocations in the build file. One runs it across everything and asks for a diff, which is a preview: it shows you what the tool wants to change without changing it. The other runs a check, which is the same operation with the intent of verification, and it exists as its own target alongside a check target for the agent documentation. The tool also has its own package directory at the top level. A code generator that can emit the mechanical parts of adding a model to a framework, with a diff mode for review and a check mode for continuous integration, is a real answer to the problem every model-supporting framework has, which is that adding a model is mostly repetitive edits that nobody wants to review by hand and nobody wants to type. Two things are worth noting about how it is wired. The preview mode being a first-class flag means the tool is designed to be run before it is trusted, which is how you should treat any generator. And the check target means the repository asserts that its generated output is in sync with its source, so drift between them is a build failure rather than something a contributor notices a year later. Alongside this there are two shell entry points at the top of the repository, one for building and one for training, which is a small signal about how many entry points a real training run needs.

A continuous integration check that the agent instructions point at files that exist

The build file contains seven targets and one of them has no counterpart in most projects. The check-agent-docs target runs a Python script from a continuous integration directory whose job is to check the paths in the agent documentation. The repository carries six separate tool configurations for coding assistants and editors at the top level, including files for three different assistants, a configuration for a repository review bot, and an agent instruction file at the root. So the project has automated the problem of those files going stale, which is a real problem: a document that tells a tool where your code is will reference paths that stop existing as the repository moves, and a tool that follows a wrong path will confidently tell you something untrue about your own codebase. Validating those references in continuous integration is a small amount of Python and it removes an entire class of misleading output. The rest of the build file is conventional and shows the project's centre of gravity. The quality and style targets run a fast linter and its formatter in check mode over four directories, being the package, the tests, the tasks and the documentation, so the documentation is linted alongside the code. The test target runs one directory. The commit target installs the pre-commit hooks and then runs them across every file, which means the hooks are enforced for a contributor who forgets to run them locally. And the build target still calls a source distribution script directly rather than going through a standard build front end, which is a small piece of legacy in an otherwise modern setup.

The performance evidence is a figure and a paper, and the paper has another name

The performance section is one image and a sentence pointing at a paper, and that is a deliberate choice with consequences for an evaluator. There is no table in the readme, no throughput figure, no scaling curve in text, and no benchmark you can run. If you want the numbers you are reading a paper, which is the right place for them and also a different kind of evidence. The paper is on a preprint server under a different name from the project, with a title describing a model-centric distributed recipe zoo, and the news section records that it was accepted to a conference in late 2025. Peer review is worth something, and worth being precise about: it tells you the method was evaluated by researchers who read it, and it tells you nothing about whether the code runs on your hardware, whether the recipes are current, or whether the maintainers will be there next year. The documentation is built by a hosted documentation service, which is a good sign for a project that will change, and the community channel is a chat group, which is worth noting for anyone who needs support in a searchable form. The list of downstream work is short in what is visible here, with one fine-tuning project from another research group, and the readme truncates after it. The version record is the most reassuring line in the whole document, and it is the last thing to look at.

Apache-2.0, twelve patch releases in a year, and a roadmap that is four issues

The project is young and moving, and the dates say so. The news section records a first release in April 2025, a technical report and an open community group in August, a first official version in September, and a paper accepted in November. The release list then shows that official line moving from zero point one point oh to zero point one point twelve over the following year, with the most recent in September 2026, and the last push to the repository was on 2026-09-28. Twelve patch releases in twelve months on a research framework, with a paper behind the design, is a healthy signal and it is the signal most worth weighing. The roadmap is unusual in a good way: instead of a document it is four linked issues, covering the next minor version, a tool for measuring balance in a vision transformer, running a validation dataset during training rather than only at the end, and reinforcement learning post-training for omni models in partnership with a separate reinforcement learning framework. Roadmap-as-issues means each item has public discussion, which is a better place to find out whether something is planned than a checklist on a readme. The licence is Apache 2 with the file referenced from the package metadata, which is the grant most organisations accept for a training framework without discussion. The author metadata lists a single corporate entry with a team address rather than an individual, which tells you who to file issues against when a recipe does not work. And the small human details are worth noting as evidence that people write this file: one roadmap item has a typo in its name, and an environment file is committed at the top of the repository, which for a training framework is a file that more often than not should not be.

Editorial conclusion

VeOmni is worth evaluating if you train more than one kind of model, because a recipe zoo is the right abstraction when the interesting decisions are per model rather than per training run, and the configuration layout mirrors the modality split so finding the right recipe is a directory listing rather than a documentation hunt. Three accelerators supported in one framework is also rarer than it sounds and matters if your hardware is not the default. Two things to check before you commit an afternoon to it. The Python window, which is two versions wide and excludes a current release, so the framework is likely to be the thing that blocks an upgrade rather than the other way round. And how you install, because the dependency metadata says in a comment that the most important version pin is handled differently depending on whether you use one installer or another, and a user of the other one is told to pin it by hand. And read the trainer claim properly rather than as marketing: the framework supports writing a plain linear script, and it also ships a basic trainer and a reinforcement learning trainer, so the real choice is whether you want a trainer at all, not whether one exists.

Frequently asked questions

Is VeOmni really free of trainer classes?

The stated principle is that linear training scripts avoid rigid, structured trainer classes, and two specific frameworks are named as the thing being avoided. The same paragraph adds that a basic trainer is supported for text-only, vision-language and omni models, and a reinforcement learning trainer is supported as a backend. So the choice is whether you write the loop yourself, not whether a trainer exists.

Which Python versions and accelerators does it support?

The package metadata allows two minor Python versions and excludes the next one. For accelerators the feature list names one vendor's graphics processors, a competing vendor's compute stack, and a domestic accelerator, and the dependency extras for those are declared mutually exclusive to the installer, with the graphics extra pinning the framework and its libraries to exact builds of a specific version and compute toolkit.

What is a recipe in this framework?

A per-model configuration file. The supported models table gives one configuration file per model, and the files are organised into a text directory, a multimodal directory and one for diffusion transformers. Adopting a listed model means copying a configuration file and changing what you need, and there is a documented guide and checklist for models that are not listed.

Which parallelism and checkpointing does it provide?

The feature list names a fully sharded data parallel backend, sequence parallelism through a named Ulysses implementation in both synchronous and asynchronous modes, expert parallelism for large mixture-of-experts models, an efficient grouped matrix multiplication kernel for those models, and a distributed checkpoint facility from the framework itself rather than a saved single state file.

How current is the project?

The repository's first release was in April 2025 and its first official version in September 2025, and the release list shows twelve patch releases since, with the most recent on 2026-09-09. The licence is Apache 2, the documentation is built by a hosted service, and the roadmap is maintained as four public issues rather than as a document.

Official sources

  1. ByteDance-Seed/VeOmni on GitHub
  2. Issues
  3. License: Apache-2.0
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/bytedance-seed-veomni.svg)](https://hysenlabs.com/projects/bytedance-seed-veomni)