BMF: a graph runtime for heterogeneous video pipelines
Cross-platform, customizable multimedia/video processing framework. With strong GPU acceleration, heterogeneous design, multi-language support, easy to use, multi-framework compatible and high performance, the framework is ideal for transcoding, AI inference, algorithm integration, live video streaming, and more.
At a glance
- What is it?
- BMF is ByteDance's multimedia framework, and what makes it more than an FFmpeg wrapper is that the interesting work happens between the codecs: scheduling a graph across CPU and GPU, converting colour spaces and pixel formats, and treating FFmpeg, NumPy, PyTorch, OpenCV and TensorRT as interchangeable frame carriers. The demos are Colab notebooks, which tells you both how to try it and how little a local run will resemble them.
- Who is it for?
- Adopt BMF if you are building a pipeline that mixes CPU and GPU work and want one graph language across Python, Go and C++, and read the GPU frame-extraction and transcoding notebooks first because they show the scheduler rather than a toy graph. Do not expect a release to match the tree: the last tag is v0.2.0 from 2025-06-27 while the last push was 2026-08-26, so you are building from master.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 36 days ago.
- What is it written in?
- Mainly C++, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What ByteDance actually built: a graph, not an encoder
BMF, the Babit Multimedia Framework, is described as a cross-platform, multi-language, customizable multimedia processing framework developed by ByteDance, and the useful way to understand it is to look at what its demos assemble rather than at the feature list.
The Edit demo implements two Python modules, `video_concat` and `video_overlay`, and combines various atomic capabilities to construct what the README calls a complex BMF Graph. The GPU transcoding demo builds a pipeline that runs entirely on the GPU through seven stages: decode, then scale, flip, rotate, crop, blur, and encode. The broadcaster demo is a service with an API for pulling video sources dynamically, controlling layout, mixing audio and streaming the result to an RTMP server, and it is presented as showing the ability to adjust the pipeline dynamically.
That is the shape of the project: a scheduler over a graph of modules, where the modules can be codecs, filters, your own C++ or Python code, or a machine-learning model. An FFmpeg filter graph does the first thing in that list. BMF's claim is the rest of it.
The scale figures in the README are the company's own and should be read as claims: more than two billion videos processed by the framework every day, and a client-side sibling, BMFLite, used in apps such as Douyin and Xigua, serving more than one billion users and processing videos and pictures trillions of times a day. The framework has, by its own account, over four years of internal testing, which is why a project at version 0.2.0 has production history behind it.
Heterogeneous execution is the design, not a bullet point
Almost every framework can call a GPU. The harder problem is the seam between the CPU and the GPU, and that seam is what this framework is organised around.
The stated capabilities are three kinds of conversion: data format conversion across popular frameworks, namely FFmpeg, NumPy, PyTorch, OpenCV and TensorRT; conversion between hardware devices, meaning CPU and GPU; and colour space and pixel format conversion. Read together, those are the three places a mixed pipeline breaks. A frame decoded by FFmpeg into a NumPy array, uploaded to the GPU for a model, and returned as a TensorRT tensor for encoding is three format conversions and one device transfer, and doing that efficiently is the whole job.
The demos are built to make that visible rather than to hide it. The GPU frame-extraction demo explicitly points at heterogeneous processing between CPU and GPU and at hardware colour space conversion, and notes that new C++ and Python modules can be added simply. The GPU transcoding and filtering demo is more specific still: common video and image filters accelerated on the GPU, plus instructions on writing GPU modules, with all seven stages running on the device.
Performance is attributed to a scheduler plus hardware support, and the README credits NVIDIA with having cooperated on a GPU pipeline optimised for video transcoding and AI inference. That is the honest framing of what you are adopting: a runtime that has been tuned against real hardware with the vendor's help, not a portable abstraction.
One design note worth carrying into your own evaluation. The framework advertises Python, Go and C++ APIs, and the multi-language claim is demonstrated rather than asserted, since the frame-extraction demo mixes languages in one pipeline and the broadcaster demo is presented as evidence of multi-language development. If your graph needs a module in a language nobody else uses, that is the feature to test first, because it is the one a schema-driven framework is least likely to handle gracefully.
Nine packages, five console scripts and a Jinja module generator
The packaging tells you how the framework is actually put together. The Python distribution installs nine subpackages:
PACKAGES = [
'bmf',
'bmf.builder',
'bmf.cmd.python_wrapper',
'bmf.ffmpeg_engine',
'bmf.hmp',
'bmf.modules',
'bmf.python_sdk',
'bmf.server',
'bmf.templates',
]The names are informative. `bmf.ffmpeg_engine` is where the codec work lives, `bmf.hmp` is the heterogeneous pipeline layer, `bmf.modules` is the module registry, `bmf.server` is the service side that the broadcaster demo builds on, and `bmf.templates` holds the scaffolding.
And five console entry points:
CONSOLE_SCRIPTS = {
'run_bmf_graph': 'bmf.cmd.python_wrapper.wrapper:run_bmf_graph',
'trace_format_log': 'bmf.cmd.python_wrapper.wrapper:trace_format_log',
'module_manager': 'bmf.cmd.python_wrapper.wrapper:module_manager',
'bmf_env': 'bmf.cmd.python_wrapper.wrapper:bmf_env',
'bmf_template_generator': 'bmf.templates.cli:main',
}`run_bmf_graph` is how you execute a graph. `module_manager` is how modules are added, which is the mechanism behind the customizability claim. `bmf_template_generator` generates a module skeleton from Jinja templates, and the package data confirms it ships `jinja_templates/**/*.j2`. `trace_format_log` exists because the conversions described above are the part that goes wrong quietly, and you need a trace to see them.
So extensibility here is a code generator plus a module manager rather than a plugin interface you have to reverse engineer from a header. That is a real difference from a C++ framework where you would write a module, register it by symbol and hope the ABI holds. The Deoldify demo shows the contract at its most flattering: a state-of-the-art colourisation algorithm wrapped as a BMF Python module in fewer than one hundred lines.
A GPU build is a different package name
There is a packaging detail here that will confuse you if you hit it without warning, and it is two lines of the build script:
if "DEVICE" in os.environ and os.environ["DEVICE"] == "gpu":
PACKAGE_NAME = PACKAGE_NAME + "_gpu"GPU support is not a feature flag on one distribution. It is a separately named package, so a GPU environment and a CPU environment do not silently diverge by dependency and you cannot accidentally satisfy one with the other.
Around that line is a set of build-time overrides, all read from the environment: a package namespace, a package name override defaulting to `BabitMF`, a version override, a package URL override, and a switch controlling whether the namespace is nested inside the package list. The default URL override points at a repository path whose name differs from the one this code lives in.
Taken together, these say something specific about the project rather than something alarming. The build is parameterised so that one source tree can be published under different names, which is what a framework used both inside a company and outside one needs. It also means the published artefact is a product of environment variables, so a reproducible build is a documented environment rather than a documented command.
The Windows side has its own translation layer, from Python platform specifiers to CMake architecture arguments, and it covers four cases:
PLAT_TO_CMAKE = {
"win32": "Win32",
"win-amd64": "x64",
"win-arm32": "ARM",
"win-arm64": "ARM64",
}Four entries is the whole mapping. Anything not in that table has no documented path through this script, which matters if you are building for a Windows architecture that arrived after the table was written.
Eight build scripts, one per target, and a Go version in a text file
The top-level tree is mostly build scripts, and reading their names is the fastest inventory of what this framework targets. There is a plain `build.sh` plus `build_aarch64.sh`, `build_android.sh`, `build_d9.sh`, `build_osx.sh`, `build_wasm.sh` and `build_win_lite.sh`, alongside `CMakeLists.txt`, `CMakePresets.json`, a `cmake/` directory, a `3rd_party/` tree, `conda-env.yaml`, a `docker/` directory and a `win_env/` directory. The README does not say what the `d9` target is, so treat that filename as an open question rather than a documented platform.
Two of those scripts are more revealing than the rest. A WebAssembly build puts the runtime in a browser, and a `win_lite` build is the desktop counterpart to BMFLite, the client-side framework described in the README as lighter and more efficient for on-device work. Both point the same way: the graph runtime and the client runtime are separate products with separate build targets, sharing concepts rather than a binary.
Then the smaller files, each solving a specific reproducibility problem. `version.sh` exists because a version has to come from somewhere authoritative. `gosdk_version.txt` pins the Go SDK version as a file, which is the only way to keep a Go dependency reproducible without pinning the Go toolchain itself. `coverage.sh` and `test_pip.sh` are the quality gates, one for coverage and one for the published wheel. `create_symbols.py` is there to make a stack trace from a stripped release binary useful, which is a maintenance cost you only notice after the first crash report from a customer.
`.codebase/` at the root and a `CONTRIBUTING.md` alongside it suggest the project expects contributions, which combined with the demo story below tells you how it expects to receive them.
Colab notebooks are the documentation, and what that costs you
The documentation strategy is unusual enough to be a real decision. The README introduces its demo section by saying that for every demo provided, the corresponding implementation and documentation are available on Google Colab, so you can experience them intuitively.
For a C++ framework with GPU acceleration, Colab solves a genuine problem: nobody installs a CUDA toolchain to read a tutorial. It also means the demos run on Google's hardware with Google's driver and package versions, which is exactly what makes them approachable and exactly what makes them a poor predictor of your machine.
The six dimensions the demos cover map onto the claimed use cases: transcode, edit, meeting and broadcaster, GPU acceleration, AI inference and the client-side framework. Two of them carry unusual weight. The broadcaster demo is the only one presented as a service with an API, which is what tells you the framework is meant to sit inside a live system rather than a batch script. And the LLM preprocessing demo is described as a prototype of how ByteDance builds video preprocessing for training data, serving billions of clip processing a day: input video is split on scene change, subtitles are detected and cropped by an OCR module, video quality is scored by a provided aesthetic module, and the surviving clips are encoded as output.
That pipeline is a better description of the framework's real purpose than the transcoding demo, incidentally. Scene detection, OCR, quality scoring and clipping are exactly the operations where a graph scheduler and a hardware abstraction earn their keep, because each stage is somebody else's model or filter and none of them agree on a pixel format.
The practical rule: use the notebooks to learn the graph syntax, then rebuild the pipeline locally before you commit to it. A notebook that runs in three minutes on a hosted GPU is not evidence that your pipeline runs in three minutes on your hardware.
Apache-2.0, and fourteen months of commits past the last tag
The licence is Apache-2.0, recorded in LICENSE at the top of the repository, with CONTRIBUTING.md alongside it and no CLAIMS about warranty that the README makes.
The release history is the thing to plan around. The three most recent tags are v0.2.0 from 2025-06-27, v0.1.0 from 2025-04-14 and v0.0.13 from 2025-03-26. The last push was on 2026-08-26. So the tree has moved more than a year past the newest release, and the version string in the packaging script is still `0.2.0`, which means it tracks the last tag rather than the source.
There is a defensible reason for the gap. The framework has four years of internal use behind it and only a short public release history in front of it, so a low version number describes the open-source cadence rather than the maturity of the thing. That is also why the version starts at 0.x for something processing two billion videos a day inside one company: the public surface is still moving.
What it means for you is concrete. Installing from a package index gets you the 2025-06-27 code. Everything you read in the current tree, including any demo notebook that only exists on master, is newer. Building from source is the only way to get the current state, and the build scripts per target are the entry point for that.
The version-to-feature mapping is the other thing to check before you adopt. A framework whose generator, scheduler and module set all move together will occasionally have a module built against a newer header than the runtime you installed, and the presence of `bmf_env` and `trace_format_log` as first-class commands suggests the project expects you to interrogate the runtime rather than infer it.
Editorial conclusion
Adopt BMF if you are building a pipeline that mixes CPU and GPU work and want one graph language across Python, Go and C++, and read the GPU frame-extraction and transcoding notebooks first because they show the scheduler rather than a toy graph. Do not expect a release to match the tree: the last tag is v0.2.0 from 2025-06-27 while the last push was 2026-08-26, so you are building from master. Check which package you install, because a GPU build ships under a different package name, and check the Windows architecture mapping, since the build script translates only four platform specifiers to CMake targets.
Frequently asked questions
What is BMF and who developed it?
BMF, the Babit Multimedia Framework, is a cross-platform, multi-language, customizable multimedia processing framework developed by ByteDance. The project says it has four years of internal testing and is used in ByteDance's video streaming, live transcoding, cloud editing and mobile pre and post processing, with more than two billion videos processed by the framework each day.
Which languages can I write BMF modules in?
The framework provides Python, Go and C++ APIs. The GPU frame-extraction demo mixes languages within one pipeline and notes that new C++ and Python modules can be added simply, and the Deoldify demo wraps an existing colourisation algorithm as a BMF Python module in fewer than one hundred lines.
How does BMF handle CPU and GPU work in the same pipeline?
Heterogeneous processing between CPU and GPU is a first-class feature rather than an optimisation, and the README lists conversion between hardware devices alongside format conversion across FFmpeg, NumPy, PyTorch, OpenCV and TensorRT, plus colour space and pixel format conversion. The GPU transcoding demo runs decode, scale, flip, rotate, crop, blur and encode entirely on the device.
How do I add my own module to a BMF graph?
The framework ships a `module_manager` console script for adding modules and a `bmf_template_generator` that scaffolds one from Jinja templates, and a graph is executed with `run_bmf_graph`. The broadcaster demo goes further and shows the pipeline being adjusted dynamically at runtime.
What is the difference between BMF and BMFLite?
BMFLite is the client-side variant, a lightweight cross-platform multimedia processing framework described as being for apps rather than servers. The README says its algorithms are used in apps such as Douyin and Xigua across live streaming, video playback and cloud games, and there is a separate `build_win_lite.sh` build script for it.
How do I try BMF without installing a GPU toolchain?
Every demo the README lists is available as a Google Colab notebook, covering transcode, edit, broadcaster, GPU acceleration, AI inference and the client-side framework. The notebooks run on hosted hardware, so use them to learn the graph syntax and then rebuild the pipeline locally before relying on it.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/babitmf-bmf)