Library / SDK
flagos-ai/FlagGems avatar
flagos-ai/FlagGems

FlagGems: a Triton operator library that swaps in behind the PyTorch ATen backend

FlagGems is an operator library for large language models implemented in the Triton Language.

1,125 stars555 forksPythonApache-2.0

At a glance

What is it?
FlagGems implements LLM operators in Triton and registers them with PyTorch's ATen backend, so existing model code keeps its API while kernels run on non-NVIDIA accelerators. Here is what the repository documents, what it leaves open, and where it is the wrong tool.
Who is it for?
Adopt FlagGems if you already run PyTorch models on a non-NVIDIA accelerator and want Triton kernels behind the familiar ATen calls, and if you can budget time for a kernel-coverage audit of your own model. Do not adopt it as a general CUDA replacement on NVIDIA hardware, and do not expect it to fix a model that torch.compile already handles well.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 6 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 25, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem FlagGems targets: chip-specific operator stacks

Every accelerator vendor ships its own kernel library, and every library exposes a slightly different set of operators with slightly different numerics. A model that runs on one chip needs porting work before it runs on another. FlagGems is the FlagOS project's answer to that fragmentation: one operator library, written in Triton, intended to run across what the README calls "over 10 supported backends." The audience is narrow and specific. It is not the person training a model on an NVIDIA GPU who is happy with the vendor stack. It is the engineer who has a Llama-2-7b or Bert-base-uncased workload and a non-NVIDIA board, and who would rather not rewrite model code to get it running. The README names those two models plus Llava-1.5-7b as sample models for testing, which tells you the intended shape of the workload: transformer inference and training, not arbitrary tensor math.

How the ATen registration actually works

The mechanism is registration, not replacement. FlagGems registers its Triton kernels with the ATen backend of PyTorch. Model code keeps calling torch.nn.functional and torch operators as before; the dispatch underneath resolves to a FlagGems kernel instead of the default one. The README states this directly: developers "can continue using their familiar Pytorch APIs while at the same time benefit from new hardware acceleration technologies."

Two design choices follow from that. First, the library is eager-mode ready and independent of torch.compile, so it does not require a graph capture pass to take effect. Second, there is a per-function runtime kernel dispatcher, which the README lists as a feature; the cost of that lookup is paid on each call rather than folded into a compiled graph. The repository also carries a C++ Triton function dispatcher, listed as work in progress, which suggests the team considers the Python-side dispatch path worth shortening. The README does not quantify either the dispatch overhead or the gain from the C++ path.

The operator set is not a fixed hand-written list. The README describes automatic pointwise operator codegen supporting arbitrary input types and layout, so pointwise operations are generated rather than enumerated. Hand-optimized kernels are described as covering "selective operators," which is the honest phrasing: not every operator in the library has had tuning attention.

Installing FlagGems and running a first model check

The README does not contain install commands. It points to a Getting Started page and a usage page on flagos-ai.github.io, and the repository ships setup.py and pyproject.toml for a normal Python build. The package name on PyPI is flag_gems, and requires-python is >=3.10.0. The build-time constraints are worth knowing before you file a bug: pyproject.toml pins setuptools below 77 and setuptools-scm below 10, with comments explaining that setuptools 77 and later emit License-File metadata that some indexes reject, and that setuptools-scm 10.x breaks pip install . with an AttributeError on ignore_egg_info_in_manifest. If your environment forces newer versions of those two packages, that is the failure you will see.

The repository's own example files are the most concrete starting point, since the README gives none. A typical first step is to enable FlagGems and then run one of the bundled model scripts, for example the BERT test:

bash
pip install flag_gems
python examples/model_bert_test.py

The enable call is the switch that turns registration on; the README does not show its exact form beyond describing registration with the ATen backend, so check the usage documentation for the current call before wiring it into a training script. For a larger check, examples/model_llama_test.py and examples/model_llava_test.py cover the other two sample models. If you are serving rather than training, examples/integration_gems_with_vllm.py and examples/deepseek_with_vllm_test.py show the vLLM integration path, and examples/pretune.py exists for pre-tuning kernels ahead of a run. The repository also ships a benchmark/ directory and a flaggems-benchmark package, so you can measure the delta on your own hardware rather than trusting a number from the README, which does not publish one.

Where FlagGems is the wrong tool

The library is only as complete as its operator coverage, and coverage is the failure mode you will actually hit. If your model calls an operator that FlagGems has not registered, dispatch falls back to whatever the backend provides, which may be a slow reference implementation or nothing at all. Nothing in the README describes a coverage report, a fallback warning, or a way to enumerate which operators are registered on a given backend. You find out by running the model.

The hand-optimized claim carries a similar caveat. The README says selective operators are hand-optimized, so performance across the library is uneven by design, and the automatic codegen path for pointwise operators is a generality mechanism rather than a tuning one. On NVIDIA hardware, where the vendor stack is already tuned, the case for FlagGems is weak: the README's pitch is about portability across accelerators, not about beating CUDA on its home turf. And if your workload is already served well by torch.compile on a supported backend, adding a second dispatch layer is cost without a matching benefit. Finally, the C++ dispatcher is explicitly work in progress, so anyone planning around that path is planning around unfinished code.

FlagGems against vendored kernel libraries

The obvious alternative is the vendor kernel library that ships with your accelerator: cuDNN and cuBLAS on NVIDIA, the equivalent stacks elsewhere. Those are closed, tuned per chip, and updated by the vendor. FlagGems takes the opposite approach on every axis. It is Apache-2.0, written in Triton rather than CUDA, and maintained as one codebase intended to serve many backends instead of one codebase per chip. The difference in practice is who absorbs the porting cost. With a vendor library you wait for the vendor to support your chip. With FlagGems you inherit a library whose kernels must be correct and reasonably fast on all of them, which is a harder engineering problem and shows up as uneven operator maturity. The README's own framing is a "develop once, run anywhere" workflow, and that is the trade: breadth of backends in exchange for depth of tuning on any single one.

Maintenance, releases and the Apache-2.0 licence

The last push to master was on 2026-06-24, the same date as the v5.3.0 release, which is tagged "FlagOS 2.1, FlagGems v5.3.0." Before that, v5.0.2 landed on 2026-04-24 and v5.0.1.rc.0 on 2026-03-26. The repository is not archived. The release cadence visible in those three entries is roughly every one to two months, with the caveat that one of them is a release candidate, so not every tag is a stable drop.

Upgrade cost is dominated by the build pins rather than the API. Because pyproject.toml holds setuptools below 77 and setuptools-scm below 10 for the reasons described in the file itself, an environment that upgrades those packages independently can break the install even though nothing in FlagGems changed. The setup.py shim adds a second wrinkle: it notes that on master the declared backend is scikit-build-core, which ignores setup.py, and that the released wheel is built by .github/workflows/release.yaml after patching the backend to setuptools.build_meta. Installing from a git checkout and installing from a released wheel are therefore not the same build path.

The licence is Apache-2.0, stated in the README and in the LICENSE file. That is a permissive licence with an explicit patent grant and a notice requirement when you redistribute. The pyproject.toml license field uses the older text form, "Apache Software License," rather than an SPDX identifier, which some automated licence scanners will flag even though the LICENSE file is unambiguous. That is a packaging detail, not a legal question, and it is not legal advice.

Editorial conclusion

Adopt FlagGems if you already run PyTorch models on a non-NVIDIA accelerator and want Triton kernels behind the familiar ATen calls, and if you can budget time for a kernel-coverage audit of your own model. Do not adopt it as a general CUDA replacement on NVIDIA hardware, and do not expect it to fix a model that torch.compile already handles well. Before committing, install the wheel in a throwaway environment, run flag_gems.enable(), and check that every operator your model actually calls is registered; the README does not document a rollback path or a coverage report, so that check is on you.

Frequently asked questions

What is FlagGems?

FlagGems is an operator library for large language models implemented in the Triton language, part of the FlagOS system software stack. It registers its kernels with the ATen backend of PyTorch so model code can keep using standard PyTorch APIs.

How does FlagGems relate to vLLM?

The repository includes examples/integration_gems_with_vllm.py and examples/deepseek_with_vllm_test.py, which show FlagGems being used alongside vLLM. The README itself does not describe the vLLM integration in prose, so the example files are the reference.

Which models can I use to test FlagGems?

The README lists Bert-base-uncased, Llama-2-7b and Llava-1.5-7b as sample models for testing, and the examples directory contains a test script for each of them.

What Python version does FlagGems require?

The pyproject.toml sets requires-python to >=3.10.0.

Official sources

  1. Official README
  2. Project repository
  3. Release notes
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/flagos-ai-flaggems.svg)](https://hysenlabs.com/projects/flagos-ai-flaggems)