# Liger-Kernel: Triton kernels that cut LLM training memory by 60%

> Liger-Kernel swaps Hugging Face model layers for fused Triton kernels, claiming 20% more training throughput and 60% less memory. Here is what it actually replaces, how to install it, and when it is the wrong tool.

**linkedin/Liger-Kernel** — Efficient Triton Kernels for LLM Training

- Repository: https://github.com/linkedin/Liger-Kernel
- Website: https://linkedin.github.io/Liger-Kernel/
- Stars: 6,635 · Forks: 613
- Language: Python
- License: BSD-2-Clause
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/linkedin-liger-kernel

## The memory wall Liger-Kernel is built to move

Fine-tuning a 7B or 8B model on a single GPU usually fails for one reason: activation memory, not weights. The README states that Hugging Face models start to OOM at a 4K context length while Hugging Face plus Liger Kernel scales to 16K, under the benchmark conditions it lists (LLaMA 3-8B, batch size 8, bf16, AdamW, gradient checkpointing, FSDP1 on 8 A100s). The target user is therefore narrow and specific: someone training or aligning an LLM with the Hugging Face stack who is hitting a context-length or batch-size ceiling, and who would rather change one import than rewrite the model.

The pitch is not a new trainer. Liger-Kernel is a collection of Triton kernels that replace individual operations inside a model you already have, plus a separate set of chunked loss kernels for post-training. The README frames the benefit as longer context lengths, larger batch sizes and massive vocabularies, which is a memory argument before it is a speed argument. If your job already fits in memory and runs at a speed you accept, the project has little to offer you.

## What the kernels actually replace in a Hugging Face model

The README lists the implemented operations as RMSNorm, RoPE, SwiGLU, CrossEntropy and FusedLinearCrossEntropy, described as Hugging Face compatible. These are the per-layer operations that run on every forward and backward pass, and they are where a naive PyTorch implementation materializes intermediate tensors that a fused kernel can keep in registers or shared memory. FusedLinearCrossEntropy is the interesting one: it merges the final linear projection with the cross-entropy computation so the full logits tensor for a large vocabulary does not have to exist in memory at once.

The repository layout confirms this is the whole product. The `src/` tree holds the kernel implementations and the patching layer, `benchmark/` holds the measurement harness, and `test/` is split by implementation backend, with separate directories for `cutedsl`, `cutile` and `cute`. The Makefile shows those backends are opt-in test targets rather than the default path, and it notes that the `cute` fused-MoE tests need a separate `liger_cute_kernels` wheel and at least two CUDA devices. That is a useful signal about where the project's maturity is concentrated: the Triton path is the default, and the CuTe DSL and cuTile paths are newer and gated behind optional extras.

## Installing Liger-Kernel and patching a model in one line

The package is published on PyPI as `liger-kernel`, with a nightly channel under `liger-kernel-nightly`. The README's installation section is the entry point, and the setup.py in the repository shows that dependencies are chosen by detected hardware: CUDA and CPU get `torch>=2.1.2` and `triton>=2.3.1`, ROCm gets `triton>=3.0.0`, XPU gets `torch>=2.6.0`, and NPU gets a pinned `torch==2.9.0` with `torch_npu==2.9.0` and `triton-ascend==3.2.2`. That pinning is worth reading before you install, because it means the project's supported matrix is not uniform across accelerators.

The README's central claim is that one line of code activates the kernels, and the post-training kernels are shown used as plain Python modules:

```python
from liger_kernel.chunked_loss import LigerFusedLinearORPOLoss
orpo_loss = LigerFusedLinearORPOLoss()
y = orpo_loss(lm_head.weight, x, target)
```

That snippet is copied from the README. The README states these chunked loss kernels reduce memory usage by up to 80% and that DPO, CPO, ORPO, SimPO, KTO and JSD are supported. If you are doing alignment work rather than supervised fine-tuning, this is the part of the library you would actually use, and it does not require patching the model at all.

The examples directory is the place to look for a full training loop: `examples/huggingface/` is described in the README as training LLaMA 3-8B about 20% faster with over 40% memory reduction on the Alpaca dataset using 4 A100s with FSDP, and `examples/lightning/` is listed with a 15% throughput figure. Run one of those scripts rather than assembling your own first, because the patch has to be applied at the right point in the process and the examples show where.

## The compatibility surface is the real constraint

The one-line patch is convenient precisely because it is invasive. It replaces operations inside a model class you did not write, which means correctness depends on the kernel matching the reference implementation's numerics closely enough for your training run. The repository has a `test/` tree and a `test/convergence` directory that the Makefile deliberately excludes from the default `test` target, which tells you convergence testing is treated as a separate, heavier activity rather than something that runs on every change.

There is a second constraint that the README is quiet about: version coupling. The setup.py pins floors and, for NPU, exact versions of `torch`, `torch_npu` and `triton-ascend`. The optional CuTe DSL extra pins `nvidia-cutlass-dsl>=4.6.0` and `apache-tvm-ffi>=0.1.0`, and a comment in setup.py explains the floor exists because older `cuda-tile` versions lack a `CompilerOptions` field and raise `unexpected keyword argument 'num_worker_warps'` at import. A library that patches another library's internals inherits a dependency on that library's internals staying stable. If your organisation upgrades `transformers` on its own schedule, you are now coordinating two upgrade calendars.

The third constraint is hardware. The published benchmark conditions are 8 A100s with FSDP1. The README does not document behaviour under tensor parallelism, pipeline parallelism, or multi-node configurations, and it does not document a rollback path if a patched run diverges. Treat the single-node case as the supported case until you have evidence otherwise.

## Liger-Kernel against Flash Attention and Unsloth

These three get compared often, and they operate at different layers. Flash Attention is an attention kernel: it makes the attention computation itself memory-efficient, and the Liger README states the kernels work out of the box with it. They are not substitutes, and the README treats Flash Attention as something you run alongside Liger rather than instead of it. If your bottleneck is attention at long context, Flash Attention is the relevant tool; if your bottleneck is the norm, rotary, activation and loss operations around attention, Liger-Kernel is.

The Unsloth comparison is a different axis. Unsloth is a fine-tuning stack that includes its own kernels and its own model loading path. Liger-Kernel's design choice is the opposite: it stays inside the Hugging Face model classes and patches them, so your trainer, your data pipeline and your checkpoint format do not change. That is the trade-off to weigh. A self-contained stack can optimise more aggressively across the whole pipeline because it controls the whole pipeline; a patch layer is easier to adopt and easier to remove, but it is bounded by what the underlying model class exposes. The README's claim that the kernels work out of the box with Flash Attention, PyTorch FSDP and Microsoft DeepSpeed is the payoff for that choice, and it is the reason the library is plausible as a drop-in rather than as a migration.

## Maintenance, licensing and what an upgrade costs

The repository is not archived, and the last push was on 2026-09-10. The most recent release is v0.8.2 from 2026-08-18, following v0.8.1 in July 2026 and v0.8.0 in April 2026. That is a steady cadence with a gap of roughly three months between the April and July releases, so plan for a few upgrades a year rather than continuous churn.

The upgrade cost is not the install. It is the re-validation. Because the library patches model internals, a new Liger release combined with a new `transformers` release is two variables changing at once. The repository gives you a way to separate them: `make test` runs the correctness suite while ignoring the convergence, cutedsl, cutile and cute directories, and `make test-convergence` is a separate target. Running the default suite after an upgrade tells you the kernels still match their reference implementations on the shipped tests. It does not tell you your fine-tune still converges, and the README does not claim it does.

On licensing: the project is BSD-2-Clause, and `pyproject.toml` points the license field at the `LICENSE` file. A permissive two-clause BSD licence is straightforward for commercial use, but the dependency chain is not covered by it. `torch`, `triton`, `transformers` and the optional NVIDIA packages carry their own terms, and the NPU path pins a specific `torch_npu` build. Check those separately; nothing in this repository's licence resolves them.

## Conclusion

Adopt Liger-Kernel if you fine-tune or align Hugging Face models on a single node with enough VRAM to be the bottleneck, and you want the memory win without rewriting your training loop. Do not adopt it if you need multi-node determinism guarantees, if you cannot pin torch and triton versions in your image, or if your model's architecture is not in the supported list. Before you commit, run one short training job with and without the monkey patch on your own hardware and compare peak memory and step time, because the published numbers come from LLaMA 3-8B on 8 A100s with FSDP1 and will not transfer directly to a different model or a different parallelism strategy.

## FAQ

### What is Liger-Kernel?

It is a collection of Triton kernels designed specifically for LLM training, implemented as Hugging Face compatible replacements for operations such as RMSNorm, RoPE, SwiGLU and CrossEntropy. The README states it can increase multi-GPU training throughput by 20% and reduce memory usage by 60%.

### What does Liger-Kernel do?

It patches a Hugging Face model so that fused Triton kernels run in place of the original per-layer operations, and it provides separate chunked loss kernels for post-training methods like DPO, ORPO and SimPO. The README describes the effect as enabling longer context lengths, larger batch sizes and massive vocabularies.

### How does Liger-Kernel compare to Flash Attention?

They are not alternatives. Flash Attention is an attention kernel, and the README states the Liger kernels work out of the box with it, so the two are used together rather than chosen between.

### What are kernels used for?

In this project, the kernels replace the per-layer operations that run on every forward and backward pass, such as RMSNorm, RoPE, SwiGLU and CrossEntropy. The README states the effect is higher training throughput and lower memory usage, which allows longer context lengths and larger batch sizes.

### What is a Triton kernel?

Triton is the language and compiler the Liger-Kernel operations are written in, and the README describes the project as a collection of Triton kernels designed specifically for LLM training. The setup.py lists triton as a dependency for the CUDA, CPU and ROCm paths.

### What is a kernel in an LLM?

In Liger-Kernel's usage, a kernel is a fused implementation of one model operation, such as RMSNorm, RoPE, SwiGLU or CrossEntropy, that replaces the Hugging Face version when the model is patched. The README lists these as the implemented Hugging Face compatible operations.

## Sources

- [License: BSD-2-Clause](https://github.com/linkedin/Liger-Kernel/blob/main/LICENSE)
- [linkedin/Liger-Kernel on GitHub](https://github.com/linkedin/Liger-Kernel)
- [Project website](https://linkedin.github.io/Liger-Kernel/)
- [README](https://github.com/linkedin/Liger-Kernel/blob/main/README.md)
- [Releases](https://github.com/linkedin/Liger-Kernel/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/linkedin-liger-kernel
