# Axolotl: A Python Framework for Fine-Tuning Large Language Models

> Axolotl is an open-source Python framework for fine-tuning large language models, supporting LoRA, QLoRA, full fine-tuning, and a growing list of parallelism strategies including tensor, context, and expert parallelism. It is YAML-configuration-driven, uv-first as of April 2026, and runs on Docker with NVIDIA GPU support.

**axolotl-ai-cloud/axolotl** — Go ahead and axolotl questions

- Repository: https://github.com/axolotl-ai-cloud/axolotl
- Website: https://docs.axolotl.ai
- Stars: 12,512 · Forks: 1,449
- Language: Python
- License: Apache-2.0
- Published: 2026-09-09 · Updated: 2026-09-09 · Language: en
- Canonical page: https://hysenlabs.com/projects/axolotl-ai-cloud-axolotl

## What Axolotl Is and Who It Is For

Fine-tuning a large language model requires assembling a training stack: a compatible version of transformers, an accelerator library, a PEFT adapter library, a data loader, and configuration for distributed training if the model does not fit on one GPU. Axolotl provides a pre-assembled stack with YAML configuration, so practitioners can describe a training run in a config file and start training without writing custom training loop code.

Axolotl targets ML engineers and researchers who need to adapt pre-trained language models to specific tasks or datasets. The framework handles supervised fine-tuning, instruction tuning, direct preference optimisation, and reinforcement learning from human feedback variants including GRPO and DPO. Its pyproject.toml lists transformers, peft, trl, accelerate, and datasets as core dependencies, pinned to specific versions to maintain compatibility across the supported model list.

The project is actively maintained, with the last push on 2026-09-28. Version v0.19.0 shipped 2026-09-10, v0.18.0 on 2026-07-17, and v0.17.0 on 2026-06-03.

## Model Support and Training Techniques

Axolotl maintains an explicit list of supported models in its documentation. Recent additions include Ling 3.0, Muse Glimmer, Mistral Small 4, Qwen3.5, Qwen3.5 MoE, Gemma 4, GLM-4.7-Flash, GLM-4.6V, GLM-4.5-Air, and Shieldstral. Each supported model has a corresponding documentation page and, in most cases, example YAML configurations in the examples/ directory.

Training technique coverage is broad. LoRA fine-tuning is the baseline. QLoRA extends this with quantized base weights. ScatterMoE LoRA applies LoRA directly to MoE expert weights using Triton kernels. SonicMoE adds fused LoRA support for another MoE execution path. NVFP4 (4-bit) MoE LoRA training is supported via ScatterMoE (W4A16) and SonicMoE (W4A4), including merge back into an NVFP4 checkpoint. BitNet 1.58-bit fine-tuning was added in June 2026. Quantization-aware training with NVFP4 is supported via the quantize_moe_experts: true configuration key.

For reinforcement learning, Axolotl implements GRPO with asynchronous execution support (noted as providing up to 58% faster steps), standard DPO, and GDPO (Generalized DPO). EAFT (Entropy-Aware Focal Training) and EBFT are also integrated.

## Installing Axolotl and Running a Training Job

Axolotl is a uv-first project as of April 2026. Docker is the recommended environment for GPU training. The docker-compose.yaml in the repository mounts the local workspace and a HuggingFace cache directory into the container:

```yaml
services:
  axolotl:
    build:
      context: .
      dockerfile: ./docker/Dockerfile-uv
    volumes:
      - .:/workspace/axolotl
      - ~/.cache/huggingface/:/root/.cache/huggingface/
```

The service requires NVIDIA GPU hardware, declared in the docker-compose.yaml deploy.resources.reservations section. The Python package itself requires Python >= 3.12 and torch >= 2.13.0 as declared in pyproject.toml. Install from source using uv or pip.

A Colab notebook is available at the link in the README for users who want to try Axolotl without a local GPU. The examples/ directory contains YAML configuration files for dozens of models and training setups. Each example directory is named for the model family it covers: examples/gemma4/, examples/qwen2_5-vl/, examples/devstral/, examples/axolotl-contribs-lgpl/, and many others.

## Distributed and Parallel Training

Axolotl supports several parallelism strategies for training models that exceed the memory of a single GPU or a single node. Context Parallelism (CP), Tensor Parallelism (TP), and Fully Sharded Data Parallelism (FSDP) can be composed within a single node and across nodes through the ND (N-Dimensional) Parallelism feature, added in July 2025. The documentation for ND Parallelism is at docs.axolotl.ai/docs/nd_parallelism.html.

Expert Parallelism (EP) for distributed MoE training via DeepEP was added in June 2026. Context Parallelism for hybrid SSM models (Nemotron-H, Falcon-H1, Bamba) was also added at that time.

DeepSpeed support is present for single-GPU to multi-GPU training with DDP, and TiledMLP is available for expanding training beyond a single GPU with DDP and FSDP support. The deepspeed_configs/ directory in the repository contains pre-built DeepSpeed configuration files.

## Limitations and Platform Constraints

Axolotl is a training framework, not an inference server. It does not provide an API server, a chat interface, or deployment tooling. Teams that need to deploy fine-tuned models for inference will need separate tools such as vLLM, TGI, or Ollama.

The pyproject.toml excludes bitsandbytes and Triton on macOS (both are conditioned on sys_platform != 'darwin'). This means QLoRA and techniques that depend on custom Triton kernels (ScatterMoE LoRA, for example) are not available on macOS. Apple Silicon Mac users can install the base package but cannot use the full training feature set.

Version pinning is strict. The pyproject.toml pins torch to >=2.13.0,<=2.14.0, transformers to ==5.17.0, peft to ==0.21.0, and trl to ==1.13.0. This provides reproducibility but means Axolotl may not work with torch versions outside the pinned range until the maintainers update the pins.

Unsloth is an alternative LLM fine-tuning tool that focuses on memory efficiency and faster training through custom CUDA kernels. It targets a similar audience of practitioners fine-tuning models on constrained GPU hardware. The two projects are not interchangeable: Axolotl has broader model coverage and parallelism options; Unsloth focuses on single-GPU efficiency.

## Maintenance, License, and Dependencies

The last push to the repository was on 2026-09-28. Version v0.19.0 shipped 2026-09-10, following v0.18.0 on 2026-07-17 and v0.17.0 on 2026-06-03. The release cadence of roughly six to eight weeks between versions reflects active development.

Axolotl is licensed under the Apache-2.0 license. Two additional contribution packages are listed in pyproject.toml: axolotl-contribs-lgpl and axolotl-contribs-mit, indicating that some community contributions are packaged separately under different licences. Teams with strict commercial licence requirements should review the LGPL package's terms separately, as LGPL imposes additional conditions compared to MIT and Apache-2.0 when linking against it.

## Conclusion

Axolotl suits ML engineers who need a configurable, actively maintained fine-tuning framework that tracks the latest models and training techniques without requiring deep framework knowledge. It is not the right choice for production inference or deployment; it is a training tool. Teams without NVIDIA GPU hardware face a harder path since bitsandbytes and Triton are excluded from macOS builds. Before adopting it, verify that your Python version is at least 3.12 and that your torch version falls within the >=2.13.0,<=2.14.0 constraint declared in pyproject.toml.

## FAQ

### What Python version does Axolotl require?

Axolotl requires Python 3.12 or later, as declared in the requires-python field of pyproject.toml. It also requires torch >= 2.13.0 and <= 2.14.0 in the current release.

### Does Axolotl support training on macOS or Apple Silicon?

Axolotl can be installed on macOS, but bitsandbytes and Triton are excluded from macOS builds via platform conditions in pyproject.toml. This means QLoRA and Triton-kernel-based techniques like ScatterMoE LoRA are not available on macOS. Full fine-tuning without those techniques may work, but the full feature set requires a Linux system with an NVIDIA GPU.

### How does Axolotl handle multi-GPU and multi-node training?

Axolotl supports DeepSpeed, FSDP, Tensor Parallelism, Context Parallelism, and Expert Parallelism. These can be composed through the ND Parallelism feature added in July 2025. Configuration is managed through YAML files and the docker-compose.yaml setup. The deepspeed_configs/ directory contains pre-built DeepSpeed configurations.

## Sources

- [axolotl-ai-cloud/axolotl on GitHub](https://github.com/axolotl-ai-cloud/axolotl)
- [License: Apache-2.0](https://github.com/axolotl-ai-cloud/axolotl/blob/main/LICENSE)
- [Project website](https://docs.axolotl.ai)
- [README](https://github.com/axolotl-ai-cloud/axolotl/blob/main/README.md)
- [Releases](https://github.com/axolotl-ai-cloud/axolotl/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/axolotl-ai-cloud-axolotl
