# GPT-NeoX: Distributed Training for Billion-Parameter Models

> EleutherAI's framework for training large language models across GPUs using Megatron and DeepSpeed. Built for researchers and labs that need to train models from scratch at scale, not for inference on existing models.

**EleutherAI/gpt-neox** — An implementation of model parallel autoregressive transformers on GPUs, based on the Megatron and DeepSpeed libraries

- Repository: https://github.com/EleutherAI/gpt-neox
- Website: https://www.eleuther.ai/
- Stars: 7,465 · Forks: 1,124
- Language: Python
- License: Apache-2.0
- Published: 2026-09-22 · Updated: 2026-09-22 · Language: en
- Canonical page: https://hysenlabs.com/projects/eleutherai-gpt-neox

## A training framework for billion-parameter models

GPT-NeoX is built for one task: training new language models with billions of parameters. The README states this directly: if you are not training models with billions of parameters from scratch, this is likely the wrong library. The framework is in widespread use across academic, industry, and government labs including Oak Ridge National Lab, Stability AI, CarperAI, Korea University, Carnegie Mellon University, and the University of Tokyo. EleutherAI created it to centralize techniques for large-scale autoregressive training and accelerate research into how those systems work. The project is not for inference: the README explicitly recommends Hugging Face transformers for teams that need to run existing models. GPT-NeoX targets research labs and organizations with access to GPU clusters who need to control training from the ground up. For teams doing inference only, fine-tuning existing models, or lacking dedicated cluster infrastructure, Hugging Face transformers is the right choice. It is not the right tool for inference, model serving, or small-scale training.

## How GPT-NeoX distributes training across thousands of GPUs

GPT-NeoX is built on NVIDIA's Megatron Language Model and augmented with techniques from DeepSpeed, plus novel optimizations from EleutherAI. The framework implements distributed training with ZeRO and 3D parallelism, allowing a single model to span thousands of GPUs across multiple machines. ZeRO reduces memory overhead; 3D parallelism divides computation across tensor, pipeline, and data dimensions. It includes advanced architectural innovations: rotary and alibi positional embeddings, parallel feedforward attention layers, and flash attention. The library includes predefined configurations for popular architectures including Pythia, PaLM, Falcon, and LLaMA 1 and 2, so you start with a known architecture rather than guessing at layer counts and hidden dimensions. Training also supports curriculum learning, which orders your dataset by difficulty during training. The framework integrates with Hugging Face tokenizers and transformers, WandB, Comet, and TensorBoard for experiment tracking and evaluation via the Language Model Evaluation Harness.

## Setting up training on your cluster

Clone the repository and install dependencies:

```bash
git clone https://github.com/EleutherAI/gpt-neox.git
cd gpt-neox
pip install -r requirements/requirements.txt
```

The framework supports multiple launch methods depending on your cluster infrastructure. For a Slurm-managed supercomputer, you write a Slurm script that calls deepy.py with your configuration path. For MPI-based clusters, GPT-NeoX handles the launch directly. For the IBM Job Step Manager, use that launcher. For local development on a single machine or small cluster, deepy.py is the entry point. First, choose or write a configuration file specifying your model architecture, parallelism settings, learning rate, batch size, and dataset. The configs/ directory includes predefined configurations for Pythia, PaLM, Falcon, and LLaMA 1 and 2, which you can customize by editing the YAML. You then run training via the launch method matching your cluster infrastructure. The framework distributes training automatically across your GPUs once you specify your parallelism strategy and data configuration. The last push to the repository was on 2026-09-04, indicating active maintenance and recent fixes.

## Checkpointing to disk and cloud storage

Training a billion-parameter model takes weeks or months, so checkpointing is essential. GPT-NeoX supports saving model state at intervals and resuming from those checkpoints, so a hardware failure, node outage, or preemption does not lose your entire progress. The framework also supports checkpointing directly to AWS S3 via the s3_path config option, which you activate in your YAML configuration. This allows you to scale across multiple machines or cloud environments without depending on local shared storage or network filesystem overhead. Cloud checkpointing is critical when training across cloud providers using spot instances that can be preempted at any time. Without cloud checkpointing, you would need to maintain a shared filesystem across all training nodes, which adds latency and complexity.

## Maintaining compatibility: versions 1.0 and 2.0

Prior to March 2023, GPT-NeoX relied on DeeperSpeed, which was based on an older DeepSpeed version (0.3.15). The project maintains two versioned releases for backward compatibility. Version 2.0 and later are built on the latest DeepSpeed and will be maintained going forward. Version 1.0 maintains snapshots of the old stable versions that GPT-NeoX-20B and the Pythia Suite were trained on. If you are starting a new project, use version 2.0 for access to the latest features and optimizations. If you need to reproduce results from existing published models like Pythia or GPT-NeoX-20B, use version 1.0 and the matching configurations to ensure compatibility.

## Emerging architectures: Mamba, Mixture-of-Experts, and RWKV

GPT-NeoX continues to add support for new model architectures beyond the transformer. On 2024-03-15, the framework added support for Mamba with tensor parallelism. On 2024-03-21, support for Mixture-of-Experts (MoE) was added, allowing sparse models where different token inputs route to different sets of expert layers. On 2024-05-21, RWKV support with pipeline parallelism arrived. This reflects the library's role as a centralized place for researchers to implement and test new training techniques at scale. Support for AMD MI250X GPUs was added on 2024-03-17, broadening hardware options beyond NVIDIA and allowing training on alternative GPU vendors. Recent updates also added Transformer Engine integration for further optimization and DPO, KTO, and reward model support for preference learning.

## Cost and vendor lock-in

GPT-NeoX is open source under the Apache 2.0 license, so there is no licensing cost. Training cost is determined by your hardware: GPU clusters at supercomputing centers, cloud providers, or private data centers. Because GPT-NeoX is software you run on your own infrastructure, you own the trained models and are not locked into a vendor's proprietary inference API. You can serve the trained model however you want: deploy it to Hugging Face, export it for inference in transformers, or integrate it into a product. The trade-off is operational overhead: managing a GPU cluster, configuring distributed training, monitoring long training runs, and troubleshooting issues with synchronization across nodes is more complex than paying for a managed model API. However, for organizations training multiple models or scaling to truly large parameters, the operational complexity becomes worthwhile compared to calling a vendor API and paying per token.

## Conclusion

GPT-NeoX is for researchers and labs training new billion-parameter models from scratch, not for teams doing inference on existing models. If you need inference, use Hugging Face transformers instead. Choose a predefined configuration (Pythia, Falcon, LLaMA), decide on your parallelism strategy (ZeRO, 3D), and verify that your cluster can launch training via Slurm, MPI, or local deepy.py before committing to a long training run.

## FAQ

### Is GPT-NeoX a large language model?

No. GPT-NeoX is a training framework, not a model itself. You use it to train your own models from scratch. EleutherAI has published trained models like Pythia and GPT-NeoX-20B using this framework, but the core library is the software for training.

### What is the difference between GPT-NeoX and GPT-3?

GPT-3 is a proprietary model trained by OpenAI. GPT-NeoX is EleutherAI's open source framework for training large language models. You can use GPT-NeoX to train models similar to GPT-3 if you have access to a GPU cluster with sufficient compute resources.

### What is GPT-NeoX?

GPT-NeoX is EleutherAI's framework for training large language models on GPU clusters. Built on NVIDIA's Megatron and DeepSpeed, it enables research labs to train billion-parameter models using distributed training with ZeRO and 3D parallelism.

## Sources

- [EleutherAI/gpt-neox on GitHub](https://github.com/EleutherAI/gpt-neox)
- [License: Apache-2.0](https://github.com/EleutherAI/gpt-neox/blob/main/LICENSE)
- [Project website](https://www.eleuther.ai/)
- [README](https://github.com/EleutherAI/gpt-neox/blob/main/README.md)
- [Releases](https://github.com/EleutherAI/gpt-neox/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/eleutherai-gpt-neox
