LitGPT: from-scratch LLM implementations for pretraining, finetuning and serving
20+ high-performance LLMs with recipes to pretrain, finetune and deploy at scale.
At a glance
- What is it?
- LitGPT ships 20+ language models as single-file implementations with no abstraction layers, plus YAML recipes for training. It suits engineers who want to read, modify and debug the model code, not just call it.
- Who is it for?
- Adopt LitGPT if you need to read and change the model code itself, or if you want LoRA, QLoRA and adapter finetuning behind a CLI and YAML configs on one to many GPUs. Do not adopt it if you only need the fastest possible OpenAI-compatible endpoint, or if you need a model that is not in its list.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 16 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What LitGPT solves for people who need to see the model code
Most LLM tooling hides the model behind a framework. You pass a config, it builds layers you never inspect, and when something goes wrong in the forward pass you are debugging someone else's abstraction. LitGPT takes the opposite position. The README describes every model as implemented "from scratch" with "no abstractions" and "single file implementations", and the pitch is explicitly about debugging: "Easy debugging with no abstraction layers".
That is the actual product. You get Llama 3, Gemma 2, Phi 4, Qwen2.5, Code Llama, Falcon and others as readable Python, plus the surrounding machinery to download weights, finetune, pretrain and serve them. The intended user is an engineer who wants to modify attention, swap a normalization layer, or trace exactly where memory goes during training. Someone who just wants an endpoint does not need this, and the README does not pretend otherwise.
The second audience is teams that need to finetune on their own data without writing a training loop. The repository ships a config_hub/ directory of YAML recipes and a CLI that consumes them, so a finetune is closer to editing a config file than writing code. Both audiences pay the same cost: you are responsible for understanding the pieces LitGPT hands you.
How LitGPT is put together: single-file models, YAML recipes, a CLI
The repository layout tells most of the story. The litgpt/ package holds the implementation, config_hub/ holds recipe YAML files, extensions/ holds optional add-ons, and tutorials/ holds the walkthroughs the README links to. The README points at tutorials/python-api.md for the full Python API.
Dependencies in pyproject.toml are split into groups. The base install pulls huggingface-hub, lightning, safetensors, tokenizers, torch and jsonargparse. The extra group adds bitsandbytes for quantization, datasets for litgpt.evaluate, litdata, litlogger and litserve for litgpt.deploy. The compiler group adds lightning-thunder and caps torch below 2.13 on Linux, with a comment explaining that thunder references named-tensor APIs removed in torch 2.13. That cap is a concrete example of the trade-off in this design: because LitGPT tracks torch closely and compiles models directly, it inherits breakage from upstream changes faster than a wrapper that pins an older runtime.
Model weights come from Hugging Face. The README's quick start loads "microsoft/phi-2" through huggingface-hub, and the repository includes a convert_hf_checkpoint path, which implies you can bring a Hugging Face checkpoint into the LitGPT format. The README does not document the conversion in the excerpt available here, so treat that as something to check in the tutorials before you plan around it.
Installing LitGPT and running a first generation
The README gives a single install command for the common case. The extra group is what brings in quantization, evaluation and deployment dependencies, so this is the install you want unless you are deliberately keeping the environment small.
pip install 'litgpt[extra]'If you would rather work from the source tree, the README shows a clone followed by either uv or pip. Note the three extras in the pip form: extra, compiler and test.
git clone https://github.com/Lightning-AI/litgpt
cd litgpt
# if using uv
uv sync --all-extras
# if using pip
pip install -e ".[extra,compiler,test]"The first real use is the Python API. The README's example loads phi-2 and generates a correction for a misspelled sentence. Expect the first run to download weights from Hugging Face, so budget time and disk for that before the model responds.
from litgpt import LLM
llm = LLM.load("microsoft/phi-2")
text = llm.generate("Fix the spelling: Every fall, the family goes to the mountains.")
print(text)
# Corrected Sentence: Every fall, the family goes to the mountains.The printed output in the README is the corrected sentence. Your exact string may differ because generation is sampled, but the shape of the call is the same for every model in the table: swap the Hugging Face identifier and the rest of the code holds. Python 3.10 or newer is required according to pyproject.toml, and the classifiers list support through 3.14.
Where LitGPT gets in the way
The no-abstractions stance is a real cost, not a slogan. Because each model is a separate from-scratch implementation, coverage is per-model and per-size. The README lists Llama 3 through 3.3 at 1B, 3B, 8B, 70B and 405B, Qwen2.5 from 0.5B to 72B, Gemma 2 at 2B, 9B and 27B, Phi 4 at 14B, CodeGemma at 7B, and so on. A model family that is not on that list is not a configuration change away. If your organization has standardized on a checkpoint outside the table, LitGPT is the wrong tool and no amount of YAML will fix it.
The second constraint is the dependency surface. The compiler extra pins torch below 2.13 on Linux because of the thunder interaction described in pyproject.toml. That means an environment already standardized on a newer torch cannot simply add the compiler extra. The base install has a looser floor (torch>=2.7), so inference and finetuning without compilation are less exposed, but anyone reaching for the compiler path should read that pin before upgrading torch across a fleet.
Third, the README is a launch page more than a manual. It links to tutorials for the Python API and lists workflows, but the README does not document rollback, checkpoint resumption semantics, or failure recovery during distributed training. If your finetune runs for days on a cluster, verify those behaviors in the tutorials before you commit, because the README will not tell you.
LitGPT versus vLLM and versus torchtune
The comparison people search for is LitGPT against vLLM, and the difference is what each one optimizes. vLLM is built around serving throughput: paged attention, continuous batching, an OpenAI-compatible server as the primary interface. You point it at weights and it serves them. LitGPT's serving path exists (the extra group includes litserve for litgpt.deploy), but the project's center of gravity is the model implementation and the training recipes. If your goal is maximum requests per second behind an HTTP API, vLLM is the better fit and LitGPT's from-scratch files are overhead you will never read.
The more interesting comparison is torchtune, which shares the philosophy of readable, native implementations but is organized around PyTorch-native recipes rather than a model zoo with a CLI. LitGPT's distinguishing feature is breadth in one place: 20+ models, a config_hub of YAML recipes, and the same CLI surface for download, finetune, pretrain and serve. The README also advertises LoRA, QLoRA and Adapter finetuning plus FSDP across 1 to 1000+ GPUs or TPUs, which is the scaling story torchtune-style toolkits leave to you. The cost of that breadth is the per-model coverage limit described above: torchtune's narrower scope means fewer models but less surface to keep aligned with upstream torch.
Maintenance, licence and the upgrade bill
The repository is not archived, and the last push was on 2026-09-09, which is recent enough that the project is being worked on now. Releases are less frequent than commits: v0.5.13 landed on 2026-06-29, preceded by v0.5.12 on 2025-12-18 and v0.5.11 on 2025-09-09. So the pattern is roughly two releases a year with ongoing commits between them. Plan upgrades around releases, not around the commit stream, and expect the changelog rather than the README to carry migration notes.
The licence is Apache-2.0, stated in the README badge and in pyproject.toml via license = { file = "LICENSE.md" }. The README frames this as "Apache 2.0 for unlimited enterprise use". That is the project's own framing, not legal advice; if your organization has a policy on model weights, remember that the licence covers LitGPT's code, while each checkpoint carries its own terms from its publisher. The README does not describe those per-model terms, so check them separately.
The upgrade cost is dominated by the torch relationship. Because LitGPT compiles and runs models directly against torch, a torch bump can break the compiler extra, as the 2.13 cap shows. Budget for testing the compiler path separately from the inference path on every upgrade.
Editorial conclusion
Adopt LitGPT if you need to read and change the model code itself, or if you want LoRA, QLoRA and adapter finetuning behind a CLI and YAML configs on one to many GPUs. Do not adopt it if you only need the fastest possible OpenAI-compatible endpoint, or if you need a model that is not in its list. Before committing, check that your target checkpoint appears in the model table, confirm which optional extra group supplies the dependencies you need, and run one finetune on your own data to see how the YAML recipe maps onto your hardware.
Frequently asked questions
What is the difference between an LLM and GPT?
LitGPT's README uses LLM as the umbrella term for the models it implements, such as Llama 3, Gemma 2, Phi 4 and Qwen2.5. The GPT family is not among the models listed there, and the README does not draw the distinction between the two terms.
Is ChatGPT a language model?
LitGPT's README does not describe ChatGPT. It does use the phrase large language models for the class of models it implements, and the package name litgpt reflects that general naming rather than any connection to ChatGPT.
How do I install LitGPT?
The README gives pip install 'litgpt[extra]' for the common case, or a source install with git clone followed by uv sync --all-extras or pip install -e ".[extra,compiler,test]". Python 3.10 or newer is required according to pyproject.toml.
Which models does LitGPT support?
The README lists Llama 3, 3.1, 3.2 and 3.3, Code Llama, CodeGemma, Gemma, Gemma 2, Phi 4, Qwen2.5, Qwen2.5 Coder, Falcon, Falcon 3, R1 Distill Llama, FreeWilly2 and Function Calling Llama 2. Coverage is per model and per size, so a family outside the table is not a config change away.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/lightning-ai-litgpt)