meta-llama/llama is a deprecated Llama 2 example with a 24 hour download URL and four unpinned dependencies
GitHub describes it as Inference code for Llama models. The repository metadata lists Python as its primary language. The metadata lists the NOASSERTION license. This article stays within the project description and details documented in the GitHub repository README.
At a glance
- What is it?
- A read of the Llama inference repository: the deprecation notice that points at five successor repositories, the model parallel values and pre-allocated cache the examples bake in, the signed link that expires after a day, and what a setup.py version of 0.0.1 with no version pins means for anyone still trying to install it.
- Who is it for?
- This repository still answers one narrow question well, how a Llama 2 checkpoint is loaded and prompted, and its MODEL_CARD.md, USE_POLICY.md and Responsible-Use-Guide.pdf remain the files to read before running the weights. It is the wrong starting point for anything newer than Llama 2, since the readme itself redirects to llama-toolchain, llama-models and llama-cookbook.
- Can I use it commercially?
- Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
- Is it still maintained?
- Probably not. The repository last received commits 20 months ago, on January 26, 2025.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The first section of the readme is a deprecation notice naming five replacements
The document opens with a note of deprecation rather than an installation guide. As part of the Llama 3.1 release the repositories were consolidated and Llama's scope widened into what the notice calls an end to end Llama Stack, and five repositories are named as the place to go instead. llama-models holds the foundation models with basic utilities, model cards, licence and use policies. PurpleLlama is the safety component, focused on safety risks and inference time mitigations. llama-toolchain carries the model development interfaces and canonical implementations, covering inference, fine-tuning, safety shields and synthetic data generation. llama-agentic-system is the standalone stack with an opinionated interface for agentic applications, and llama-cookbook is the community driven scripts and integrations repository. The instruction on where to file an issue follows the same logic: put it on one of those five, not here.
No GitHub releases and a last push on 2025-01-26 leave nothing to pin
The repository has no GitHub releases at all, and the last push to main landed on 2025-01-26. It is not marked archived, so the code is still there to read and still runs against Llama 2 checkpoints, but nothing has moved in it for a long stretch and the readme says as much by deprecating the whole document, whose second heading is literally the parenthesised word Deprecated next to Llama 2. What that leaves for someone arriving now is a tree with no version history to select from. There is no tag to install, no release notes to read, and nothing in the readme about which commit a tutorial was written against.
The packaging makes that gap concrete. The setup script declares the distribution name as llama at version 0.0.1, with packages discovered automatically and install requirements read straight out of requirements.txt. A version of 0.0.1 that has never moved is not a version anyone can upgrade from or down to, so a reproducible install means pinning a commit hash, and the state of that commit is whatever the last push left behind on 2025-01-26. For a repository whose own header says to use other projects, that is expected rather than a defect, but it does mean a new reader who finds this page first has to notice the deprecation before treating the code as current.
An editable install pulls torch, fairscale, fire and sentencepiece with no versions
The install step is one command, run in a conda environment that already has PyTorch and CUDA available.
pip install -e .What that resolves to is the contents of requirements.txt: torch, fairscale, fire and sentencepiece, four lines with no version constraints at all. The setup script turns that file into install_requires, so the environment ends up with whatever the resolver chooses on the day you run it, and a reader returning to the same instructions in a year can get a different set of builds without changing a character. The conda environment with PyTorch and CUDA is a precondition the readme states but the install does not create, so the four requirements are an addition to an environment you assemble yourself, not a working setup.
Licensing is stated in two places with different names. The header comment on the setup script says the software may be used and distributed according to the terms of the Llama 2 Community License Agreement, while the repository's own licence field records no assertion. The tree also carries MODEL_CARD.md, USE_POLICY.md and a Responsible-Use-Guide.pdf, which is where the terms and the acceptable use policy actually live.
The weights arrive through a signed link that expires in 24 hours
Nothing in the repository downloads weights by itself. You visit the Meta website, accept the licence, wait for an email containing a signed URL, and hand that URL to a script. The prerequisites are named as well: wget and md5sum have to be installed, and the script is invoked as its own command from the cloned directory.
./download.shThree operational details decide whether this works. The readme says the links expire after 24 hours and after a certain number of downloads, so a link saved in a setup document or a CI secret is a link with a clock on it. The documented failure is an HTTP error, `403: Forbidden`, and the stated remedy is to re-request a link rather than to change anything local, which is a confusing first symptom for anyone debugging a pipeline. And the readme warns against using the Copy Link option on the email, asking instead that the link be copied manually, because the browser's copy gives a different address than the one the script expects.
Model parallel values are fixed by size, and the cache is allocated before you need it
The examples are launched with torchrun, and the number of processes is not a tuning choice. A table gives the model parallel value per model size: 1 for the 7B, 2 for the 13B, and 8 for the 70B. The chat example shows the smallest of those.
torchrun --nproc_per_node 1 example_chat_completion.py \
--ckpt_dir llama-2-7b-chat/ \
--tokenizer_path tokenizer.model \
--max_seq_len 512 --max_batch_size 6The other two arguments in that line are placeholders the readme asks you to replace with your checkpoint directory and tokenizer path, and the same pattern repeats for the completion example, which points at the non-chat checkpoint and uses a shorter sequence length and smaller batch.
torchrun --nproc_per_node 1 example_text_completion.py \
--ckpt_dir llama-2-7b/ \
--tokenizer_path tokenizer.model \
--max_seq_len 128 --max_batch_size 4The reason to note the numbers is stated plainly: all models support sequence lengths up to 4096 tokens, but the cache is pre-allocated from the max_seq_len and max_batch_size values you pass, so those two arguments are a memory decision made before the first token, not a quality setting to raise later. A batch of six at 512 tokens is what fits on the hardware the example assumes.
Chat formatting is exact, and the reference is a line number on main
The fine-tuned chat models are trained for dialogue, and the readme is blunt that the expected features and performance only appear if a specific formatting is followed: the INST and <<SYS>> tags, the BOS and EOS tokens, and the whitespace and line breaks between them. It recommends calling `strip()` on inputs to avoid double spaces, which is the kind of detail that produces visibly worse output rather than an exception when it is wrong. The formatted variant is pointed at a specific function, chat_completion, with a link that lands on a line number inside llama/generation.py on the main branch.
That link is the fragile part. A permalink into a line of a branch that stopped moving on 2025-01-26 describes the code as of that push, and a reader who pip-installs the repository, copies the function into their own code, or follows the link after the file has been touched will not necessarily be looking at the same formatting rules. The pretrained, non-chat models have the opposite requirement, since they are not fine-tuned for chat or question answering and must be prompted so the expected answer is the natural continuation of the prompt, which is what example_text_completion.py is there to demonstrate.
Safety filtering lives in the cookbook, not in this example code
The readme mentions that additional classifiers can be deployed to filter inputs and outputs judged unsafe, then sends the reader to the llama-cookbook repository for an example of wiring a safety checker into inference code. That placement is the point. Nothing in this repository's two example scripts filters anything, and the `llama/` package holds the model code rather than a safety layer, so a deployment copied from here runs with whatever behaviour the checkpoint itself has and no additional guard.
The policy material is present at the top level, though. MODEL_CARD.md, USE_POLICY.md, Responsible-Use-Guide.pdf and UPDATES.md sit next to the code, and the readme's own closing section on the risks of the technology says testing to date has not and could not cover all scenarios. The gap is between that documentation and the runtime: reading the use policy is possible in this repository, implementing the classifier is a second repository's example script, and UPDATES.md is the file to check for what changed after launch. A team that adopts these examples for a deployed service should read all four before serving anything.
Editorial conclusion
This repository still answers one narrow question well, how a Llama 2 checkpoint is loaded and prompted, and its MODEL_CARD.md, USE_POLICY.md and Responsible-Use-Guide.pdf remain the files to read before running the weights. It is the wrong starting point for anything newer than Llama 2, since the readme itself redirects to llama-toolchain, llama-models and llama-cookbook. Before installing, expect a four-package editable install with no version pins, a signed download link valid for 24 hours, and no GitHub release to pin to, so record the commit you use.
Frequently asked questions
How do I run a Llama model locally using this repository?
Clone into a conda environment that already has PyTorch and CUDA, run `pip install -e .`, request the weights from the Meta website, pass the emailed URL to `./download.sh`, then launch an example with torchrun using the model parallel value for your model size, 1 for the 7B, 2 for the 13B, 8 for the 70B.
Should I still use the meta-llama/llama repository for new work?
The readme opens with a deprecation notice that redirects to five repositories, llama-models, PurpleLlama, llama-toolchain, llama-agentic-system and llama-cookbook, and the document itself is headed as deprecated for Llama 2, with the last push on 2025-01-26 and no GitHub releases.
What does installing the Llama inference code actually pull in?
Four unpinned packages from requirements.txt, torch, fairscale, fire and sentencepiece, installed as an editable install by a setup script that declares the name llama at version 0.0.1 and reads its install requirements from that file, on top of a PyTorch and CUDA environment you have to provide.
Why does downloading Llama 2 weights fail with 403 Forbidden?
The download URL is a signed link sent by email that the readme says expires after 24 hours and after a certain number of downloads, so the stated remedy is to re-request a link. The readme also asks that the link be copied manually from the email rather than with the browser's Copy Link option.