Llama Cookbook: Meta's Official Repository for Building with Llama Models
Welcome to the Llama Cookbook! This is your go to guide for Building with Llama: Getting started with Inference, Fine-Tuning, RAG. We also show you how to solve end to end problems using Llama model family and using them on various provider services
At a glance
- What is it?
- Llama Cookbook is Meta's official reference repository for developers building with the Llama model family, covering inference setup, fine-tuning, RAG, third-party provider integrations, and a growing set of end-to-end use-case notebooks.
- Who is it for?
- Llama Cookbook is the authoritative starting point for engineers integrating any version of the Llama model family, from Llama 2 through Llama 4. The repository is organized well enough that a reader can go directly to the `getting-started/` or `end-to-end-use-cases/` directory most relevant to their task without reading the whole collection.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 133 days ago.
- What is it written in?
- Mainly Jupyter Notebook, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
DEEP OPEN-SOURCE ANALYSIS
What Llama Cookbook Is and Who Should Use It
Llama Cookbook is a Jupyter Notebook-heavy repository that serves as the companion reference for engineers who want to run, fine-tune, or build applications with Llama models from Meta. It is not a standalone library in the typical sense. The installable `llama-cookbook` package provides the `src/llama_cookbook/` utilities carried over from the earlier llama-recipes project, but most of the value is in the notebooks and integration guides rather than importable Python code.
The target audience is machine learning engineers and application developers who already have access to Llama model weights through Hugging Face (at meta-llama/), the Llama API, or a third-party provider like AWS, Azure, or Google. The repository assumes familiarity with Python, PyTorch, and the Hugging Face ecosystem. It does not teach foundational deep learning.
The repository was previously called llama-recipes. According to the README's FAQ, the rename to llama-cookbook happened before the current main branch; an `archive-main` branch preserves the repository state from before the refactor.
Repository Structure and Navigation
The top-level layout divides content into four main directories:
`getting-started/` contains reference notebooks for inference and fine-tuning. The README highlights a notebook for getting started with the Llama API, one for Llama 4 Scout's 5-million-token long context capability, and one for Llama 4 Maverick.
`end-to-end-use-cases/` holds multi-notebook projects spanning specific application domains. Named examples include a WhatsApp bot powered by Llama 4, a research paper analyzer using Llama 4 Maverick, a book character mind map builder, and a SearchQA integration.
`3p-integrations/` provides getting-started guides and use-case recipes for running Llama through third-party providers. This is the section to check before setting up your own inference stack if you plan to use a managed provider.
`src/` contains the original llama-recipes library code along with a fine-tuning FAQ in `src/docs/`. The `recipes/` directory holds additional miscellaneous recipes.
The `requirements.txt` lists core dependencies: `torch>=2.2`, `accelerate`, `transformers>=4.45.1`, `peft`, `datasets`, `bitsandbytes`, `faiss-gpu` (for Python below 3.11), `sentence_transformers`, and several other libraries for evaluation, document processing, and the Gradio demo UI.
Installing and Using the Package
The installable package is named `llama-cookbook` and requires Python 3.8 or newer. The `pyproject.toml` defines several optional extras: `vllm` for the vLLM inference backend, `langchain` for LangChain integration (which pulls in `langchain_openai`, `langchain`, and `langchain_community`), `tests` for the pytest mock dependencies, `searchqa` for SearchQA data materialization, `docs` for the MkDocs documentation site, and `auditnlg` for the audit-NLG evaluation utilities.
Most notebooks in the repository require a GPU for practical use. The fine-tuning examples in particular depend on bitsandbytes for quantization and PEFT for parameter-efficient methods like LoRA. The base `requirements.txt` pulls in `torch>=2.2`, `transformers>=4.45.1`, `accelerate`, `peft`, `datasets`, `bitsandbytes`, `sentence_transformers`, and a number of evaluation libraries.
Cloning the repository directly is necessary to access the notebooks, end-to-end use cases, and third-party integration guides. These are not distributed through the pip package. The README points to the Llama API notebook (`getting-started/build_with_llama_api.ipynb`) as the recommended first stop for new users.
Fine-Tuning Coverage and the Fine-Tuning FAQ
Fine-tuning is a major theme of the repository. The `getting-started/finetuning` directory contains reference implementations. The `src/docs/` subdirectory holds a fine-tuning FAQ that the main README FAQ points to for questions about model adaptation.
The `requirements.txt` includes the standard set of libraries for parameter-efficient fine-tuning: PEFT for LoRA and adapter methods, bitsandbytes for 4-bit and 8-bit quantization, and `accelerate` for distributed training. The `torch>=2.2` constraint means older PyTorch installations need upgrading before running fine-tuning notebooks.
The repository covers fine-tuning for both text and multimodal use cases, given that Llama 4's Scout and Maverick models are vision models. The README highlights Llama 4 Maverick specifically for the research paper analyzer example, which involves processing document content.
Llama 4 Recipes and Long-Context Support
The most recent additions in the README highlight Llama 4-specific recipes. The getting-started notebook for Llama 4 Scout demonstrates 5-million-token long context usage with the Llama API. The `end-to-end-use-cases/` section includes a WhatsApp integration and a research paper analyzer as illustrative full-stack examples.
The Llama API is referenced repeatedly in the README as a starting point for new users: `getting-started/build_with_llama_api.ipynb` is listed as the primary onboarding notebook. This implies that the simplest path to running Llama 4 recipes uses the hosted Llama API rather than a local model deployment.
Licensing for Llama models is version-specific. The README lists separate license files and Acceptable Use Policies for Llama 4, Llama 3.3, Llama 3.2, Llama 3.1, Llama 3, and Llama 2. Each version has its own terms. Using a model from a different version than the one listed for a recipe may require verifying the applicable license separately.
Limitations and Cases Where This Repository Falls Short
Llama Cookbook is not a production deployment guide. It demonstrates how to run and fine-tune models in a notebook environment. Converting a notebook experiment into a production service with monitoring, rate limiting, cost tracking, and fault tolerance requires work that the repository does not cover.
The repository has no stable API contract. Notebooks are illustrative rather than maintained library code. A notebook that worked at the time of its last commit may depend on a specific version of `transformers` or a specific Hugging Face model card that has since changed.
The last push was on 2026-05-19. For teams working with the most recently released Llama model variants or provider integrations added after that date, the repository may be missing relevant examples. The `archive-main` branch also means that users who followed older llama-recipes tutorials may encounter structural differences when switching to the current repository layout.
Fine-tuning Llama models requires significant GPU memory. The README does not specify minimum hardware requirements; individual notebooks include their own requirements, but this is not consolidated anywhere in the documentation.
Llama Cookbook vs. OpenAI Cookbook
OpenAI Cookbook is the equivalent reference repository for OpenAI's APIs, covering GPT models, embeddings, fine-tuning, and function calling with Python examples. The structural analogy is close: both are notebook-heavy reference repositories maintained by the model provider.
The practical difference is the starting point. OpenAI Cookbook assumes API access through an API key and focuses on API integration patterns. Llama Cookbook covers both managed API usage (through the Llama API) and self-hosted deployment, including local weight loading, quantization, and fine-tuning on your own hardware. Engineers who want to run models locally or modify model weights at the weight level have a reason to use Llama Cookbook that OpenAI Cookbook cannot address.
Maintenance, Licensing, and Contribution
Llama Cookbook is maintained by a team at Meta. The `pyproject.toml` lists authors with meta.com email addresses. The last push was on 2026-05-19, which means the repository has been relatively quiet in the four months preceding the date of this article.
The package code in `src/` is licensed under MIT. The model weights accessed through the recipes are not MIT licensed. Each Llama version has its own custom license with an Acceptable Use Policy. The README links to each model's LICENSE file in the meta-llama/llama-models repository. Using a fine-tuned model commercially requires checking both the base model's license and any terms imposed by the provider distributing it.
Contributions are accepted through pull requests. The CONTRIBUTING.md file in the repository describes the process and code of conduct.
Editorial conclusion
Llama Cookbook is the authoritative starting point for engineers integrating any version of the Llama model family, from Llama 2 through Llama 4. The repository is organized well enough that a reader can go directly to the `getting-started/` or `end-to-end-use-cases/` directory most relevant to their task without reading the whole collection. Teams starting from a provider rather than from raw Llama weights should check the `3p-integrations/` folder for provider-specific examples before setting up their own inference stack. The last push was on 2026-05-19; for the most current Llama 4 recipes, verify the notebook file timestamps before running them.
Frequently asked questions
How do I install llama-cookbook?
Run `pip install llama-cookbook` for the base package. Optional extras for vLLM, LangChain, and testing tools are available via pip extras syntax. Cloning the GitHub repository is necessary to access the notebooks and integration guides, which are not part of the pip package.
What is the relationship between llama-cookbook and llama-recipes?
The README FAQ explains that llama-recipes was renamed to llama-cookbook. The `archive-main` branch in the repository preserves the state before the refactor. The current main branch reflects the post-refactor structure with `getting-started/`, `end-to-end-use-cases/`, and `3p-integrations/` as the top-level content directories.
Does llama-cookbook cover fine-tuning Llama models?
Yes. The `getting-started/finetuning` directory contains fine-tuning reference notebooks, and `src/docs/` holds a fine-tuning FAQ. The requirements include PEFT, bitsandbytes, and accelerate, which are the standard libraries for parameter-efficient fine-tuning with quantization.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/metainternal-llama-cookbook)
Community notes