# ed-donner/llm_engineering: an eight-week course repository you run yourself

> The repository is the code and setup companion to Ed Donner's LLM engineering course, organised as week1 through week8 with a uv-managed environment. It is teaching material, not a library, and the README's own warning about model size is the first thing to read.

**ed-donner/llm_engineering** — Repo to accompany my mastering LLM engineering course

- Repository: https://github.com/ed-donner/llm_engineering
- Stars: 7,541 · Forks: 7,358
- Language: Jupyter Notebook
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/ed-donner-llm-engineering

## What the llm_engineering repository actually is

This is a course companion, not a package you import. The README describes it as the repo to accompany a mastering LLM engineering course, and the top-level layout confirms the shape: week1 through week8, plus guides, setup, extras, assets and community-contributions. Each week is a module, and the README states that the projects build on each other, culminating in an agentic AI solution in week8 that draws on the earlier weeks.

The audience is therefore narrow and specific. It is someone following the lectures who wants to run the same cells, inspect the objects, and then change the code. The README is explicit that the instructor executes the code rather than typing it, and that the learner should work along or afterwards. If you are looking for a retrieval library, an agent framework or a CLI, this repository is the wrong artefact. There is no console entry point in pyproject.toml and no published distribution name beyond the local project name llm-engineering.

The dependency list is the clearest signal of what the course covers. It spans langchain and langchain-chroma, chromadb, sentence-transformers, transformers, torch, openai, anthropic, google-genai, google-generativeai, groq, ollama, litellm, gradio, plotly, matplotlib, scikit-learn, xgboost, pandas and wandb. That is a broad tour rather than a minimal stack, which fits a teaching goal and works against anyone hoping for a lean runtime.

## How the weeks and the environment fit together

The mechanism is straightforward: one repository, one environment, eight week folders. environment.yml and pyproject.toml sit at the top level alongside uv.lock, so the intended workflow is a single resolved environment that all weeks share. The presence of both a conda-style environment file and a uv lock file means the course material supports more than one setup path, and setup/SETUP-new.md is the document the README sends you to for the full environment.

Data flow is notebook-driven. You open a week folder, run cells, and the cells call out to whichever provider the lesson uses. The provider list in pyproject.toml matches that: OpenAI, Anthropic, Google, Groq, Ollama and LiteLLM as a router. The README also points to guides/09_ai_apis_and_ollama.ipynb for the free alternative path, describing it as containing exact code for Ollama, Gemini, OpenRouter and more. So the abstraction layer for model choice is a guide notebook plus LiteLLM, not a configuration file you edit once.

Two details in pyproject.toml deserve attention because they constrain the environment. The Python floor is >=3.11, and protobuf is pinned to exactly 3.20.2 while most other dependencies use >= ranges. A hard pin on protobuf inside a large dependency graph is the kind of thing that produces resolver conflicts when you add your own packages, and the repository does not document a workaround. dataloaders and datasets are also pinned, with datasets==3.6.0.

## Installing the environment and running the first Ollama cell

The README puts an instant-gratification exercise before the full setup, and it deliberately uses llama3.2 rather than the newer llama3.3. The README states the reason plainly: llama3.3 at 70B parameters is too large for most home computers, and it notes that several students have missed the warning. Step one is installing Ollama from ollama.com, with a note that Windows may need administrator permissions.

```bash
ollama run llama3.2
```

On a smaller machine the README gives a reduced variant instead:

```bash
ollama run llama3.2:1b
```

If the model does not start, the README says to run the server in a second terminal and retry, and that Windows users may need an admin PowerShell:

```bash
ollama serve
```

After that first run, the README directs you to setup/SETUP-new.md for the real environment across all platforms. The repository ships uv.lock and pyproject.toml, so a uv-based install is the path the files imply, and environment.yml covers the conda route. The README does not spell out the exact install command in the text reproduced here, so treat SETUP-new.md as authoritative rather than guessing at flags.

For anyone who cannot get Ollama working locally, the README provides a Google Colab notebook as a cloud fallback, free with a Google account. That is the only hosted option the README offers directly.

## The cost model, and why the free path matters

The README is unusually direct about spend. It says API costs are optional, that the frontier-model experiments cost a few cents at a time, and that the whole course should not require more than a couple of dollars. It also warns that some providers, naming OpenAI, require a minimum credit around $5 or the local equivalent. Week 7 is flagged as the one point where you might choose to spend more; the README mentions spending about $10 there and says it is not necessary.

This matters because the repository mixes paid and free providers in the same dependency list. Nothing in pyproject.toml tells you which weeks need which key. The README's answer is guide 9 in the guides folder, which it says contains the exact code for Ollama, Gemini, OpenRouter and more. If you want to complete the course without a card on file, that notebook is the entry point, and the Ollama quick start is the proof that a local model is enough for the first day.

The risk here is not the money, it is the monitoring. The README asks you to watch your API usage and links to provider dashboards, but it does not describe any budget cap or alerting inside the repository. There is no cost guard in the code layout. A runaway loop in a notebook bills against your key, and the repository's only defence is the README's advice.

## Where this repository is the wrong tool

Three cases stand out. First, if you need a maintained library with a versioned API, this is not it. The version is 0.1.0, there are no retrieved releases, and the README frames everything as educational. The README even says the projects are first and foremost designed to be educational. Adopting course notebooks as production scaffolding means inheriting dependencies like gradio, plotly, wandb and xgboost that your service probably does not need.

Second, the environment is heavy. torch, transformers, sentence-transformers, chromadb and scikit-learn in one resolution is a large install, and the exact protobuf pin at 3.20.2 will fight any project that needs a newer one. If you only want to call a hosted model, you are paying for a lot of unused surface area.

Third, the support channel is external. Questions go to Udemy or to ed@edwarddonner.com, per the README, and there are no retrieved releases to track. Anyone who needs a changelog, a deprecation policy or a security contact will not find one in the repository. The community-contributions folder exists and the README invites pull requests, but that is contribution, not maintenance commitment. Note also that the README's FAQ links point at an avatar URL and a resources page rather than in-repo documentation, so the answers live off-repository.

## Compared with a framework-first approach such as LangChain alone

The obvious alternative is to skip the course structure and learn the same ground through a framework's own documentation, LangChain being the one this repository leans on most. The difference in approach is real. LangChain's documentation is organised by capability: you look up a retriever, a chain, an agent, and you assemble. This repository is organised by week, and the README says the projects build on each other, so week8's agentic solution reuses what earlier weeks introduced.

That ordering is the whole point and also the limitation. A framework's docs let you jump to the piece you need today. A course sequence makes you carry the earlier context. If your goal is to ship a retrieval feature next sprint, working through week1 to week4 first is overhead. If your goal is to understand why the retrieval step behaves the way it does, the sequence is the value.

The dependency list shows the repository is not purist about this either. It pulls in langchain, langchain-community, langchain-experimental, langchain-openai, langchain-anthropic, langchain-ollama, langchain-huggingface, langchain-chroma and langchain-text-splitters, plus litellm as a separate router. So the course teaches LangChain while also showing a provider-agnostic path through LiteLLM and raw SDKs. That breadth is a teaching choice, and it is why the environment is as large as it is.

## Licence, maintenance and what to verify before you start

The repository is MIT licensed, with the LICENSE file at the top level. MIT is permissive: you can reuse the code, including in commercial work, provided you keep the copyright notice and the licence text. That applies to the code in the repository, not automatically to any third-party course content, slides or assets that the README links to on external sites. Nothing here is legal advice; if you plan to redistribute the notebooks or the assets folder, read the LICENSE file yourself and check the terms of any external resource the README points to.

On maintenance, the last push was on 2026-09-05. The repository is not archived. There are no retrieved releases, so there is no version history to read and no upgrade path documented beyond pulling the latest commit. Because the environment is lock-file driven, an upgrade means re-resolving pyproject.toml, which is where the protobuf pin and the datasets pin are most likely to bite.

Before you commit time, verify four things: that your Python is at least 3.11, that you have read setup/SETUP-new.md rather than inferring steps, that you can run llama3.2:1b locally or have the Colab fallback open, and that you are comfortable with the provider keys the later weeks assume. If any of those four is a blocker, the course material will not resolve it for you.

## Conclusion

Adopt it if you are working through Ed Donner's course and want the notebooks, environment file and setup guides in one place; skip it if you want an installable package, since pyproject.toml declares no library surface and the README points elsewhere for questions. Before you spend anything, read setup/SETUP-new.md, check your Python version against the >=3.11 floor in pyproject.toml, and confirm your machine can run llama3.2:1b rather than the 70B llama3.3 the README warns against.

## FAQ

### What is LLM engineering?

The repository does not define the term directly; it teaches it through eight weekly modules that move from a local Ollama model to an agentic AI solution in week8. The dependency list shows the working surface: prompt and model calls through OpenAI, Anthropic, Google, Groq and Ollama, retrieval with chromadb and langchain-chroma, and fine-tuning tooling with transformers and torch.

### How can I become an LLM engineer?

The README frames the path as eight weeks of building projects that build on each other, with the mantra that the best way to learn is by doing. It tells you to run each cell, inspect the objects, then tweak the code, and it invites pull requests to the community-contributions folder.

### Is LLM actually an AI?

The repository does not address that question. Its material is about building with large language models rather than about the terminology debate, so there is nothing here to answer it.

### What is the difference between LLM engineering and AI engineering?

The repository does not draw that distinction. Its scope is stated as mastering LLM engineering, and the week folders cover model calls, retrieval and an agentic solution, with no section comparing the two terms.

## Sources

- [ed-donner/llm_engineering on GitHub](https://github.com/ed-donner/llm_engineering)
- [Issues](https://github.com/ed-donner/llm_engineering/issues)
- [License: MIT](https://github.com/ed-donner/llm_engineering/blob/main/LICENSE)
- [README](https://github.com/ed-donner/llm_engineering/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/ed-donner-llm-engineering
