nlp-with-transformers/notebooks: What the Book's Example Code Actually Runs Today
Jupyter notebooks for the Natural Language Processing with Transformers book
At a glance
- What is it?
- The companion notebooks for Natural Language Processing with Transformers are pinned, partly unmaintained, and still the fastest way to run the book's PyTorch chapters. Here is what works, what broke, and how to set it up.
- Who is it for?
- Adopt these notebooks if you are working through the book's PyTorch chapters and want the code to match the text; clone the repository, create a Python 3.10 environment, and install from requirements.txt rather than installing transformers yourself. Do not adopt them if you need the TensorFlow code paths, if Chapter 7's Haystack pipeline is the part you care about, or if you want a library to depend on rather than teaching material.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 125 days ago.
- What is it written in?
- Mainly Jupyter Notebook, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem these notebooks solve: runnable code that matches the printed page
A book about transformer models is not much use if the listings in it have drifted away from the libraries they call. The repository holds the example code for Natural Language Processing with Transformers, published by O'Reilly, as a set of Jupyter notebooks numbered to match the chapters: 01_introduction.ipynb through 11_future-directions.ipynb, plus a second question-answering notebook. The audience is a reader who has the book open and wants to execute the same pipeline, on the same data, with the same hyperparameters, without retyping code from a screenshot.
That framing matters more than it sounds. The notebooks are not a library and they are not a framework. They are a fixed target: a specific transformers version, a specific datasets version, and a set of transitive pins that exist only so the book's code still imports. The requirements file says so directly, listing pyarrow<14, huggingface_hub<0.11, numpy<1.24, tokenizers<0.12, pandas<2 and protobuf<3.21 as pins added because newer releases break transformers 4.16.2. Anyone expecting a maintained toolkit will be disappointed. Anyone following the book will find the environment they need.
How the repository is organised, and why the pins are the interesting part
The layout is flat. Eleven chapter notebooks sit at the top level, alongside utils.py, install.py, plotting.mplstyle, environment.yml, and per-chapter requirement files for the one chapter that needs its own stack. A data/ directory and a SageMaker/ directory handle datasets and the AWS path. There is no package to import and no CLI to install; the unit of execution is the notebook.
The dependency strategy is the part worth understanding. Rather than tracking current releases, requirements.txt pins transformers[torch,sentencepiece,vision,optuna,sklearn,onnxruntime]==4.16.2 and datasets[audio]==1.16.1, then adds pins purely to keep them working on Python 3.10 or newer. The comments name the failure each pin prevents: pyarrow 14 and later drops PyExtensionType, which breaks datasets 1.16.1; huggingface_hub 0.11 and later drops the cached_download API that transformers 4.16.2 calls; numpy 2.x is ABI-incompatible with pyarrow 13; protobuf 3.21 and later break the sentencepiece and pegasus _pb2.py files shipped inside transformers. Chapter-specific pins follow the same logic, down to nltk==3.8.1 because 3.9 has a WordNetLemmatizer bug, and py7zr<0.20 because newer versions removed the SevenZipFile.readall method the samsum loader uses.
This is a deliberate trade: the environment is reproducible but frozen, and every pin is a small debt. The README is honest about the two places where the debt has already come due. TensorFlow sections are no longer tested, and the README states that code such as TFAutoModel.from_pretrained(..., from_pt=True) in Chapter 2 was written against TF 2.x with classic Keras 2 and breaks against current Keras 3. Chapter 7 is worse off: it depends on farm-haystack 0.9 or 1.4 with Elasticsearch 7.x and pydantic 1.x, a combination the README describes as incompatible with Python 3.10 and deprecated upstream. The PyTorch path is what the maintainers say they maintain.
Installing the environment and running the first notebook
The README points readers at cloud platforms first, on the grounds that most chapters need a GPU to run in reasonable time and the hosted environments ship with CUDA. Each chapter row in the README table carries Colab, Kaggle, Gradient and SageMaker Studio Lab badges that open the notebook directly from the repository, so the fastest first run needs no local setup at all.
If you would rather work locally, the repository ships an install script at its root:
python install.pyThe README does not spell out what install.py does beyond its presence in the repository, so treat requirements.txt as the authoritative list and read the script before running it on a machine you care about. Installing from the requirements file is the documented dependency source:
pip install -r requirements.txtOnce that finishes, open the introduction chapter in Jupyter and run the cells in order. The first cells import transformers and datasets and download a model checkpoint, so expect network access and a pause while the weights arrive. If the imports succeed, the environment matches what the book assumes; if you see an error about cached_download or PyExtensionType, a transitive pin has been overridden by something else in the environment. The README does not document a rollback or uninstall procedure for this environment, so keep the environment isolated from anything else you depend on.
Where these notebooks are the wrong tool
The clearest limitation is Chapter 7. The README states plainly that it is no longer maintained and that its stack, farm-haystack with Elasticsearch 7.x and pydantic 1.x, is incompatible with Python 3.10. The repository does include a second file, 07_question_answering_v2.ipynb, alongside requirements-chapter7-v2.txt and environment-chapter7.yml, but the README's guidance is to read the chapter for the concepts and follow current Haystack documentation for runnable code. If retrieval-augmented question answering is your actual goal, this repository is a reading aid, not a starting point.
The TensorFlow situation is a second boundary. The book contains optional TensorFlow paths, and the README says they were written against TF 2.x with classic Keras 2 and break against Keras 3. There is no supported migration path in the repository; the reader is told to install a compatible older TensorFlow and Keras themselves. If your team standardises on TensorFlow, the PyTorch notebooks will not map cleanly onto your stack.
There is a third, quieter issue: a pinned environment is not a production environment. The pins exist to keep a 2023-era transformers release importable, and several of them hold back packages that other parts of a modern project will want. Nothing here is designed to be vendored into an application, and the repository layout shows no test suite or CI configuration that would tell you a pin has stopped resolving. The last push to the repository was on 2026-05-29, so the pins are not abandoned, but they are also not something to build a service on.
How this differs from the Hugging Face course notebooks and from Haystack
The obvious alternative for learning the transformers library is the Hugging Face course, whose notebooks are maintained alongside the library itself. The difference is one of intent. Course notebooks track current releases and are updated when APIs change, which means the code you run is close to what you would write today, but it does not correspond line for line to a printed book. These notebooks do the opposite: they freeze a version so the text and the code agree, and they accept that some chapters stop working as the ecosystem moves.
For the retrieval material specifically, the alternative named in the README is Haystack itself. The README directs readers to the current Haystack documentation and notes that Deepset's project moved to haystack-ai 2.x with a substantially different API. That is a real difference in approach, not a rename: the chapter's retriever plus reader pipeline against Elasticsearch is the older design, and the current library is a different codebase. If you want to build a question-answering system rather than understand one, the current Haystack is the thing to read.
A third comparison is with the transformers library's own documentation and model pages. Those are the reference for a specific model or a specific pipeline call, and they are updated continuously. The notebooks are the reference for a worked example end to end, including data preparation, training loop, and evaluation, which the model pages do not attempt.
Licence, maintenance and the cost of upgrading
The repository is licensed under Apache-2.0, which permits commercial use, modification and redistribution provided the licence and notices are preserved. That covers the notebook code in this repository. It does not automatically cover the book text, the datasets the notebooks download, or the model checkpoints they pull from the Hugging Face Hub, each of which carries its own terms. Check those separately if you plan to reuse anything beyond the code itself. None of this is legal advice.
On maintenance, the repository is not archived and the last push was on 2026-05-29. That is recent enough that the pins are still being touched, but the README's own warnings about TensorFlow and Chapter 7 show that maintenance is selective rather than uniform. The PyTorch path is the maintained one.
The upgrade cost is the interesting question, because the repository has effectively chosen not to upgrade. Moving to a current transformers release would mean rewriting the notebooks against APIs that have changed, and the pins in requirements.txt exist precisely to avoid that work. If you want newer library behaviour, you are not upgrading these notebooks; you are porting them, chapter by chapter, and re-verifying the outputs against the book. That is a real project, not a version bump.
Editorial conclusion
Adopt these notebooks if you are working through the book's PyTorch chapters and want the code to match the text; clone the repository, create a Python 3.10 environment, and install from requirements.txt rather than installing transformers yourself. Do not adopt them if you need the TensorFlow code paths, if Chapter 7's Haystack pipeline is the part you care about, or if you want a library to depend on rather than teaching material. Before you commit time, verify three things: that your Python is 3.10 or newer, that a GPU is available for the training chapters, and that the chapter you need is not one of the two the README flags as unmaintained.
Frequently asked questions
Do I need a GPU to run the nlp-with-transformers notebooks?
The README states that most chapters require a GPU to run in a reasonable amount of time, and recommends cloud platforms because they come with CUDA pre-installed. The Colab, Kaggle, Gradient and SageMaker Studio Lab badges in the README table open each notebook directly on those platforms.
Why do the nlp-with-transformers notebooks pin transformers to 4.16.2?
The requirements file pins transformers[torch,sentencepiece,vision,optuna,sklearn,onnxruntime]==4.16.2 so the book's code matches the printed listings. It then adds transitive pins such as pyarrow<14 and huggingface_hub<0.11 specifically to keep that release importable on Python 3.10 or newer.
Is the TensorFlow code in the nlp-with-transformers notebooks still tested?
No. The README states that TensorFlow sections are no longer tested, that they were written against TF 2.x with classic Keras 2, and that they break against current Keras 3. The PyTorch path is the one the maintainers say they maintain.
Can I still run Chapter 7 of the nlp-with-transformers notebooks?
The README states that Chapter 7 is no longer maintained because it relies on farm-haystack 0.9 or 1.4 with Elasticsearch 7.x and pydantic 1.x, a stack it describes as incompatible with Python 3.10 and deprecated upstream. It suggests reading the chapter for the concepts and following the current Haystack documentation for runnable code.
What licence applies to the nlp-with-transformers notebooks?
The repository is licensed under Apache-2.0. That covers the notebook code here, but the book text, the datasets the notebooks download and the model checkpoints they fetch each carry their own terms.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/nlp-with-transformers-notebooks)