HuggingLLM (ButterflyBook): a Chinese-language course for building on LLM APIs
HuggingLLM, Hugging Future.
At a glance
- What is it?
- HuggingLLM is a Datawhale open course that teaches ChatGPT-era API work (embeddings, classification, generation, reasoning) to programmers who are not NLP specialists. It is a book and notebook collection, not a library, and the README is explicit about what it refuses to cover.
- Who is it for?
- Adopt HuggingLLM if you can already program and want a structured path through embedding search, classification, summarization and reasoning with an LLM API, and you read Chinese. Skip it if you want to train or fine-tune models from scratch, or if you need an English-language curriculum.
- Can I use it commercially?
- Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
- Is it still maintained?
- Yes. The repository last received commits 105 days ago.
- What is it written in?
- Mainly Jupyter Notebook, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 28, 2026, and from our analysis. They are not legal advice.
DEEP OPEN-SOURCE ANALYSIS
What HuggingLLM is, and the reader it is written for
HuggingLLM is a Datawhale project whose README calls it a course introducing the principles, use and applications of ChatGPT, aimed at lowering the barrier so that people outside NLP or algorithm roles can build something with an LLM. The repository is a Jupyter Notebook project, and the README splits the material in two: docs/ holds the ebook version, content/ holds the notebooks in their original and iterated forms. The stated audience is people who are interested in ChatGPT, want to apply it to a new service or an existing problem, and have some programming background. The README also lists who should not bother: anyone researching low-level algorithm details such as how PPO is implemented or whether it could be swapped for NLPO or ILQL, anyone who wants to build a ChatGPT from scratch, and anyone after other technical details. That is an unusually direct scope statement, and it is the single most useful thing on the page. If you are an algorithm engineer looking for training internals, this is the wrong repository and the README says so before you clone anything.
The chapter structure: four usage guides plus context chapters
The table of contents is organised as a short theory chapter followed by four usage guides and three context chapters. Chapter 1 covers LM, Transformer, GPT and RLHF at a survey level. Chapters 2 through 5 are the practical core: similarity matching (embeddings, API use, QA tasks, clustering, recommendation), sentence and word classification (NLU basics, API use, document QA, classification and entity recognition fine-tuning, dialogue applications), text generation (summarization, error correction, machine translation) and text reasoning (what reasoning is, importing ChatGPT, testing its reasoning ability, calling it, and comparing ChatGPT with GPT-4). Chapters 6 to 8 cover engineering practice (evaluation, safety, networking), limitations (factual errors, real-time updates, resource cost) and commercial applications across search, office, education, gaming, music, retail, advertising, media, finance, healthcare, design, film and industry. The README states the chapters are relatively independent, so you can read one or work through all of them. The practical consequence is that this is closer to a set of standalone tutorials than a progressive curriculum, and the embedding chapter does not assume you finished the RLHF chapter.
Installing nothing: the setup is an API key and a notebook
There is no package to install for the project itself. The README's learning guide says you need working access to the OpenAI API and the ability to call gpt-3.5-turbo, or a domestic Chinese LLM API, or an open-source model. It also says you can have no algorithm experience but should have some programming background or project experience, and it budgets 2 to 3 days per usage-guide chapter, 6 to 8 hours each, with text reasoning as the exception. The repository itself is cloned or read online; the docs site at datawhalechina.github.io/hugging-llm is the ebook build. For the domestic-API path, the README gives two worked examples. Zhipu GLM installs its SDK and is called like this:
pip install zhipuaiThe Python call constructs a ZhipuAI client with an API key, sends a system and a user message, and streams the response. The model name in the README's example is glm-4, with a note to check the official documentation for the model you need. Qwen takes a different SDK and a different call shape:
pip install dashscopeThe dashscope example uses dashscope.Generation.call with dashscope.Generation.Models.qwen_max, an api_key argument, result_format='message', stream=True and incremental_output=True, and it checks response.status_code against HTTPStatus.OK before appending to the output string. Both examples end with a printed answer, and the README shows the model's reply about Datawhale as the expected output. What you should see after running either is a streamed answer accumulated into one string. Note that the README tells you to fill in your own API key in both snippets; there is no key management, no .env file and no config file documented. That is fine for a tutorial and thin for anything else.
Where the material stops: no evaluation harness, no cost model
The engineering-practice chapter lists evaluation, safety and networking as topics, but the README does not describe a harness, a test set, a scoring script or any reusable component you could drop into a pipeline. The chapters are teaching material. If you need to measure whether a prompt change regressed your classification accuracy, this repository will explain the idea and not hand you the tooling. The same applies to cost: chapter 7 names resource consumption as a limitation, and the README gives no token accounting, no price table and no budget guidance. The learning guide's time estimates are for a human reading, not for a service. A second limitation is language. The README, the chapter titles and the linked content are in Chinese, and the two API examples are written for Chinese providers. If your team works in English, the docs site will be a translation exercise before it is a tutorial. A third constraint is that the API surface is the moving part. The README pins its examples to gpt-3.5-turbo, glm-4 and qwen_max, and it explicitly tells you to consult official documentation for model names, which is an admission that the snippets age faster than the prose.
How it compares with LangChain and the Hugging Face course
LangChain is the obvious alternative, and the difference is categorical. LangChain is a Python library you install and call; it provides chains, retrievers, memory and provider abstractions so you can assemble an application. HuggingLLM is a set of notebooks and prose that teaches you what to assemble and why, then leaves the assembly to you. If your problem is that you already know what a retrieval pipeline should do and you want less boilerplate, LangChain is the right shape. If your problem is that you do not yet know why an embedding-based QA system works or when classification beats generation, the notebooks are the right shape. The Hugging Face course is the other comparison point, and the split is language and provider focus. The Hugging Face material is in English and centred on the Hugging Face ecosystem; HuggingLLM is in Chinese and, per its README, deliberately centred on API usage with OpenAI, Zhipu GLM or Qwen, with a stated emphasis on building applications rather than on model internals. Neither is a superset of the other, and the choice is mostly about which language you read and which provider you can actually call.
Maintenance, licence and what the repository does not state
The repository is not archived, and the last push was on 2026-06-16, which is recent enough that the project is still being touched. There are no retrieved releases, so there is no versioned changelog to track and no upgrade path to plan around: you consume whatever the main branch holds. The README's own note that docs/ is the ebook version but not the final version, with possible small edits by the editor, and that content/ holds the initial and iterated notebooks, means the two trees can drift. There is no documented migration note for that drift. The licence is recorded as NOASSERTION, which means GitHub could not match the LICENSE file to a known licence identifier. The repository does contain a LICENSE file at the top level, but the README's licence section is not reproduced here, so the actual terms are not something this article can state. Read the LICENSE file in the repository root yourself before reusing the notebooks or the book text in anything you ship, and treat the NOASSERTION label as a prompt to look rather than a verdict. This is a description of the metadata, not legal advice.
Editorial conclusion
Adopt HuggingLLM if you can already program and want a structured path through embedding search, classification, summarization and reasoning with an LLM API, and you read Chinese. Skip it if you want to train or fine-tune models from scratch, or if you need an English-language curriculum. Before committing time, open the docs site, confirm the chapter you need exists in both content/ and docs/, and check which API your account can actually call, since the README's examples assume OpenAI, Zhipu GLM or Qwen access.
Frequently asked questions
What is HuggingLLM?
HuggingLLM is a Datawhale open course, also called ButterflyBook, that introduces the principles, use and applications of ChatGPT. Its README says it targets people who are interested in ChatGPT and want to apply it, with some programming background, rather than NLP or algorithm specialists.
Is the HuggingLLM course free?
The material is published openly on GitHub and as an ebook at datawhalechina.github.io/hugging-llm, and the README does not describe a paid tier. What is not free is the API access the course assumes: the learning guide requires working OpenAI API access, a domestic Chinese LLM API, or an open-source model, and you supply your own API key.
How do you use an LLM API with HuggingLLM?
The README gives two worked paths. For Zhipu GLM you install the zhipuai SDK, create a ZhipuAI client with your API key, and send a system and user message with stream=True; for Qwen you install dashscope and call dashscope.Generation.call with qwen_max, result_format='message' and incremental_output=True. Both examples accumulate the streamed chunks into a single string and print it.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/datawhalechina-hugging-llm)
Community notes