Model or dataset
voidful/hung-yi-lee-skill avatar
voidful/hung-yi-lee-skill

hung-yi-lee-skill: a Claude skill that answers in Hung-yi Lee's teaching style

蒸餾李宏毅老師的skill,結合Karpathy的LLM,Fable 5和GPT6 加持 以及 本人親自訪談

1,292 stars128 forksHTMLLicense varies

At a glance

What is it?
This repository distils a Taiwan machine learning lecturer's teaching structure, verbal tics and concept graph into a SKILL.md that an AI coding assistant reads. The interesting part is the calibration interview; the weak part is that almost all of it is in Traditional Chinese.
Who is it for?
Adopt it if you already run a coding assistant that reads a SKILL.md and you want machine learning explanations in Traditional Chinese that follow a lecturer's structure rather than a generic persona prompt. Skip it if you need English output, if you need coverage beyond the 27 cached transcripts, or if you cannot read the SKILL.md and references/ files to judge the calibration claims for yourself.
Can I use it commercially?
Not without permission. GitHub finds no licence file in the repository, and without a licence all rights are reserved by default: you may read the code but not reuse it. Check the README, or ask the authors, before using it.
Is it still maintained?
Yes. The repository last received commits 52 days ago.
What is it written in?
Mainly HTML, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 2, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What the skill actually distils, and who it is for

Most attempts at making a chatbot talk like a specific teacher stop at a persona prompt: tell the model it is person X and hope the tone sticks. This repository argues that the difference is source material. The README states the skill was built from 478 YouTube video metadata records, 27 full transcripts, 8 topic pages and 4 curated research references, and that these produced a knowledge graph of 916 nodes and 10 concept communities. The claim is not that the model knows more; it is that the model has a documented path back to a specific lecture for a specific answer.

The intended user is someone who already reads Hung-yi Lee's course material and wants an assistant that explains machine learning the way those lectures do. The README's own example prompts are in Traditional Chinese and ask for a transformer explanation in the lecturer's style, or for his reading of an AI safety report. That framing matters: this is a niche skill for a Chinese-language audience that follows the course, not a general purpose teaching assistant. If you have never watched the lectures, the value proposition is thinner, because the skill's authority rests on traceability to transcripts you would not recognise.

The three-stage build: frequency mining, interview script, then the lecturer himself

The README describes the construction in three steps, and each step exists to fix a limit in the previous one. The first stage fed 27 transcripts through a model the README calls Fable 5 (`claude-fable-5`) to count verbal habits rather than guess at them. The numbers given are concrete: `比如說` appears 609 times, `假設` 518, `而已` 160, `就結束了` 36. Whether or not you care about those particular particles, the method is the point. Frequency counts are checkable in a way that an impression of someone's style is not.

The second stage addressed what the README calls the negative space problem. A transcript only records what was said on camera, so it cannot tell you which jokes were tried and dropped or what the lecturer deliberately avoids. The response was to generate an interview script with a multi-agent workflow, drawing on SKILL.md, the transcripts and golden/negative examples. That script lives at `references/interview-protocol.md` and reportedly contains quote-reaction prompts, reverse-generation prompts and ten A/B comparison questions grounded in the corpus.

The third stage is the one that separates this project from a scraping exercise. The authors interviewed the lecturer, and his answers were placed above everything inferred from transcripts. The README's table of corrections is the most useful page in the repository: the skill had assumed a phrase like 「熱騰騰的文件」 was a good opener, and the lecturer said he would not use that word. It had assumed a school-daily-life analogy for a scary topic, and he said that was too plain and that he would reach for a more absurd anime comparison. It had assumed converting a number into "one working day" was concrete enough, and he said it was not, that the cost and scarcity have to be felt. Where transcript inference and the interview disagree, the interview wins.

Installing the skill and running a first graph query

The README's install section is short. It clones the repository, installs the Python dependencies, and then tells you to place the directory somewhere your AI coding assistant can read it and point the assistant at `SKILL.md`. There is no package on PyPI, no Docker image and no installer script.

bash
git clone https://github.com/voidful/hung-yi-lee-skill.git
cd hung-yi-lee-skill
pip install -r requirements.txt

The dependency list is small and worth reading before you run it, because it tells you what the tooling actually does: `youtube-transcript-api>=1.2,<2` for fetching captions, `networkx>=3.4` and `python-louvain>=0.16` for building and partitioning the graph, and `yt-dlp` as a fallback fetcher. Nothing here is a model runtime. The skill itself is text plus a graph; the language model is whatever assistant you point at it.

After cloning, the README says to let the assistant read `SKILL.md`. The example prompts it gives are of this shape:

code
> 用李宏毅老師的風格幫我解釋 transformer
> 老師會怎麼看這份 AI safety report?
> 什麼是 self-supervised learning?像老師上課那樣講

What you should see is an answer that opens with an intuition, states the black box input/output/objective, then opens the mechanism, then names the trap, then recaps. The README lists that five-part structure explicitly, and it is the part of the skill most likely to be visible in a single response. If the answer opens with a definition and a bulleted list, the skill is not being applied.

The knowledge graph has its own CLI, which is the piece you can exercise without involving a language model at all:

bash
python3 scripts/hungyi_kb.py graph query "attention mechanism"
python3 scripts/hungyi_kb.py graph query "語音模型"
open wiki/graph/graph.html

The first two commands query the persisted graph; the third opens the interactive visualisation. The README also documents a sync command for channel metadata, though the excerpt cuts off mid-line, so treat the sync path as documented but not fully specified here.

The knowledge graph is the strongest artefact and the least explained one

The graph report gives numbers that are unusually specific for a project like this: 916 nodes, 3,664 edges, 10 communities, with edges split into 1,621 EXTRACTED and 2,043 INFERRED. That split is the honest detail. Roughly 44 percent of the relationships in the graph were inferred rather than extracted from the corpus, and the README does not say what produced the inference or how it was validated. A user querying the graph has no way to tell from the command line whether a given edge came from a transcript or from a model's guess.

The community breakdown is more legible. ML Fundamentals holds 396 nodes, Diffusion And Generation 116, Speech And Audio 81, Evaluation 79, Agents 72, Model Editing 33. The god-node table lists ML Fundamentals at 385 connections, 語言模型 at 251, Llama at 101, Transformer at 83. Those figures suggest a graph weighted heavily toward introductory material, which fits a corpus drawn from a lecture series. It also means the Agents and Model Editing communities are small enough that graph queries in those areas are likely to return thin results. The README does not document a coverage threshold or a fallback when a query hits a sparse community.

Where this breaks down: language, coverage and the licence question

The first limitation is language. The README is in Traditional Chinese with an English translation at `README.en.md`, and the skill's own examples, verbal markers and teaching-structure vocabulary are Chinese. The style markers the project counted (`比如說`, `假設`, `而已`) are Chinese particles. A user who wants English explanations of machine learning is asking the skill to do something its source corpus does not support, and the README does not claim otherwise.

The second is corpus size. Twenty-seven full transcripts against 478 video metadata records is a small fraction, and the README says the transcripts are "持續擴充中" (still being expanded). Any topic outside those 27 transcripts has no traceable source, which undercuts the skill's central promise that answers can be traced to a specific lecture. The graph's 916 nodes may cover more ground than the transcripts, but nodes without transcript backing are exactly the ones you cannot check.

The third is the licence. The repository metadata shows no licence, and the README does not state one either. The README does say the lecturer authorised the description of his teaching style and that the skill "模仿教學方法,不冒充本人" (imitates the teaching method, does not impersonate the person). That is a statement about intent, not a grant of rights, and it does not tell a downstream user what they may do with the transcripts, the persona files or the graph data. Anyone planning to redistribute or build on this should resolve that question directly rather than infer it from the README's framing.

A fourth, smaller issue: the README names model identifiers (`claude-fable-5`, and mentions of Fable 5 and GPT6 in the repository description) that a reader cannot verify from the repository contents. The build pipeline that produced the distillation is described in prose, not shipped as a runnable script. You can inspect the outputs; you cannot rerun the process.

Against a plain persona prompt, and against a retrieval setup

The README anticipates the obvious comparison and answers it directly: a persona prompt tells the model "you are person X, answer in their style," which the README calls cosplay. The stated differences are three. First, every answer can be traced to a lecture and a timestamp. Second, the concept relationships come from the corpus rather than from the model's own associations. Third, the teaching structure (intuition, black box, open the box, trap, recap) was induced from transcripts rather than written by hand.

That is a fair description of the gap, but it is not the only alternative. A retrieval-augmented setup over the same transcripts would give you the traceability without the style layer, and it would not require the graph or the interview. The trade-off runs the other way too: a retrieval system answers with what the lecturer said, while this skill answers with how he would approach a question he never covered. The interview calibration is what makes the second claim defensible, and it is also what makes the project hard to fork, since you cannot redo the interview.

A third option is simply reading the source lectures. The skill's own five-part structure is a compressed version of what the lectures already do. For someone who watches the course, the skill's value is speed and coverage, not new information.

Maintenance cost and what the repository asks you to keep current

The last push to the default branch was on 2026-08-11, which is recent enough that the repository is not stale, but the README describes an ongoing expansion rather than a finished artefact. The transcript count is stated as growing, and `AGENTS.md` is described in the repository layout as a "Wiki maintenance schema," which implies the wiki pages are meant to be regenerated rather than hand-edited. If you fork this, you inherit that maintenance loop: fetching captions through `youtube-transcript-api`, rebuilding the graph with `networkx` and `python-louvain`, and keeping the topic and series pages consistent with the schema in `AGENTS.md`. The README does not document a migration path for the graph format, so a change to `graph.json`'s structure would be a manual fix.

The upgrade surface is otherwise small. There are no releases, no version tags and no changelog in the repository, so adopting this means tracking the `main` branch. Pin a commit if you care about reproducibility, because there is nothing else to pin to.

Editorial conclusion

Adopt it if you already run a coding assistant that reads a SKILL.md and you want machine learning explanations in Traditional Chinese that follow a lecturer's structure rather than a generic persona prompt. Skip it if you need English output, if you need coverage beyond the 27 cached transcripts, or if you cannot read the SKILL.md and references/ files to judge the calibration claims for yourself. Verify first that the repository's stated licence is acceptable for your use, since the README does not name one, and check whether the concept you care about appears in graph.json before trusting a graph query to return it.

Frequently asked questions

Who is Hung-yi Lee, and what does this skill have to do with him?

The repository does not give a biography, but it is built from his course material: 478 YouTube video metadata records, 27 transcripts and 8 topic pages. The README states the skill imitates his teaching method and does not impersonate him, and that the description of his style was authorised.

How do I install hung-yi-lee-skill?

Clone the repository, install the Python dependencies from requirements.txt, and place the directory where your AI coding assistant can read it. The README says pointing the assistant at SKILL.md is what starts the skill.

Does hung-yi-lee-skill work in English?

The README is in Traditional Chinese with an English translation, and the skill's examples, style markers and teaching vocabulary are Chinese. The corpus the style was mined from is Chinese lecture transcripts, so English output is outside what the repository supports.

What is the knowledge graph inside hung-yi-lee-skill?

It is a graph of 916 nodes and 3,664 edges across 10 communities, extracted from the course corpus. The README splits the edges into 1,621 EXTRACTED and 2,043 INFERRED, and there is a CLI at scripts/hungyi_kb.py for querying it.

What licence does hung-yi-lee-skill use?

The repository metadata shows no licence and the README does not name one. The README only states that the lecturer authorised the description of his teaching style, which is a statement about intent rather than a licence grant.

Official sources

  1. Issues
  2. README
  3. voidful/hung-yi-lee-skill on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/voidful-hung-yi-lee-skill.svg)](https://hysenlabs.com/projects/voidful-hung-yi-lee-skill)