# hung-yi-lee-skill distils one lecturer's teaching voice from 478 videos, and more than half its knowledge graph is inferred rather than extracted

> hung-yi-lee-skill is an agent skill that reproduces a named lecturer's teaching method, built from 27 full transcripts, 478 video metadata records, and an interview the repository says was conducted with him. The graph statistics are the most informative part: 916 nodes and 3,664 edges, of which 2,043 are inferred rather than extracted, and a community table headed as ten rows that lists six.

**voidful/hung-yi-lee-skill** — 蒸餾李宏毅老師的skill，結合Karpathy的LLM，Fable 5和GPT6 加持 以及 本人親自訪談

- Repository: https://github.com/voidful/hung-yi-lee-skill
- Stars: 1,292 · Forks: 128
- Language: HTML
- License: not declared
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/voidful-hung-yi-lee-skill

## More than half the graph edges are inferred, not extracted

The graph statistics table is four lines and the fourth is the one that matters. There are 916 nodes, 3,664 edges, and 10 communities. Of the edges, 1,621 are marked EXTRACTED and 2,043 are marked INFERRED. That is 56 percent inferred, which means the majority of the structure connecting those concepts was not read out of the transcripts at all and instead came out of whatever process builds the graph. The repository is explicit elsewhere that relationships between concepts are extracted from the course corpus rather than guessed by a language model, so the two claims sit in tension unless the extracted-versus-inferred split is the distinction being made, which the table never spells out. Two smaller mismatches sit in the same area. The community table is headed as ten communities and lists six, with node counts of 396, 116, 81, 79, 72, and 33, which sums to 777 and leaves 139 of the 916 nodes unassigned. The god-node table also lists six rows, but those are connectivity degrees rather than node counts, so ML Fundamentals appears as 396 nodes in one table and 385 connections in the other.

## The interview outranks everything the transcripts suggested

This is the part of the repository that is unusual and the part most worth reading. A multi-agent workflow generated an interview script from the skill file, the transcripts, and the golden and negative examples, containing quote-reaction prompts, reverse-generation prompts, and ten A/B comparison questions rooted in real corpus. The repository then says it interviewed the person, and placed his answers in the highest-authority layer of the skill, a first-person calibration block, with an explicit rule: anything inferred from transcripts that conflicts with what he said loses to what he said. The correction table lists four such conflicts. The skill had assumed an opening built on a hot, freshly-written document was vivid; he does not use that word. It had assumed frightening things were best made concrete with everyday analogies; he considers those too plain and prefers stranger anime metaphors. It had assumed converting a number into one working day was concrete enough; he says it has to convey how expensive and how scarce the thing is. And the three rules the skill had guessed at were replaced by three he stated himself: content needs context rather than a running account, there must be a punchline, and teaching a method should lead the student to ask how the method was arrived at.

## Twenty-seven transcripts out of four hundred and seventy-eight videos

The corpus is described in two places and the numbers only partly line up. The header claims 478 YouTube videos, 27 complete transcripts, 8 topic pages, and 4 curated research references. The sources section adds a fifth figure the header does not mention, 203 series pages. So the material behind the voice is 27 transcripts out of 478 videos, roughly one video in eighteen, with the transcripts themselves marked as still expanding. The curated references are four Markdown files with predictable roles: a persona file defining the teaching persona and voice, a spirit file for the deeper philosophy, a work file for the technical scope, and a sources file listing where everything came from. That set is what the voice claims to rest on beyond the transcripts themselves, and it is four documents rather than a bibliography of research.

## The description names two ingredients the README never mentions

The repository description credits a language model by Karpathy, Fable 5, GPT6, and a personal interview. The README body names only Fable 5, and it names it twice, written as `claude-fable-5`, once for frequency exploration over the 27 transcripts and once for generating the interview script. Karpathy and GPT6 appear nowhere in the document, and no file in the tree is attributed to them. The frequency work is the one part with hard numbers attached: a word counted 609 times against 14 for its plainer synonym, another 518 times, another 160, and a fourth 36, plus one catchphrase. The project is explicit that these were counted rather than added from impression, which is the right way to build a voice profile from a corpus. The missing attribution is the part you cannot check, since it describes how the artifact was produced rather than what is in it.

## Three dependencies are bounded and yt-dlp is not

requirements.txt has four entries and only three of them have a version range. `youtube-transcript-api` is held between 1.2 and 2, `networkx` needs 3.4 or newer, and `python-louvain` needs 0.16 or newer. The fourth, `yt-dlp`, has no bound at all. That is the package most likely to break underneath you, since it tracks changes in video hosts, so a fresh install months from now can resolve to a different behaviour than the one the transcripts were fetched with. python-louvain is what produces the community list, which makes the ten-versus-six discrepancy in the graph section a question about that library's output as much as about the extraction step. Installation itself is three commands:

```bash
git clone https://github.com/voidful/hung-yi-lee-skill.git
cd hung-yi-lee-skill
pip install -r requirements.txt
```

Then activation is manual: the directory has to sit somewhere your coding assistant will read, and the assistant reads `SKILL.md` to start.

## One CLI, eight subcommands, and a structure listing that stops mid-tree

Everything is driven through `scripts/hungyi_kb.py`, with `scripts/hungyi_graph.py` behind it as the graph engine. The subcommands are sync-metadata for the channel records, sync-transcripts with a limit and a title filter, compile to build the wiki, search with a limit, graph build, graph query, graph report, and lint as a health check. A second entry point opens the interactive visualisation, `wiki/graph/graph.html`. That last file is also the reason GitHub reports the repository's primary language as HTML, in a project whose entry point is a Markdown skill file and whose tooling is Python. The structure listing in the README stops partway through the `wiki/` directory, after three files, so the graph directory that the visualisation command points at is never described there. The root also carries `agents/`, `outputs/`, and `assets/`, which the listing does not reach. Two READMEs are maintained, the Traditional Chinese original and an English one.

## The skeleton is five steps and the value is skepticism

What the skill actually distils is a shape. Every answer is meant to follow five stages derived from the transcripts: intuition first in one sentence, then the black box with its input, output, and objective, then opening the box to explain the mechanism, then the traps including common misconceptions, limitations, and a debugging view, and finally a short recap. Alongside that sit language markers for anticipating a student's question, teaching outside in, a self-questioning rhythm, and bridging to something covered elsewhere in the course. The four stated principles are the more transferable part. Benchmark skepticism: the number is not the answer, you have to ask what it is measuring. Knowledge honesty: separate fact from inference and say so when unsure. Concrete analogy: turn an abstraction into a picture from daily life. And intuition before mathematics, with the standard being whether an undergraduate can follow it.

## Conclusion

hung-yi-lee-skill suits someone who wants a specific pedagogical shape, intuition before mathematics then mechanism then traps, rather than a general assistant with a tone setting, and the transcript-provenance discipline is the part worth stealing. It is a poor fit if you need reproducible coverage, since 27 transcripts out of 478 videos is the actual corpus, and a poor fit if you need to know what is licensed, because the repository carries no licence file at all. Before using it, read the extracted-versus-inferred edge split and decide whether an inferred graph is what you are paying for, note that `yt-dlp` is the one dependency with no version bound, and treat the interview material as the project's own account rather than something you can verify from the tree.

## FAQ

### What is hung-yi-lee-skill and who is it for?

It is an agent skill that reproduces a named lecturer's teaching method rather than merely a tone, built for coding assistants that read a SKILL.md file. Its pitch is that the artifact captures a way of thinking about machine learning topics rather than a collection of quotations, and its teaching skeleton runs intuition, black box, mechanism, traps, and recap.

### How big is the knowledge graph inside hung-yi-lee-skill?

916 nodes, 3,664 edges, and 10 communities. Of those edges, 1,621 are marked extracted and 2,043 are marked inferred, so more than half the structure is inferred rather than read from the transcripts. The community table lists only six of the ten, with node counts that sum to 777 of the 916.

### What corpus was hung-yi-lee-skill built from?

478 YouTube video metadata records, 27 complete transcripts described as still expanding, 8 topic pages, and 203 series pages. Four curated Markdown references sit alongside: a persona file, a spirit file, a work file, and a sources file. The transcripts are therefore about 27 of the 478 videos.

### How is hung-yi-lee-skill installed and used?

Clone the repository, `cd hung-yi-lee-skill`, and run `pip install -r requirements.txt`, which pulls youtube-transcript-api, networkx, python-louvain, and yt-dlp. Then place the directory somewhere your coding assistant reads and have it load `SKILL.md`. The bundled CLI is driven with `python3 scripts/hungyi_kb.py` using subcommands like sync-metadata, sync-transcripts, compile, search, graph build, graph query, graph report, and lint.

### What license is hung-yi-lee-skill under?

It is not stated. No licence value is recorded in the repository metadata and no LICENSE file appears among the top-level entries, which include the two READMEs, SKILL.md, AGENTS.md, requirements.txt, and the raw, references, wiki, scripts, agents, assets, and outputs directories. The project also describes itself as built from transcripts of a third party's lectures, so the absence is worth resolving before you redistribute anything.

## Sources

- [Issues](https://github.com/voidful/hung-yi-lee-skill/issues)
- [README](https://github.com/voidful/hung-yi-lee-skill/blob/main/README.md)
- [voidful/hung-yi-lee-skill on GitHub](https://github.com/voidful/hung-yi-lee-skill)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/voidful-hung-yi-lee-skill
