# The ELI5 skill ships a prompt, a Python grader, and one undated pass rate

> ELI5 is a Claude Code skill that reshapes an explanation for the audience named in the prompt, plus the Python harness that scores it. The harness is where the repository's real machinery sits, and the single number it publishes rests on twelve assertions.

**DreambigOu/ELI5** — ELI5 — A Claude Code skill that explains anything to anyone: kids, managers, engineers, parents. Adapts tone, vocabulary, and analogies to match the audience.

- Repository: https://github.com/DreambigOu/ELI5
- Website: https://andrewou.pages.dev/posts/building-an-eli5-skill-for-claude/
- Stars: 1,614 · Forks: 91
- Language: Python
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/dreambigou-eli5

## The skill is prompt text and the only Python in the tree is the grader

The repository's top level holds five entries: `.gitignore`, `LICENSE`, `README.md`, `eli5-workspace/` and `skills/`. Python is recorded as its primary language, and the Python is `eli5-workspace/run-evals.py`, the script that scores answers. What a user installs is the `skills/` directory, and everything that directory is described to do arrives as prose: the skill detects the target audience from your prompt and calibrates vocabulary, analogies, tone, depth and framing. No file in the tree parses an audience out of anything. There is no classifier, no keyword table, no regular expression, and no config file that maps an age or a job title onto a tone. The five axes are therefore instructions the model is asked to follow, and the only instrument in the repository that can report whether it worked is another model call. Measuring the skill and being the skill are the same operation, which is worth holding on to for every number that follows.

## One of the five documented invocations names no audience

The usage section is five lines in total:
```
ELI5 what a database index is
Explain this code to my manager
Break down how git merge conflicts work for a 5th grader
Explain this error to my mom
Simplify this for a designer
```
Four of the five name the listener. The first one, a bare topic request, does not, and the documentation never says what happens in that case: no default age, no default register, nothing about whether the skill asks a follow-up question or settles on a middle voice and answers anyway. Now compare that with what the test suite feeds it. Every case in `evals.json` carries an explicit `audience` field, the worked example sets `"audience": "Age 15"`, and the first test in the sample run is named `explain-db-index-age5`. The graded prompts all state their audience, so the default case that the first documented invocation stands for is the one case the suite does not grade. Anyone who copies that first line takes the least tested path through the skill.

## The audience columns overlap and depth decides how long the answer gets

The supported-audience table has four columns that are not independent: Ages, Grade Levels, Job Roles and Relationships. A fifth grader is age 10 or 11, so the first two columns describe the same listener with no rule for reconciling them. A manager in their forties is also 40+ and also a parent, and the table gives no way to say which of those wins when the axes disagree. Depth is the axis that fixes length: short and sweet for simple audiences, nuanced for grad students. That reads sensibly until the assertions are placed next to it. For the first test the with-skill arm is graded on jargon, on picking a book and page analogy over a phone book, on short conversational sentences, and on a warm tone containing particular slang. Three of four criteria are about manner. In the recursion case only two of the four assertions check whether the answer is right, namely that a function calling itself is explained and that a base case is mentioned; the other two ask which references get used as analogies and whether the voice avoids what one assertion calls cringey. The headline rate pools both kinds of criterion, so it is not a correctness rate, and the depth axis means a short answer can earn the same marks for a prompt where a longer one would be graded differently.

## Installing takes two shell lines and the second one assumes two directories exist

Installation is a clone followed by a copy:
```bash
git clone https://github.com/DreambigOu/ELI5.git
cp -r ELI5/skills/eli5 ~/.claude/skills/eli5
```
The second line assumes two things the documentation never creates: a clone directory named `ELI5`, which is what GitHub hands you for this repository name so long as nobody clones somewhere else, and an existing `~/.claude/skills/`. No `mkdir -p` appears anywhere. There is also no second way in. The top level carries no `package.json`, no `requirements.txt`, no `pyproject.toml` and no plugin manifest, so there is no packaged form of this skill and nothing to install except the copy. The evaluation commands are relative paths that have to run from the clone itself:
```bash
python eli5-workspace/run-evals.py
python eli5-workspace/run-evals.py --test=1
python eli5-workspace/run-evals.py --with-skill-only
python eli5-workspace/run-evals.py --grade-only
```
The stated prerequisites are the Claude Code CLI and the skill sitting at `~/.claude/skills/eli5/`, so the workspace has to stay on disk after you copy the skill out of it, and the two halves of the setup are coupled in a way neither section spells out. One dependency is still unnamed: the grader calls Claude, and nothing here names a client library, an environment variable or a key that pays for those calls.

## The grader is a Claude call and one assertion asks it to have taste

The script's own summary says it runs each prompt twice, once with the skill and once without, auto-grades every output against its assertions using Claude, and prints a pass rate summary comparing the arms. Both arms are written by the same kind of model and both are then judged by that same kind of model, which is the cheapest arrangement to run and the hardest one to lean on, since nothing in the loop is independent of the thing being measured. The assertions are not all checkable properties either. One asks for a casual tone that is not cringey, and rules out the register it calls fellow kids energy. Another looks for the literal words huuuge and super duper fast in an answer meant for a five-year-old. A grader that is itself a language model has to arbitrate both, with no rule saying when enough slang is enough. Each case then leaves a `grading.txt` holding what the documentation calls detailed evidence, which is the grader's own reasoning written back out. The printed sample also looks edited rather than captured, because the same first criterion appears as no technical jargon present in one arm and as no technical jargon found in the other:
```
--- Test 1: explain-db-index-age5 ---
  [with skill]
    PASS  #1 — No technical jargon present
  [baseline]
    PASS  #1 — No technical jargon found
```
A captured run would not reword its own criteria between arms.

## 83.3% against 41.6% works out to twelve assertions

The results table gives 83.3% with the skill, 41.6% without, and a delta of +41.7%. Those figures are not averages of anything rounded: they are ten of twelve and five of twelve, and the delta column is the plain subtraction of the two columns, so it reads as points rather than as the near doubling of the baseline that a percentage change would be. At four assertions per test case, twelve graded assertions means three prompts, which matches the identifiers on show, `explain-db-index-age5` for the first run and `explain-recursion-teenager` carrying id 3. At that size a single prompt moving from zero passes to four passes swings the headline by 33 points. The manager result quoted underneath, 0% at baseline and 50% with the skill, describes a category with at most a couple of graded assertions behind it. Output accumulates in `eli5-workspace/iteration-N/` with N incrementing on every run, and the table defers to a checked-in `eli5-workspace/eval-results.md`. Neither file reference carries an iteration number, a date, a grader prompt or a model name. The default branch's most recent commit is dated March 18, 2026, and the repository publishes no GitHub releases, so the one figure a reader would quote has no commit attached to it.

## `--test=1` picks a position while `id: 3` picks an identity

The instructions for adding a case say to add an entry to the `evals` array in `eli5-workspace/evals.json`, and the worked example carries `"id": 3`. The documented way to run one case is `python eli5-workspace/run-evals.py --test=1`, and the sample header prints an ordinal next to a name. Nothing says whether `--test` matches the `id` field or the array position. The two agree only while ids are handed out in order and new cases land at the end. Insert one at the top of the array and every numeric flag in a shell history re-points at a different prompt without error, because grading still runs and still prints a summary. The `name` field is described as a directory-friendly identifier for storing results, so each name is also a path inside an iteration folder, and renaming a case after the fact leaves the older folder exactly where it was. `--grade-only` closes the loop by rescoring stored outputs without re-running anything, into an auto-incremented folder, so a rescore and a fresh run end up sitting side by side with nothing in the summary to say which pass rate came from which.

## Conclusion

Use this skill when the audience is already named in the prompt and you only want the register changed. Skip it when the audience has to be inferred, because that is the case its own suite never grades. If you plan to quote the 83.3% figure, open `eli5-workspace/eval-results.md` and match it to an iteration folder first, because the results table records no date, no grader and no commit. The default branch's last commit is dated March 18, 2026, and no releases sit behind it, so read the harness as a snapshot you can inspect rather than a scorer you can take on faith.

## FAQ

### What does the ELI5 Claude Code skill change in an answer?

It asks the model to adjust five things for the audience named in the prompt: vocabulary, analogies, tone, depth and framing. The skill is installed by copying `skills/eli5` into `~/.claude/skills/eli5`, and nothing in the repository performs the audience detection itself.

### How does the ELI5 skill decide who the audience is?

From the wording of the prompt. The supported list covers ages from 5 to 40+, grade levels, job roles and relationships, and the cases in `evals.json` pass the audience explicitly, for example `Age 15`. The documentation does not say what happens when no audience is named.

### What pass rate does ELI5 report for answers written without the skill?

41.6%, against 83.3% with the skill. Those two numbers work out to five of twelve assertions and ten of twelve, and the setup asks for four assertions per test case, which puts the suite at three prompts.

### How do you run the ELI5 evaluations?

Run `python eli5-workspace/run-evals.py` from the repository root with the skill installed at `~/.claude/skills/eli5/`. The flags include `--test=1` for a single case, `--with-skill-only` to drop the baseline arm, and `--grade-only` to rescore stored outputs.

### When was the ELI5 repository last changed?

The most recent commit on the default branch is dated March 18, 2026, and the repository publishes no GitHub releases. Evaluation output accumulates under `eli5-workspace/iteration-N/` with N incrementing on each run.

## Sources

- [DreambigOu/ELI5 on GitHub](https://github.com/DreambigOu/ELI5)
- [Issues](https://github.com/DreambigOu/ELI5/issues)
- [License: MIT](https://github.com/DreambigOu/ELI5/blob/main/LICENSE)
- [Project website](https://andrewou.pages.dev/posts/building-an-eli5-skill-for-claude/)
- [README](https://github.com/DreambigOu/ELI5/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/dreambigou-eli5
