# Auto-Empirical-Research-Skills: What the AERS Skill Library Actually Ships

> AERS bundles agent skills for empirical social science work, from data cleaning to journal submission, with a numeric benchmark you can run against your own agent. The catalog is broad and the rigor tooling is real, but the licence and the benchmark scope both need checking before adoption.

**brycewang-stanford/Auto-Empirical-Research-Skills** — 🔬 A curated collection of 23,000+ agent skills for empirical research across 8 social science disciplines. | 精选 23,000+ AI Agent 技能库，覆盖8大社会科学学科的实证研究。CoPaper.AI 20分钟完成一篇可复现的规范实证论文，并支持用户上传 Skills。-- Maintained by CoPaper.AI from Stanford REAP.

- Repository: https://github.com/brycewang-stanford/Auto-Empirical-Research-Skills
- Website: https://copaper.ai
- Stars: 4,457 · Forks: 533
- Language: Stata
- License: NOASSERTION
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/brycewang-stanford-auto-empirical-research-skills

## The gap AERS is aimed at: research agents that improvise

An agent asked to run an event study will produce something. Whether that something matches Callaway and Sant'Anna, whether the standard errors are clustered the way the field expects, and whether the resulting table would survive a referee are separate questions. AERS exists to answer them by handing the agent a named procedure instead of a blank page.

The repository describes itself as a curated collection covering eight social science disciplines, curated by CoPaper.AI from Stanford REAP. The audience is narrow on purpose: empirical researchers in economics, political science, sociology, psychology, education, public administration, international relations and communication, working with an agent that can load skill folders. If you write theory papers or do qualitative fieldwork, the pipeline stages below will not map onto your process.

## How the skill folders and the nine-stage pipeline fit together

The unit of distribution is a folder containing a SKILL.md file. The README states that a copied folder must carry its own SKILL.md, and warns that for some collections the file sits one level down, so you copy that level instead. The agent reads the description field and selects a skill on its own; naming a method explicitly is the fallback when automatic selection misses.

Above the folders sits a nine-stage pipeline: topic refinement, literature review, data acquisition, identification strategy, estimation, robustness auditing, publication-grade tables and figures, writing and peer review, then AIGC reduction and submission. The README claims each stage is covered by specific skills and that a human can intervene at any step, edit the method or add robustness checks, then let the pipeline resume.

The trust surface is the part worth reading closely. There are 19 numeric benchmark tasks whose gold values are recomputed from real data on each run, plus 42 behavioural eval scenarios carrying 217 rubric items. Of those scenarios, 9 have pass/fail fixture pairs that the README says demonstrate the harness can tell a correct answer from a wrong one, and all 6 critical scenarios are in that set. That is a small number stated plainly, which is more useful than a large number stated vaguely.

## Installing AERS and running one real task

The fastest path the README gives is to hand the repository URL to Claude Code or Codex and state the scope. The scope word decides where files land: directory scope writes to .claude/skills/ under the current working directory, project scope writes to the project root so the folder can be committed, and global scope writes to ~/.claude/skills/ or ~/.codex/skills/.

```text
帮我安装 https://github.com/brycewang-stanford/Auto-Empirical-Research-Skills
装到「全局」（~/.claude/skills/），我想在所有项目里都能用
```

For a managed install, the README documents a plugin marketplace route on Claude Code v2.1 and later. Each plugin maps to a pipeline, so you install only the stack you use rather than the whole catalog.

```bash
claude plugin marketplace add brycewang-stanford/Auto-Empirical-Research-Skills
claude plugin install aer-skills@auto-empirical-research-skills
claude plugin install empirical-analysis-stata@auto-empirical-research-skills
```

If you want a single skill, clone with submodules and copy the folder. The recursive flag matters because the collections are submodules, and the README notes that the folder you copy must contain SKILL.md at its top level.

```bash
git clone --recurse-submodules https://github.com/brycewang-stanford/Auto-Empirical-Research-Skills.git
cd Auto-Empirical-Research-Skills
cp -R skills/00.1-Full-empirical-analysis-skill_Python  .claude/skills/
```

To use it, open a new session and describe the task in plain language. The README's own example asks for a panel event study with Callaway and Sant'Anna estimation, HonestDiD robustness and journal-grade tables. What you should see is the agent selecting a matching skill from the catalog rather than writing the estimator from scratch. If it does not, name the method and the skill directly.

There is a separate scoring tool. Installing the package from a checkout gives you the aers-score console script, which runs an agent against the numeric benchmark.

```bash
pip install -e .
aers-score tasks
```

The pyproject.toml is explicit that this package contains only the aers_score console script. It does not package the skills catalog, the eval harness or the benchmark datasets, and it declares no dependencies on purpose so it installs in the minimal environment an agent author already has. The exam is resolved from a checkout at runtime.

## Where the repository is thin, and where it is honest about it

The benchmark's numeric coverage is Python-centric. The pinned scientific stack in requirements.txt is numpy, pandas, matplotlib, statsmodels, linearmodels and python-pptx, and the comments tie the linearmodels floor to the NumPy 2 ABI and the python-pptx entry to a defence-deck checker added alongside a submodule advance. Stata is the repository's primary language by GitHub's classification, but nothing in the repository shows a Stata grader in benchmark/. If you work primarily in Stata or R, verify what the benchmark actually grades before treating a score as meaningful for your stack.

The licence is the sharper problem. The README carries a CC BY-SA 4.0 badge and pyproject.toml declares CC-BY-SA-4.0, while the repository's licence field reads NOASSERTION. CC BY-SA is a share-alike content licence, not the permissive software licence many engineering teams assume when they pull a skills folder into a private repository. Whether the share-alike obligation attaches to your derived skill files is a question for your own counsel; the point here is that the repository does not resolve it for you, and the two declarations do not agree.

The maintenance picture, from the two facts available: the repository is not archived, and the last push was on 2026-09-07. The only tagged release is v2026.07 from 2026-07-02, described as the first tagged release. There is no published upgrade path or deprecation policy beyond the note that README-zh-CN.md is deprecated and now a redirect placeholder.

## AERS against a general-purpose research skill collection

The obvious alternative is a general research skill pack, the kind assembled for literature search, summarisation and citation management across any field. The difference is verification. A general pack gives the agent retrieval and writing habits; it has no opinion about whether your difference-in-differences specification is identified, and it will not tell you that your robustness table is stale.

AERS couples the skills to a grader. The catalog statistics are checked by a readme-stats validator under make validate, and the benchmark recomputes gold values from real data rather than storing them. That coupling is the whole argument: a skill that produces a wrong number gets caught by the same repository that shipped the skill. The cost is scope. A general pack will help a historian summarise archives; AERS is built around estimators, identification and journal tables, and outside that it has less to offer.

## Frequently asked questions

The FAQ entries below are drawn from the repository's own documentation and from the questions people search for around this project. They cover what the skills are, how they install, and what the benchmark does and does not grade.

## Who should install AERS, and what to check first

The repository is not archived and the last push was on 2026-09-07, so it is current, but the only tagged release is v2026.07 from 2026-07-02 and no upgrade policy is documented. Treat the main branch as the thing you install.

Before you commit a team to it, run make validate in a checkout and read LICENSE rather than the README badge, because the two disagree on the licence. If you plan to score your own agent, install the package from a checkout with pip install -e . and run aers-score tasks to see the exam before you see the score. The contribution guide asks contributors to run make check locally, which covers catalog validation, link checking, unit tests, the eval harness and the benchmark; that same command is the cheapest way to find out whether the repository's own gates pass on your machine.

## Conclusion

Adopt AERS if you run an agent-assisted empirical workflow in economics, political science, sociology or a neighbouring field and you want a catalog plus a benchmark you can point your own agent at. Do not adopt it if you need a permissively licensed drop-in library, if your work sits outside the eight named disciplines, or if you expect the benchmark to cover Stata and R the way it covers Python. Before installing, run the catalog validator and read the license file, because the repository declares CC BY-SA 4.0 in its README badge while GitHub reports the licence as NOASSERTION.

## FAQ

### What skills are needed for an AI researcher?

AERS does not answer this as a list of personal competencies. It ships agent skills organised by research stage, from topic refinement through submission, and distributes them as folders containing a SKILL.md file that the agent selects by description.

### What are the four main types of research skills?

The repository does not group skills into four types. Its own grouping is a nine-stage pipeline, and the catalog is browsable by method, stage, language and licence through the skill search page.

### Can you give me an example of a research skill in Auto-Empirical-Research-Skills?

The README's usage example asks the agent for a panel event study with Callaway and Sant'Anna estimation, HonestDiD robustness and journal-grade tables. That request maps onto skills in the estimation and robustness stages of the pipeline.

### What are the most important research skills in AERS?

The repository does not rank skills by importance. It does mark 6 eval scenarios as critical and states that all 6 are among the 9 scenarios proven to distinguish correct from incorrect answers by pass/fail fixture pairs.

## Sources

- [brycewang-stanford/Auto-Empirical-Research-Skills on GitHub](https://github.com/brycewang-stanford/Auto-Empirical-Research-Skills)
- [Issues](https://github.com/brycewang-stanford/Auto-Empirical-Research-Skills/issues)
- [Project website](https://copaper.ai)
- [README](https://github.com/brycewang-stanford/Auto-Empirical-Research-Skills/blob/main/README.md)
- [Releases](https://github.com/brycewang-stanford/Auto-Empirical-Research-Skills/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/brycewang-stanford-auto-empirical-research-skills
