# Twelve skills and six hooks inside fcakyon/phd-skills

> A Claude Code plugin whose selling point is interception rather than commands, written by a researcher with 2000+ citations after 300+ sessions of watching models misread research work. Here is what each skill triggers on, what the hooks block, and where the documentation runs out.

**fcakyon/phd-skills** — PhD Research Skills for Claude Code: paper reproduction, experiment design, paper review, result comparison and more.

- Repository: https://github.com/fcakyon/phd-skills
- Stars: 412 · Forks: 36
- Language: Shell
- License: MIT
- Published: 2026-09-15 · Updated: 2026-09-15 · Language: en
- Canonical page: https://hysenlabs.com/projects/fcakyon-phd-skills

## Guardrails that fire when you never invoke them

Most of this plugin is not something you type. The research guardrails run silently, and the README draws the line itself: other plugins hand you more commands, this one hands you guardrails. Nothing in the visible table has a slash name, because nothing in it is a slash command. Four of the entries are shell scripts under plugin/scripts/, and each maps onto one failure mode. remote_inplace_edit_guard.sh catches in-place edits to git-tracked source over SSH, the situation where an agent patches a file on a remote machine and the working tree stops matching the commit you reviewed. outbound_artifact_reminder.sh catches unverified commands or paths in outbound teammate messages, which is what a half-checked command becomes once it travels to a human in chat. jargon_scrub.sh catches project-internal jargon in commits and docs. timezone_scrub.sh catches timezone tokens that do not match the system clock. One entry is not a script at all: plugin/hooks/prompts/stop_research_peer.md reviews conclusions against actual artifacts using a fresh-context research peer, which reads as a stop hook rather than a lint. The last visible row is a pre-flight checklist on long ML training launches, and its script name is cut off mid-word. The framing behind that list is stated in the opening: the author reports 2000+ citations, 5 patents, 300+ Claude Code sessions and thousands of hours of PhD research, and says every guardrail in the plugin traces to a real mistake.

## Twelve skills keyed to phrasing instead of names

Nothing in the skill table asks you to memorize a command. Each row pairs a sentence with the file that handles it, so activation depends on how you phrase the request. Ask why is X failing / diverging / OOMing and Debug runs. Ask compare run A to baseline and Compare runs. Ask launch a new training run or kick off training and Launch runs. The research half works the same way: design an ablation study reaches Experiment Design, check if my numbers match the code reaches Paper Verification, review my methods section for consistency reaches Paper Writing, analyze dataset bias reaches Dataset Curation, prepare code for open-source release reaches Research Publishing, what will reviewers ask about this? reaches Reviewer Defense, setup latex for CVPR reaches LaTeX Setup, find related papers on X reaches Literature Research, and reproduce this arxiv paper reaches Reproduce. That is twelve skills, each a directory under plugin/skills ending in SKILL.md. The set spans one research cycle rather than a single stage: debugging and launching sit in the same table as paper writing, dataset curation and open-source release preparation. Two agents sit a layer above them. Claude delegates to paper-auditor, which cross-checks paper claims against code and data from an isolated worktree and remembers patterns across sessions, and to experiment-analyzer, which reads wandb, neptune, tensorboard, mlflow or local results before handing off to the compare and debug skills. /phd-skills:help is the only place the whole set is shown at once.

## Compare fixes the epoch, debug withholds the verdict

Two of these skills exist because of how the author lost time, and the README is specific about the failures. Compare aligns runs at the same epoch, so the sentence compare run alpha to baseline does not turn into a comparison of whatever numbers happened to land in two log files. Epoch is the whole argument: a run read at step 3000 against a baseline read at step 30000 produces a gap that looks like a result. Debug is built the same way, and the usage section shows it auto-triggering on why is my loss diverging? and running evidence-first probes. The failure that shaped it is precise: an experiment was declared diverging on the strength of a proxy metric that had not converged, and the run was killed before downstream evaluation would have shown the truth. Another item in the same list is that a generated figure was never once looked at, only its numbers trusted, which is the reason a stop hook asks a fresh-context research peer to check conclusions against actual artifacts. Among the visible guardrails, none targets local file deletion, and the lost-work list includes rm -rf run against a path hallucinated from memory that took local checkpoints with it. Whether a hook intercepts that case is not answered by the visible table. The rest of that failure list is equally specific: a 50-hour training job was restarted without diffing the config against the reference run, which cost three days; a full dataset was analyzed when a specific 4k/2k/2k split had been asked for; a test was reported as covering a bug that had never actually been verified; and a typo, typing done? where dont? belonged, launched an unwanted upload of thousands of files.

## /phd-skills:reproduce takes one arxiv identifier and keeps going

The reproduce command carries the longest chain of work in the plugin, and the usage section describes it in a single line: /phd-skills:reproduce arxiv 2508.12345 reproduces a paper from an arxiv URL through replication runs. The identifier in that example is an arxiv style id, and what follows it is the part that matters for planning, because the command is described as continuing past retrieval into the replication runs themselves. Three neighbouring pieces cover the checking a reproduction needs. The xray command audits a paper against code and data across five parallel dimensions. The factcheck command verifies BibTeX entries and cited claims against DBLP. The Paper Verification skill, which activates on check if my numbers match the code, covers the narrower case where the claim under test is a number rather than a reference. Two agents can be handed the same job: paper-auditor cross-checks claims against code and data inside an isolated worktree, and experiment-analyzer works from tracker output instead of a paper. The difference shows up in the file layout rather than in prose. Commands live under plugin/commands as single markdown files such as plugin/commands/xray.md, while skills live under plugin/skills as directories containing a SKILL.md, and agents live under plugin/agents.

## fortify takes a venue, gaps takes a topic

Two commands take arguments, and they are written the way a shell signature would be. /phd-skills:fortify [venue] selects the strongest ablations and anticipates reviewer questions, so a venue name is optional and the analysis is not. /phd-skills:gaps <topic> has the opposite shape: the topic is required, and the command adds web confirmation to the gap analysis it produces. The reviewer-facing skill sits next to fortify and answers what will reviewers ask about this? by name, while literature-research answers find related papers on X. Grouped this way, these are the paper-cycle commands rather than the experiment-cycle ones. fortify and Reviewer Defense both work on anticipating criticism, factcheck works on citations, and gaps works on whether a contribution is already covered somewhere. What is not written down is how far each one reaches. The brackets around venue are the only hint that a recognized set exists, and there is no venue list, no threshold for what counts as a strong ablation, and no account of how the web confirmation in gaps is performed or recorded. LaTeX Setup is the odd one out in this group, since setup latex for CVPR is a formatting task rather than an analysis.

## Two install lines and an optional 30-second tour

Installation is two lines, and both concern the plugin registry rather than a package manager:
```
claude plugin marketplace add fcakyon/phd-skills
claude plugin install phd-skills@phd-skills
```
The repository root holds .claude-plugin/, LICENSE, README.md and plugin/, and the primary language is Shell, which matches the guardrail scripts under plugin/scripts/. No dependency step, environment variable, or configuration file appears anywhere on the setup path. One prerequisite sits before both lines: Claude Code has to be open in the project directory you want covered, since the commands and skills act on that session. The next step is optional: the README says the plugin works correctly the moment it is installed, and that /phd-skills:setup gives a 30-second tour of what was auto-detected while letting you opt into extras. Those extras are notifications, an allowlist, and LaTeX. Notifications cover task completion and background agents and forward to ntfy, Slack, or email once setup has run, so skipping the tour is also how you stay without them. The recurring-check pattern in the usage section leans on a /loop invocation that reads experiment logs on a 30 minute interval and notifies when metrics beat the baseline or when loss starts to diverge, which is the alerting face of the same pre-flight check that guards a long launch.

## The guardrail table ends mid-row, and setup has no documented knobs

Two properties of the documentation matter before you depend on it. First, it is cut off. The research guardrail table stops mid-row, on a pre-flight checklist for long ML training launches whose script name is truncated after long_job_la, so the visible list is a floor rather than a complete inventory, and whatever followed it in the source is not visible here. Second, the installation story is complete while the configuration story is absent. No file states which hooks are enabled by default, how the allowlist offered during setup is populated, or how to switch an individual guardrail off. What can be checked from the repository itself is licensing and cadence. A single MIT LICENSE file sits at the root and covers the repository, with no competing license to interpret and no per-artifact split to untangle. Releases are tagged v1.0.0 on 2026-03-11, v1.1.0 on 2026-03-14, and v1.3.0 on 2026-05-09, where v1.3.0 carries the line about catching AI mistakes before they cost weeks of compute, and no v1.2.0 appears in the tag list. The default branch is main, the last push landed on 2026-09-16, and the project is not archived. No homepage is set, so the author link in the README is the only external pointer given.

## Conclusion

This fits a researcher who already runs Claude Code against wandb or tensorboard and wants the checklist enforced instead of remembered, and it fits nobody who wants a paper summarizer or a writing assistant out of the box. Before installing, read the hook scripts under plugin/scripts/ yourself: the visible guardrail list is truncated mid-row, no configuration knobs are documented, and the entire value here is in what gets blocked before the compute is spent.

## FAQ

### Does fcakyon/phd-skills install anything besides the Claude Code plugin?

Installation is two commands, claude plugin marketplace add fcakyon/phd-skills and claude plugin install phd-skills@phd-skills. The README documents no dependency step, environment variable, or configuration file.

### What does the compare skill in fcakyon/phd-skills actually align?

It aligns runs at the same epoch, and the experiment-analyzer agent reads results from wandb, neptune, tensorboard, mlflow, or local output before handing off to the compare and debug skills.

### How do notifications work in fcakyon/phd-skills?

Task completion and background agent notifications forward to ntfy, Slack, or email, and you opt in during /phd-skills:setup, where the extras also include an allowlist and LaTeX.

### What stops a bad training launch in fcakyon/phd-skills?

The launch skill runs a pre-flight checklist when you ask it to start a training run, and a separate hook applies a pre-flight checklist on long ML training launches.

### Which external services do fcakyon/phd-skills commands call?

factcheck verifies BibTeX entries and cited claims against DBLP, and the gaps command adds web confirmation to the literature gap analysis it produces.

## Sources

- [fcakyon/phd-skills on GitHub](https://github.com/fcakyon/phd-skills)
- [Issues](https://github.com/fcakyon/phd-skills/issues)
- [License: MIT](https://github.com/fcakyon/phd-skills/blob/main/LICENSE)
- [README](https://github.com/fcakyon/phd-skills/blob/main/README.md)
- [Releases](https://github.com/fcakyon/phd-skills/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/fcakyon-phd-skills
