Model or dataset
UditAkhourii/neuroarxiv avatar
UditAkhourii/neuroarxiv

NeuroArxiv: arXiv Prior Art for Coding Agents

A skill to kill from-scratch coding — Claude checks real arXiv prior art before it designs a new architecture.

433 stars45 forksTypeScriptMIT

At a glance

What is it?
NeuroArxiv is a TypeScript skill for Claude Code and Codex CLI that fetches real arXiv papers before a model designs a new architecture, then commits to one cited recommendation rather than handing back a list.
Who is it for?
Engineers working on non-trivial architecture choices in Claude Code or Codex CLI should install NeuroArxiv before the design phase and run it with the specific problem statement. It is the wrong choice for a quick script, a debugging question, or any problem where arXiv has no relevant literature.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 14 days ago.
What is it written in?
Mainly TypeScript, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What NeuroArxiv Does Before the First Design Decision

Coding agents will happily design a caching layer, a leader election protocol, or a consensus algorithm from scratch if you ask them to. That output often looks coherent. The problem is that decades of research on those same problems exist on arXiv, and the agent is not consulting any of it.

NeuroArxiv addresses that gap. It is a skill for Claude Code and Codex CLI that runs before design work, not instead of it. When you invoke it with a problem statement, it searches arXiv by category, reads each result in isolation, and then commits to one recommended prior-art path. The intended audience is engineers working on non-trivial architecture choices: caching strategies, distributed coordination, inference scheduling, statistical methods. Anything where a published mistake in a 2019 paper is directly relevant to what they are about to build.

The project description puts this plainly: before Claude designs something new, it checks arXiv first. The skill is not a search wrapper. Search hands back a list of sources. NeuroArxiv hands back a decision grounded in them.

Isolate, Score, Converge: The Five-Step Mechanism

The README documents five numbered steps. The first maps the problem onto three to five arXiv categories plus search terms. The second issues real HTTP requests to export.arxiv.org, one request per category, with no language model involved. The rate-limiter is described as courtesy-rate-limited and deterministic.

Step three is the divergent half. Each paper is read in a separate, isolated language model call. The prompt for that call sees exactly one abstract, never the others. This isolation is the design choice that makes the convergent step meaningful: because no single read can anchor on what another found, the scoring step that follows operates on independent opinions rather than a consensus that formed too early.

Step four scores each paper on relevance, practicality, and rigor. Papers are then clustered by architectural angle. Step five, convergence, picks one cluster as the recommended path, names why the others were not selected, synthesizes a first step, and flags known failure modes in the prior art.

The README describes source-skepticism behavior as the distinguishing output. An internal evaluation showed NeuroArxiv flagged a withdrawn proof it had cited, and caught a benchmark validated at only one context length. Neither a cold LLM call nor an undisciplined web-plus-arXiv agent produced any such flags across five cross-domain problems. The README gives the headline result as zero flags versus zero flags versus seven, out of five problems each. The full methodology, per-problem transcripts, and scoring are in EVALS.md and bench/deep-tech-eval-transcripts.md.

Convergence is the explicit design departure from open-ended literature search tools. The README names this: NeuroArxiv does not hand back four papers and ask you to decide. It commits to one recommendation, states why the runner-ups lost, and names what to watch for even in the paths not taken.

Installing and Running the Skill

The README describes a single-command install that requires no cloning:

bash
npx github:UditAkhourii/neuroarxiv install

The installer detects which agents are present and places the skill in the appropriate directory. Claude Code receives it at ~/.claude/skills/neuroarxiv; Codex CLI receives it at ~/.codex/skills/neuroarxiv. Both can be targeted with --all, or a specific agent forced with --claude or --codex. The CLAUDE_CONFIG_DIR and CODEX_HOME environment variables are respected for non-default locations.

After restarting the agent, NeuroArxiv runs in Claude Code as /neuroarxiv followed by the problem statement. Two quickstart examples from the README:

bash
neuroarxiv "cache LLM completions across requests without serving stale answers"
neuroarxiv "leader election for a queue with flaky nodes" --papers 6

The --papers flag controls how many papers the fetch step retrieves. For a full local checkout:

bash
git clone https://github.com/UditAkhourii/neuroarxiv.git
cd neuroarxiv
npm install
npm run build
node dist/cli.js install

Node 18 or above is required. Runtime dependencies are @anthropic-ai/claude-agent-sdk, p-limit, and zod. The skill at skills/neuroarxiv/SKILL.md runs the same fetch-diverge-converge loop using WebFetch when invoked inside Claude Code without a separate CLI install.

Where arXiv-Only Coverage Falls Short

The README reports this directly. In the internal evaluation, NeuroArxiv lost to the undisciplined web-plus-arXiv condition on citation breadth in two of five problems. arXiv covers physics, mathematics, computer science, quantitative biology, statistics, and related fields. It does not index software engineering blog posts, vendor documentation, or proprietary research. For problems where practical prior art lives in conference proceedings not submitted to arXiv, or in grey literature, the search will miss it.

A second failure mode is hallucinated citations. The skill's anti-patterns section names this explicitly. Every paper ID and link in the output is real arXiv metadata returned by the actual API, but the claim made about a paper is still a language model synthesis of a short abstract. The recommendation can be based on an abstract that the full paper contradicts or qualifies. The README instructs the read prompt not to quote more than a few words verbatim; that discipline reduces but does not eliminate synthesis errors.

A third constraint is scope. NeuroArxiv targets architecture-level decisions. It is the wrong tool for debugging a specific error, writing a utility function, or choosing between two libraries that are not subjects of academic research.

Compared to a General Research Agent

The comparison constructed in the README is between NeuroArxiv and a capable agent with general web plus arXiv access but no isolation discipline. The README calls this the undisciplined condition and confirms it is a genuinely grounded baseline: a sample of its citations was verified against the real arXiv API before scoring.

The web-plus-arXiv agent wins on citation breadth in some cases because it can reach Google Scholar, vendor documentation, and indexed blog posts that arXiv does not cover. A team that wants a wide survey before a design decision will get more sources from a research agent using general web search.

What NeuroArxiv adds is the commitment step. The README is explicit that not handing back a list of papers to decide from is the deliberate design departure from open-ended research tools. When the team needs a single recommendation with stated reasons and named failure modes, the isolate-then-converge loop is specifically built for that output. When the team needs breadth and will synthesize the sources themselves, a general-purpose research agent is the more appropriate choice.

Maintenance and License

The repository carries a MIT license. The last push was on 2026-09-17. Version 0.1.0 is the current state, and the repository has no GitHub release tags. There is no changelog file in the repository. The CI workflow runs under .github/workflows/ci.yml. The package name in package.json is neuroarxiv-agent, with a bin entry mapping the neuroarxiv command to dist/cli.js.

Editorial conclusion

Engineers working on non-trivial architecture choices in Claude Code or Codex CLI should install NeuroArxiv before the design phase and run it with the specific problem statement. It is the wrong choice for a quick script, a debugging question, or any problem where arXiv has no relevant literature. The citation net is narrower than general web search by design. Before trusting the output for a consequential decision, read the EVALS.md scorecard to understand where the evaluation was grounded and where it was not.

Frequently asked questions

How does NeuroArxiv avoid inventing paper titles or citations?

The README states that every paper ID and link in the output comes from real arXiv metadata returned by the actual export.arxiv.org API. The sourcing step is deterministic HTTP with no language model involved. The skill's anti-patterns section explicitly names hallucinated citations as the failure mode to watch for in the synthesis step that follows.

Does NeuroArxiv work with Codex CLI as well as Claude Code?

Yes. The installer detects both Claude Code and Codex CLI and places the skill in the correct directory for each. You can force both with --all, or target a specific agent with --claude or --codex regardless of what is detected.

What does the --papers flag control?

It sets how many papers the arXiv fetch step retrieves per category. The README shows it used as --papers 6 in a quickstart example for a leader election problem.

Official sources

  1. Issues
  2. License: MIT
  3. Project website
  4. README
  5. UditAkhourii/neuroarxiv on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/uditakhourii-neuroarxiv.svg)](https://hysenlabs.com/projects/uditakhourii-neuroarxiv)