# scholaraio: four dependencies at the core, six agent configs around it

> ScholarAIO wraps a coding agent in an academic harness with skills, a CLI, a local library WebUI and citation-traceable writing workflows. The interesting split is packaging: keyword search is in the base install, and every semantic, topic and PDF capability is an extra.

**ZimoLiao/scholaraio** — Scholar All-In-One: A research infrastructure for AI agents

- Repository: https://github.com/ZimoLiao/scholaraio
- Website: https://zimoliao.github.io/scholaraio/
- Stars: 576 · Forks: 78
- Language: Python
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/zimoliao-scholaraio

## Four runtime dependencies, and everything semantic is an extra

The base install is deliberately thin. The dependency list is requests, pyyaml, defusedxml and beautifulsoup4, four packages that cover HTTP, configuration, XML parsing and HTML scraping. Nothing that computes a vector, loads a model or reads a PDF is in that list.

All of it is extras. `embed` adds sentence-transformers, numpy and faiss-cpu. `topics` adds bertopic and pandas, and depends on `embed` by naming the package's own extra. `pdf` adds pymupdf. `import` adds endnote-utils and pyzotero for reference managers. `office` adds markitdown with docx, pptx and xlsx support, plus python-docx, python-pptx and openpyxl. `mineru-cloud` adds a MinerU API client. `full` aggregates embed, topics, import, pdf and office and adds modelscope and curl-cffi.

That split has a consequence for the feature table. Hybrid search, semantic retrieval and topic discovery are advertised in the same table as keyword lookup, but the machinery behind the first two lives in `embed` and the third in `topics`. Following the quick start exactly, install and setup, gives you a library that can fetch, parse metadata, index text and search by keyword, and not much more until you add an extra. The self-referential extras are worth knowing about too, since `topics` and `full` both reach back into `embed` rather than listing sentence-transformers themselves.

## Six agents get their own instruction files in one repository

The tree is the clearest statement of what this project is. There is a .claude directory, a .claude-plugin directory, a .cursor directory, a .qwen directory, a .agents directory, a .clinerules file, a .cursorrules file and a .windsurfrules file. On top of that sit AGENTS.md, AGENTS_CN.md, CLAUDE.md, a clawhub.yaml, a hooks directory and a skills directory.

That is instruction and registration material for at least six coding agents, in one repository, plus two documentation languages. The README frames this as the point rather than a distraction: the agent supplies reasoning and orchestration, and ScholarAIO supplies the durable context and the operational contracts around it. The contracts are the skills and the CLI, and the README describes the value as stable ways to search, read, organize, cite, write and verify rather than a better prompt.

```bash
git clone https://github.com/ZimoLiao/scholaraio.git
cd scholaraio
pip install -e .
scholaraio setup
```

The recommended path is exactly those four commands, followed by opening the repository in Codex, Claude Code or another supported agent. In that arrangement the agent gets the fullest experience available: bundled instructions, local skills, the CLI, the repository knowledge map in docs/DESIGN.md and the whole codebase as context. Plugin-based setups are documented as alternatives rather than equivalents.

The cost is duplication that has to be kept in step. Six configuration surfaces and two copies of the agent instructions, AGENTS.md and AGENTS_CN.md, mean a change to the workflow can land in one place and miss another, and there is no mechanism visible in the tree that keeps them aligned. The plugin route is the alternative path, documented separately for Claude Code plugins and for Codex or OpenClaw skill registration.

## Two parsers, and the flagship one is a cloud extra

The first row of the feature table promises deep structure extraction, converting PDFs into structured Markdown while preserving formulas, figures and layout as much as possible. That promise sits next to a packaging choice that leaves the reader to work out which parser does it.

Two parsing paths exist in the dependencies. The `pdf` extra is pymupdf, a local library that reads and writes PDF pages. The `mineru-cloud` extra is a MinerU API client, which is a network call to somebody else's service. Neither is in the base install, and the visible text does not say which one produces the structured Markdown the first row describes. If your data cannot leave the network, that ambiguity decides your installation before you have read any of the features.

The second row widens the scope beyond papers: journal articles, theses, patents, technical reports, standards and lecture notes, handled as four inbox categories with tailored metadata handling rather than one uniform import path. Four categories for six document kinds means the grouping is by provenance rather than by file type, which is the right axis for metadata quality but adds a decision for the importer.

## config.yaml is committed, config.local.example.yaml is not

Configuration is split into a tracked default and an untracked local example. config.yaml sits in the repository, and config.local.example.yaml sits beside it as the template for machine-specific overrides. Anything you do not want in git belongs in the local file, and the tracked file is the shipped default rather than your personal setup.

The backup feature follows the same instinct. It offers two shapes: legacy data-only rsync plans, or a manifest-validated full-instance backup with one-click restore covering local config, data, workspaces, published outputs and control state. The manifest is the interesting word, because a restore that is validated against a manifest is the difference between recovering an instance and discovering that the restore silently skipped a directory.

Metadata Scrub takes the incremental view of the same problem: review and repair low-quality titles, authors and years for non-standard documents, then mark records as reviewed so future passes skip them. That marker is what stops a cleanup job from re-litigating a title you already fixed by hand, and it is the only state in the system that records a human decision.

## Upgrade paths split at 1.4, with a check command between them

The 2.0 release is described as a product-boundary and compatibility release rather than a data migration. The guidance is specific about who does what. Users on 1.4 or 1.5 keep their current data layout: upgrade the package, run `scholaraio setup check`, and rebuild indexes when appropriate. Users on 1.3 or earlier still have to complete the explicit runtime migration described in docs/getting-started/upgrading-to-2.0.md.

So there are three upgrade cohorts rather than one, and a setup check that distinguishes them. The visible text promises the 2.x compatibility promise is documented in that same file, alongside the removed surfaces, which is the part to read before upgrading rather than after something fails.

The version record itself is tidy. pyproject declares 2.0.0, the newest tag is v2.0.0 published 2026-08-17, and the beta before it was v2.0.0-beta.1 on 2026-07-21, with v1.5.0 on 2026-05-24 underneath that. The last push to main is 2026-09-25, so the branch is running ahead of the tag.

## Retrieval is layered, and one row admits it can be slow

Reading is deliberately staged. Layered Reading starts at metadata or the abstract and moves into conclusions or full text only when you ask, so browsing a library does not mean parsing every PDF in it. Hybrid Search combines full-text and vector retrieval, with an optional line-addressable evidence chunk search for the case where you need the exact source snippet rather than the document.

Everything returned is meant to be traceable. The academic writing workflows cover literature review, guided single-paper reading, paper sections, citation check, rebuttal, gap analysis, poster packages and technical reports, and the guarantee attached to all of them is that every citation traces back to your own library. Publisher PDF Fetch is described in the same spirit: it fetches DOI or publisher PDFs through the user's legal network context, with a direct campus-network mode, rather than working around access controls.

Two rows describe restraint instead of capability. Grounded Scientific Tool Use means the agent reads versioned official documentation at runtime rather than guessing scientific-software commands. Bounded Tool Adapters means external tools stay optional, isolated, testable, and subject to a 2.x integration gate. Neither sells a feature; both are statements about what the project declines to automate, and they are the reason the harness claims to be an all-in-one workflow rather than a distribution of every scientific package.

## A local WebUI, a docs site, and a logo that is still a placeholder

The interface is local. The WebUI filters records by field, runs keyword, semantic or unified retrieval, copies canonical BibTeX, shows audit status and Markdown summaries, and opens PDFs inline or in the operating system default viewer. The WSL detail is unusually specific: on WSL it launches a stable Windows edit mirror and persists embedded annotations back to the canonical library PDF, which is a workaround for the one place where a Linux-hosted library and a Windows editor disagree about which file is real.

The rest of the repository is documentation and governance. mkdocs.yml and a docs directory hold the site, including docs/DESIGN.md as a repository knowledge map and docs/getting-started/ for agent setup and the 2.0 upgrade. CONTRIBUTING.md, CODE_OF_CONDUCT.md, SECURITY.md, CHANGELOG.md, CITATION.cff and STRATEGY.md are all present, alongside a gui directory, hooks, scripts and tests.

One thing is unfinished. The README opens with an HTML comment reading TODO replace with actual logo when available, and the image slot under the title is still empty, so the top of the README renders as a heading with a gap above it. It is a cosmetic leftover in a project that otherwise labels itself production stable, and it is the kind of thing that tells you the repository is maintained by a person rather than a release process.

## Conclusion

ScholarAIO suits a researcher who already works inside a coding agent, keeps their own PDF library, and wants citations traceable to documents they hold rather than to a chat log. It does not suit someone who needs semantic retrieval or topic discovery without choosing an extra, or whose PDFs cannot leave the network and who has not decided which parser to use. Before installing, pick the extras deliberately, since the default install is keyword-only, decide between the local pymupdf path and the MinerU cloud extra on privacy grounds, read the 2.0 upgrade note to find out which cohort you are in, and run `scholaraio setup check` rather than assuming your data layout carried over.

## FAQ

### How do I install and set up scholaraio?

Clone the repository, change into it, run pip install -e . and then scholaraio setup. After that you open the repository in Codex, Claude Code or another supported agent, which gives it the bundled instructions, local skills, the CLI and the knowledge map in docs/DESIGN.md.

### Which extras does scholaraio need for semantic search?

The embed extra adds sentence-transformers, numpy and faiss-cpu, and the topics extra adds bertopic and pandas on top of it. Neither is in the base install, whose only dependencies are requests, pyyaml, defusedxml and beautifulsoup4, so hybrid and semantic retrieval need an explicit extra.

### Does scholaraio 2.0 change my existing data layout?

Not for 1.4 or 1.5 users, who are told to upgrade the package, run scholaraio setup check and rebuild indexes when appropriate. Users on 1.3 or earlier still have to complete the explicit runtime migration described in docs/getting-started/upgrading-to-2.0.md.

### Which coding agents does scholaraio support?

The repository carries instruction and registration files for several of them, including .claude, .claude-plugin, .cursor, .qwen, .agents, .clinerules, .cursorrules and .windsurfrules, plus AGENTS.md, CLAUDE.md and clawhub.yaml. Setup for Claude Code plugins and for Codex or OpenClaw skill registration is covered in docs/getting-started/agent-setup.md.

### What PDF parsers does scholaraio use?

Two optional paths exist: the pdf extra installs pymupdf for local work, and the mineru-cloud extra installs a MinerU API client. Both are optional, and neither is in the base install, so the choice depends on whether your documents can be sent to a remote service.

## Sources

- [License: MIT](https://github.com/ZimoLiao/scholaraio/blob/main/LICENSE)
- [Project website](https://zimoliao.github.io/scholaraio/)
- [README](https://github.com/ZimoLiao/scholaraio/blob/main/README.md)
- [Releases](https://github.com/ZimoLiao/scholaraio/releases)
- [ZimoLiao/scholaraio on GitHub](https://github.com/ZimoLiao/scholaraio)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/zimoliao-scholaraio
