# arxiv-mcp-server: the arXiv MCP server that reads LaTeX instead of scraping abstracts

> blazickjp/arxiv-mcp-server is a local Model Context Protocol server for agent literature work. It reads author-submitted LaTeX one section at a time, exports BibTeX from arXiv metadata, and keeps papers on disk.

**blazickjp/arxiv-mcp-server** — A local MCP server for agent literature work. Original-LaTeX section reads, BibTeX from arXiv metadata, and topic watches. Papers stay on disk. Search is optional.

- Repository: https://github.com/blazickjp/arxiv-mcp-server
- Website: https://www.pulsemcp.com/servers/blazickjp-arxiv-mcp-server
- Stars: 3,182 · Forks: 258
- Language: Python
- License: Apache-2.0
- Published: 2026-09-09 · Updated: 2026-09-09 · Language: en
- Canonical page: https://hysenlabs.com/projects/blazickjp-arxiv-mcp-server

## What arxiv-mcp-server actually solves for an agent

Most arXiv tooling hands a model an abstract and a title. That is enough to sort papers and not enough to write about them. The README frames the project around what it calls the literature loop: paper ID, then outline, then one section, then citations. Each step is a separate call, and the model is expected to walk the loop rather than swallow a paper whole.

The intended user is someone running an MCP-capable agent or IDE who needs to reason over a paper's actual argument. Reading the author-submitted LaTeX one section at a time keeps the context window bounded and keeps the model on the section it was asked about. Pulling BibTeX out of arXiv metadata rather than guessing at a citation string removes a class of fabricated references. Topic watches stay on disk, so a recurring interest does not have to be re-described in every session.

The project describes itself as a local MCP server, and the README is explicit that search, source retrieval, citation graphs and downloads call external services. The local part is narrower than the name suggests: what stays on your machine is the literature loop, not the whole pipeline.

## The data flow: stdio, a paper directory, and optional search

The server speaks MCP over stdio by default. Your client launches it as a subprocess, the two exchange MCP messages, and the client's model sees a set of tools. Downloaded papers land in a directory, and the README gives ~/.arxiv-mcp-server/papers as the default. That directory is the state: the watches and the stored papers live there, which is why the project can say papers stay on disk.

Search is the optional leg. The README states plainly that search, source retrieval, citation graphs and downloads call their respective external services, so a fully offline session is not what this is. The local claim is about where the reading loop runs and where the papers end up, not about avoiding the network.

The dependency list in pyproject.toml tells you the rest of the shape. The mcp package is pinned to >=1.27.0,<2.0.0, so the server tracks the 1.x protocol line. The arxiv package handles the metadata side. aiohttp and httpx cover network calls, aiofiles covers disk writes, and uvicorn, starlette and sse-starlette are present because the server can also be reached over SSE rather than stdio. PDF handling is not in the base install at all: pymupdf4llm and pymupdf-layout sit behind an optional extra named pdf. If you want PDFs converted, that is a deliberate add-on, not the default path.

## Installing arxiv-mcp-server and wiring the first call

The README's default install is a single uvx command. No repository clone and no Python environment setup are required, but uv must be present because uvx comes from it.

```bash
uvx arxiv-mcp-server
```

Running that starts the server. Clients that accept the mcpServers JSON shape, which the README names as Claude Desktop and Kiro, take this block. The type is stdio, the command is uvx, and the argument is the package name.

```json
{
  "mcpServers": {
    "arxiv": {
      "type": "stdio",
      "command": "uvx",
      "args": ["arxiv-mcp-server"]
    }
  }
}
```

After the client restarts, the arXiv tools should appear in its tool list. The README notes that other clients may use a top-level servers object, TOML, or their own settings UI, so the JSON above is a starting point rather than a universal answer.

To keep papers somewhere other than the default, append the storage flag to args. The path must be absolute.

```json
{
  "mcpServers": {
    "arxiv": {
      "type": "stdio",
      "command": "uvx",
      "args": ["arxiv-mcp-server", "--storage-path", "/absolute/path/to/papers"]
    }
  }
}
```

Claude Code has a one-line helper instead of hand-edited JSON. The user scope makes the server available across projects.

```bash
claude mcp add --transport stdio --scope user arxiv -- uvx arxiv-mcp-server
```

Codex and Hermes follow the same pattern with their own subcommands, and the README gives `codex mcp add arxiv -- uvx arxiv-mcp-server` and `hermes mcp add arxiv --command uvx --args arxiv-mcp-server` respectively. For Claude Code and Codex there is also a plugin route that installs the MCP connection together with a bundled arXiv research skill, registered from the repository as a marketplace. Verify a direct install with `claude mcp get arxiv` or `codex mcp get arxiv`; after installing the Claude Code plugin, restart the client or run `/reload-plugins`.

## The npm trap and other installation boundaries

The README carries an explicit warning that deserves more attention than it usually gets: an unrelated npm package uses the same name, so the server must not be installed with npm, pnpm, or npx. Anyone who reaches for `npx arxiv-mcp-server` out of habit will get something else. The supported package is on PyPI as arxiv-mcp-server==0.7.2.

There is a Dockerfile in the repository, which builds on ghcr.io/astral-sh/uv:python3.11-bookworm-slim, runs uv sync --frozen, and finishes on python:3.11-slim-bookworm with the entrypoint `python -m arxiv_mcp_server`. That is a build recipe for the project, not a documented deployment story. The README does not describe running the container as a service, does not publish an image name, and does not cover ports or environment variables for a networked deployment. If you want the server shared across a team rather than launched per client, the documentation is silent on how that should work.

Version support is another boundary. pyproject.toml sets requires-python to >=3.11 and classifies the project as Development Status 4 - Beta. The README's uvx path sidesteps that entirely, which is the reason to prefer it.

Finally, the README's own framing draws a line: search is optional. If your workload is mostly "find me papers about X", you are using the external services and the local loop is doing comparatively little work for you.

## Where the local-first design costs you

The strongest limitation is the one the project states about itself. Search, source retrieval, citation graphs and downloads all call external services. The papers stay on disk, but the discovery path does not, and neither does the fetch. On a restricted network, or in a setting where outbound calls need approval, the server's useful surface shrinks to whatever is already in the paper directory.

The LaTeX-first reading model has a second cost. Author-submitted LaTeX is the source the project prefers, and that is a real advantage for equations and section structure. It is also not universal. The README does not document what happens when a paper has no usable LaTeX source, and there is no fallback described. PDF support exists as an optional extra, which means the default install cannot read a PDF at all. If your reading list includes older papers or venue-published work outside arXiv, this is the wrong tool for that portion.

Beta status is worth taking at face value. Three releases landed inside three days in August 2026 (v0.7.0 on 2026-08-22, v0.7.1 on 2026-08-23, v0.7.2 on 2026-08-24), and the last push to the repository was on 2026-08-26. That is a fast-moving surface. The README does not document rollback, does not describe a migration path between minor versions, and does not say what happens to a paper directory written by an older release. Back up ~/.arxiv-mcp-server/papers before upgrading if the watches in it matter to you.

## How this differs from a general search MCP server

The obvious alternative is a general web or scholarly search MCP server: one that exposes a search tool and returns titles, abstracts and links. The difference is where the work happens. A search server answers "what exists" and hands the synthesis to the model, which then has to reconstruct the paper's argument from an abstract. arxiv-mcp-server answers "what does section 4 say", and it answers from the author's own LaTeX.

The practical consequence shows up in citations. A search-oriented server gives the model a metadata record it may paraphrase loosely. This server exports BibTeX from arXiv metadata, so the citation comes from the record rather than from the model's memory of the record. For anyone producing a bibliography, that is the difference between a plausible entry and a correct one.

A second alternative is a PDF pipeline: fetch the PDF, convert it to markdown, and let the model read the whole thing. The repository supports that shape through the pdf extra, so the two are not mutually exclusive. The trade-off is context. Converting a full paper to markdown puts the entire document in front of the model at once; the section-by-section loop deliberately does not. Which is better depends on whether you are asking a narrow question or trying to summarise a paper you have not read.

A third option is simply downloading papers yourself and pointing the agent at the files. That works until you need the metadata, the citation export, or a watch that survives the session. Those three things are the reason this project exists as a server rather than a script.

## Maintenance, licence, and what an upgrade costs

The repository is not archived, and the last push was on 2026-08-26, which is recent enough that the project is being worked on rather than parked. The release cadence around v0.7.x suggests the maintainer ships fixes quickly, and the MCP Registry badge in the README points at 0.7.2 as the latest listed version. There is a test workflow configured in .github/workflows/tests.yml, and the repository carries CONTRIBUTING.md and SECURITY.md alongside the usual pre-commit configuration.

The licence is Apache-2.0, declared in pyproject.toml with a LICENSE file at the repository root. That is a permissive licence with an explicit patent grant, and it is the same licence the README badge advertises. Nothing here suggests a separate commercial tier or a usage restriction on the server itself. The dependencies carry their own licences, and the README does not discuss them, so that is a question for your own dependency review rather than something this project answers. This is not legal advice.

Upgrade cost is the part the documentation leaves open. There is no changelog in the repository, no deprecation policy, and no statement about compatibility between 0.7.x releases. The mcp dependency is capped below 2.0.0, so a future protocol major version would be a breaking change by construction. The practical upgrade path is to re-run uvx, which fetches the current published version, and to check that your client still lists the tools afterward. If you have pinned arxiv-mcp-server==0.7.2 in a container image, you are the one deciding when to move.

## Conclusion

Adopt arxiv-mcp-server if you are wiring an MCP client such as Claude Code, Codex, Hermes, VS Code or Kiro into a reading loop where the paper has to be read section by section and cited correctly, and if you are comfortable that the package is versioned 0.7.2 and classified Beta. Do not adopt it if you want a hosted, multi-user paper service, if you need PDFs as the primary source, or if you plan to install it with npm, pnpm or npx: the README states an unrelated npm package uses the same name. Verify first that your client accepts the mcpServers JSON shape or has its own documented path, that ~/.arxiv-mcp-server/papers is a location you can write to or that you have passed --storage-path, and that your Python runtime is 3.11 or newer when you install from source rather than through uvx.

## FAQ

### What is the point of an MCP server?

An MCP server exposes tools to a model client over a defined protocol. In this project the server runs locally over stdio by default and gives the client a literature loop: read a paper's LaTeX one section at a time, export BibTeX from arXiv metadata, and keep topic watches on disk.

### What is a MCP data server?

The README describes this project as a local MCP server for agent literature work rather than a data server. What it keeps locally is the literature loop and the downloaded papers, in a directory that defaults to ~/.arxiv-mcp-server/papers; search, source retrieval, citation graphs and downloads call external services.

### What is arXiv and what is its purpose?

arXiv is the source of the author-submitted LaTeX and the metadata this server works from. That is what the server reads sections from and what it exports BibTeX from, which is why the README calls the metadata authoritative rather than relying on a model to reconstruct a citation.

### How trustworthy is arXiv?

The README does not assess arXiv's trustworthiness. It does say the server reads author-submitted LaTeX and exports BibTeX from authoritative arXiv metadata, so the citation side comes from the record rather than from the model's memory of it.

## Sources

- [blazickjp/arxiv-mcp-server on GitHub](https://github.com/blazickjp/arxiv-mcp-server)
- [License: Apache-2.0](https://github.com/blazickjp/arxiv-mcp-server/blob/main/LICENSE)
- [Project website](https://www.pulsemcp.com/servers/blazickjp-arxiv-mcp-server)
- [README](https://github.com/blazickjp/arxiv-mcp-server/blob/main/README.md)
- [Releases](https://github.com/blazickjp/arxiv-mcp-server/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/blazickjp-arxiv-mcp-server
