# Paper-Agent: A Multi-Agent Workbench for Academic Paper Search and Automated Literature Review

> Paper-Agent is a Python workbench that takes a research topic, searches arXiv, OpenAlex, and Semantic Scholar, reads full-text papers through a progressive pipeline, runs layered analysis, and produces a structured literature review. The 2.0 version is a full rewrite using LangGraph, Vue 3, uv package management, and SQLite session persistence.

**Tswoen/Paper-Agent** — Paper-Agent 是一个面向科研人员和学生的智能论文检索与调研工具。项目基于多智能体协作架构（LangGraph），通过自然语言处理（NLP）、自动化搜索，帮助用户高效查找学术论文、分析文献内容，并进行论文调研。Paper-Agent 支持多平台集成、关键词搜索、自动分析、论文调研，提升了学术研究的效率。适用于论文写作、学术调研、科研项目管理等多种场景，是学术调研的理想助手。

- Repository: https://github.com/Tswoen/Paper-Agent
- Stars: 459 · Forks: 45
- Language: Python
- License: not declared
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/tswoen-paper-agent

## What Paper-Agent Solves and Who It Is For

A standard literature survey involves searching multiple databases, opening dozens of abstracts, deciding which papers to read in full, extracting relevant claims, finding the conceptual structure across papers, and finally writing a coherent overview of the field. Each step is time-consuming, and interruptions, such as a PDF download that fails or a model timeout, can lose partial progress.

Paper-Agent addresses this by automating the pipeline from a research topic input to a structured literature review document, while making each step observable in real time through Server-Sent Events. Version 2.0, which is a complete rewrite of the earlier 1.x codebase, uses LangGraph for the agent workflow, FastAPI for the backend API, Vue 3 and TypeScript for the frontend, SQLite and the local filesystem for session persistence, and ChromaDB for vector-based full-text retrieval. The last push was on 2026-09-08.

The target audience is researchers and graduate students who need to survey an unfamiliar field quickly, with the requirement that the review be traceable to specific papers and include structured analysis of research trends, consensus points, open questions, and temporal evolution of the field.

## Five-Agent LangGraph Pipeline

The workflow is implemented as a LangGraph graph with five specialized agents. SearchAgent generates keyword and subtopic plans from the research topic and runs searches across arXiv, OpenAlex, and Semantic Scholar simultaneously, deduplicating and scoring results by relevance and recency. ReadAgent reads abstracts first to assess relevance before committing to full-text processing. Papers that pass the relevance check are downloaded as PDFs, converted to Markdown, split into chunks, and written to the ChromaDB vector store.

AnalyseAgent performs analysis in two passes: a per-subtopic pass that examines each paper's contribution to a specific angle, then a global synthesis pass that produces structured output covering research state, consensus, contested points, gaps, temporal evolution, and outlook. WritingOutlineAgent generates a chapter outline and maps evidence to sections. WritingAgent writes each section against the evidence, searches for additional supporting papers if evidence is insufficient, and performs a review pass with limited revision iterations.

The README documents a fallback mechanism: if a dependency required for full-text processing is unavailable, the system saves a recovery checkpoint. After fixing the configuration, the pipeline can continue from that checkpoint without reprocessing the papers it already handled.

## Installing Paper-Agent and Starting the Services

Paper-Agent uses uv for Python dependency management. In the project root directory:

```bash
uv init
uv venv --python 3.12
uv sync
npm run front:install
```

This installs Python dependencies and the Vue 3 frontend package. Python 3.12 or higher is required. The pyproject.toml dependencies include langchain (via langgraph>=1.2.6), fastapi>=0.116.0, chromadb>=1.5.9, pypdf>=5.0.0, anthropic>=0.111.0, and openai>=2.43.0.

Start the backend:

```bash
uv run python main.py
```

The backend starts on 127.0.0.1:8000 with hot reload in development mode. Start the frontend in a second terminal:

```bash
npm run front:dev
```

The frontend opens at http://127.0.0.1:5173/ and proxies /api and /webui requests to port 8000. For network access from other devices on the local network, the README documents an alternative frontend command: npm run front:dev:network.

## Model Configuration and Provider Backends

Paper-Agent supports multiple LLM providers configured in config/model.json. Supported backend types are openai, openai_compat, anthropic, and anthropic_compat. An example openai_compat provider configuration looks like:

```json
{
  "providers": {
    "my_provider": {
      "backend": "openai_compat",
      "api_key_env": "OPENAI_API_KEY",
      "api_base": "https://api.openai.com/v1",
      "extra_headers": {},
      "extra_body": {}
    }
  }
}
```

The compat variants allow routing to any API endpoint that speaks the OpenAI or Anthropic protocol, including local models. api_key and api_key_env are interchangeable; using api_key_env is recommended so the key does not appear in the config file.

The README describes three agent tiers (luna_agent for search, default_agent for reading and writing, solar_agent for analysis), each of which can be assigned a different model. This allows cheaper models to handle lightweight tasks like abstract reading while reserving a stronger model for the synthesis and writing steps. If luna_agent or solar_agent are not configured, they fall back to default_agent. A missing default_agent makes the configuration non-functional.

## Limitations: License, ChromaDB Cold Start, and Token Cost

The repository carries no license file. The pyproject.toml lists no license field and there is no LICENSE file in the top-level entries. This means the default copyright applies and usage rights are undefined without clarification from the maintainer. Teams that need a clear open source grant before adopting a tool should verify this before deploying Paper-Agent in an institutional or team context.

ChromaDB, which stores the chunked full-text vectors, initializes a new collection per session. For a long survey covering many papers, the vector store grows large and ingestion time becomes significant. The system does not degrade gracefully when ChromaDB is unavailable: the README documents the fallback checkpoint for network failures or PDF download issues, but not for vector store failures specifically.

The token cost of running a full-text analysis pipeline across many papers can be substantial. The README acknowledges this by allowing different model tiers per agent stage, but it does not provide guidance on typical token usage for a given number of papers or topic breadth. The workbench displays actual token usage per stage, which at least makes cost visible.

## Comparison with Elicit and Manual Database Search

Elicit is a web-based tool for academic literature search that uses AI to extract key findings from papers and display them in a structured table. It is hosted, requires no installation, and focuses on structured extraction from abstracts rather than full-text analysis. Elicit does not write a narrative literature review; it provides a structured view of individual paper claims that a researcher uses as input for their own writing.

Paper-Agent's difference is the full pipeline from search to written review. It reads full text, builds a vector store, performs layered analysis across subtopics, generates an outline with evidence mappings, and produces draft sections. The result is a narrative draft, not a table. The trade-off is setup complexity: Elicit requires a browser; Paper-Agent requires Python 3.12, Node.js, a working LLM provider, an embedding provider, and the ChromaDB vector search infrastructure.

For a researcher who needs a quick structured comparison of a few dozen papers, Elicit or Semantic Scholar's native paper recommendations are simpler. Paper-Agent is more appropriate when the goal is to produce a structured narrative review of a field, when the researcher wants to control which LLM processes the papers, or when the institution requires data to remain local rather than being sent to a third-party service.

## Conclusion

Paper-Agent is a practical choice for researchers and graduate students who spend significant time on literature surveys and want a tool that handles the search-to-draft pipeline end to end. It is not a plug-and-play solution: the model configuration requires a working LLM provider and embedding service before the workbench produces useful output, and full-text PDF processing requires the pdf and vector search dependencies to be available. Before relying on it for a real survey, verify that the model provider supports the token volumes that full-text analysis demands. The repository carries no license file, so check the repository directly for usage rights before adopting it in a team or institutional context.

## FAQ

### What paper databases does Paper-Agent search?

Paper-Agent searches arXiv, OpenAlex, and Semantic Scholar through built-in connectors. Results are unified into a common PaperDocument format, then deduplicated and scored by relevance, source, date, and exclusion keywords before the reading pipeline processes them.

### Can Paper-Agent recover if the backend is restarted mid-pipeline?

Yes. The README describes a recovery mechanism: if a required dependency is unavailable during the reading phase, the system saves a checkpoint. After fixing the configuration and restarting, the pipeline resumes from that checkpoint without reprocessing papers already handled in the current session.

### Does Paper-Agent support local LLMs instead of cloud providers?

Yes. The openai_compat and anthropic_compat backend types allow any API endpoint that speaks those protocols, including locally hosted models. The api_base field in the provider configuration points to the local endpoint, and api_key_env can be set to any available environment variable.

## Sources

- [Issues](https://github.com/Tswoen/Paper-Agent/issues)
- [README](https://github.com/Tswoen/Paper-Agent/blob/main/README.md)
- [Tswoen/Paper-Agent on GitHub](https://github.com/Tswoen/Paper-Agent)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/tswoen-paper-agent
