# Argus: a RAG platform that grades its evidence before it answers

> A Java 21 and Vue 3 knowledge base platform where retrieval runs twice, vector and keyword, merged with RRF, and an answer is only produced when the evidence passes four sufficiency levels. The pipeline is the interesting part; the release history is thinner, with no tagged releases and no licence file at the root.

**DevYangJC/Argus** — 🌱 Argus 是一个基于 RAG 架构的开源知识库平台，后端采用 Java 21 + Spring Boot + MyBatis-Plus + PostgreSQL/pgvector，前端采用 Vue 3 + TypeScript + Element Plus，AI 层基于 Spring AI Alibaba（通义千问）+ ReactAgent 图引擎，以 MinIO + Elasticsearch 为存储与检索引擎。

- Repository: https://github.com/DevYangJC/Argus
- Stars: 362 · Forks: 76
- Language: Java
- License: not declared
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/devyangjc-argus

## Four grades of evidence, and a refusal below the top one

The pipeline is drawn as a straight line with no shortcuts:

```text
文档上传 → 智能解析 → 文本切片
    ↓
向量嵌入（PGvector HNSW） + 关键词索引（Elasticsearch IK）
    ↓
用户提问 → 查询规划（LLM） → 混合检索（RRF 融合）
    ↓
证据评估（四级充分度） → LLM 生成 → 引用溯源
```

Query planning is an LLM step that chooses between DIRECT, REWRITE and DECOMPOSE and can fan out to at most three parallel retrievals, which is the project's answer to a vague question. The evidence check is the part worth copying: NONE, WEAK, PARTIAL and SUFFICIENT, with a refusal when the evidence does not reach the top level.

Every answer that survives carries citation snippets, the source document and a relevance score. The project describes itself as not being a search box wrapped around a chat model, and the grading step is where that claim is cashed: a wrong answer is prevented by refusing, not by a longer prompt.

## Retrieval runs on two channels and merges with RRF

Two search paths run in parallel and are fused with Reciprocal Rank Fusion. The semantic path uses PostgreSQL with pgvector, an HNSW index and COSINE_DISTANCE over 512-dimension embeddings, produced by the text-embedding-v3 model on DashScope. The exact path uses Elasticsearch 8.x with the IK Chinese analyser and BM25, scored in two stages with bool and rescore queries.

The reason for two channels is that each one fails differently. A vector index catches paraphrase and concept drift; BM25 with Chinese segmentation catches the exact product code or the exact term your users type. Fusing ranked lists rather than scores is the practical part, because the two systems do not produce comparable score scales.

Fusion is not the last step. Cluster aggregation and neighbour window expansion run afterwards to pull adjacent chunks in, because a retrieved fragment with no surrounding context tends to support a confident answer it does not actually contain.

## Memory is compressed three times before the model is called

Long conversations are handled by a BEFORE_MODEL hook that injects context before each call, in the order compact summary, session memory, recent messages. Underneath it sit three compression levels.

L1 is an incremental LLM summary that keeps key facts and decisions. L2 is a compact summary of history that discards redundant detail. L3 is runtime truncation, described as the last line of defence, applied when the token count passes 50000. Each level is cheaper and losier than the one before, so the ladder degrades quality in a fixed order rather than dropping whatever happens to overflow.

Output is streamed over SSE with delta de-duplication and an AGENT_MODEL_FINISHED fallback for the last chunk. The agent side runs on Spring AI Alibaba's ReactAgent graph engine with a think, tool call, generate chain, and two modes, CHAT for plain conversation and KB_SEARCH for retrieval, switchable inside one session. The agent decides whether to call the retrieval tool at all, at most once per turn so a turn cannot burn several searches.

## groupId filters both retrieval paths, and the role names disagree

Isolation is enforced in the query, not in the interface. Vector retrieval and Elasticsearch retrieval both carry a groupId filter, so a cross-group leak is prevented at the database and search layers rather than by hiding rows in the client.

Authentication is a JWT pair: an access token with a 15 minute lifetime and a refresh token held in an httpOnly cookie with rotation recorded in the database. Passwords are BCrypt hashed with a forced change on first use, and AOP operation logging records the sensitive calls.

One thing to resolve before writing a permission model against this code: the security section names three roles as Admin, Group Owner and Member, while the module section names the in-group roles as Owner, Manager and Member. Two lists of three that do not match is a small documentation inconsistency with real consequences for an access control system, so ask the maintainers which set the code implements.

## Upload is three phases, and SHA-256 skips the repeat

Document handling is built around resumable transfers. The protocol has three phases, init, chunk upload and complete, and it includes instant-upload detection keyed on SHA-256, so re-uploading a file the server already has does not push the bytes twice.

Parsing covers PDF, DOCX, MD and TXT with automatic encoding detection, handled by Apache PDFBox at 2.0.31 and POI at 5.2.5. From there the ETL pipeline runs asynchronously through Spring events with @Async and @Retryable, seven steps from parsing to indexing, with Spring Retry's @Recover for the failure path. Documents are held in MinIO, which is S3 compatible, and multi-part objects are merged with composeObject; the storage layer is assembled conditionally through @ConditionalOnProperty, which is why MinIO is listed as optional.

Soft delete is in the document module alongside preview and download, so removal is a state rather than a row.

## Standing it up means five services and one API key

The stated requirements are JDK 21, Node.js 20.19 or newer for the frontend build, PostgreSQL 16+ with the pgvector extension, Elasticsearch 8.x with the IK analyser plugin, MinIO if you want object storage, and a DashScope API key shared by the chat and embedding models.

The schema goes in first:

```bash
psql -h <host> -U <user> -d <database> -c "CREATE EXTENSION IF NOT EXISTS vector;"

psql -h <host> -U <user> -d <database> -f sql/schema.sql
```

MinIO is a container with a console on 9001 and a default bucket name of argus-rag-documents:

```bash
docker run -d --name minio \
  -p 9000:9000 -p 9001:9001 \
  -e MINIO_ROOT_USER=minioadmin \
  -e MINIO_ROOT_PASSWORD=minioadmin \
  minio/minio server /data --console-address ":9001"
```

The Elasticsearch step in the same quick start runs single-node with xpack.security.enabled=false. That is fine on a laptop and wrong on a shared host, since the search node holding your private document chunks would accept anonymous queries.

## No tagged release, no licence file, master branch

The repository has no GitHub releases at all, so there is no version to pin and no changelog to read. The default branch is master and the last push is dated 2026-08-16, with the repository not archived.

The tree is six directories and a few files: Argus-backend/, Argus-frontend/, sql/, docs/, AGENTS.md, CLAUDE.md, README.md, README_EN.md and .gitignore. There is no LICENSE file among them, and the repository metadata names no licence either. For a platform whose purpose is holding a company's private documents, that is the first question to ask, not a formality.

Two smaller details are worth knowing. sql/schema.sql is the single entry point for the database, which is convenient for a fresh install and unhelpful for a migration story, since no versioning tool appears in the tree. And the API documentation is served by Knife4j with SpringDoc at /doc.html, with online debugging, which makes exploring the endpoints the fastest way to learn the interface.

## Conclusion

Argus fits a team that needs answers traceable to documents it controls, and it does not fit someone who wants a weekend project, since standing it up means five services, a DashScope key and an IK analyser plugin. Before deploying it over real documents, settle the licence, which the repository does not state, and check the Elasticsearch node, because the documented quick start starts it with security disabled.

## FAQ

### What does Argus do when the knowledge base has no good answer?

It refuses. Evidence is graded on four levels, NONE, WEAK, PARTIAL and SUFFICIENT, and generation only runs on sufficient evidence. Every answer that is produced carries citation snippets, the source document and a relevance score.

### What has to be installed before Argus runs?

JDK 21, Node.js 20.19 or newer for the frontend build, PostgreSQL 16+ with the pgvector extension, Elasticsearch 8.x with the IK Chinese analyser plugin, optionally MinIO for object storage, and a DashScope API key shared by the chat and embedding models.

### How does Argus combine vector search and keyword search?

Two channels run in parallel and are merged with Reciprocal Rank Fusion. Semantic matching uses PGvector with an HNSW index and COSINE_DISTANCE over 512-dimension embeddings, while exact matching uses Elasticsearch with IK segmentation and BM25, followed by cluster aggregation and neighbour window expansion to avoid fragmented context.

### Can Argus hold a long conversation without filling the context window?

Yes. Memory is compressed in three stages, an incremental session summary, a compact summary of history, and runtime truncation when the token count exceeds 50000. A BEFORE_MODEL hook injects compact summary, then session memory, then recent messages before each model call.

### What licence is the Argus knowledge base platform released under?

The repository names no licence identifier and ships no LICENSE file at the top level, next to Argus-backend/, Argus-frontend/, sql/ and docs/. Confirm the terms with the maintainers before deploying it over documents you care about.

## Sources

- [DevYangJC/Argus on GitHub](https://github.com/DevYangJC/Argus)
- [Issues](https://github.com/DevYangJC/Argus/issues)
- [README](https://github.com/DevYangJC/Argus/blob/master/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/devyangjc-argus
