Argus: A Java RAG Knowledge Base Platform with Hybrid Retrieval and ReactAgent
🌱 Argus 是一个基于 RAG 架构的开源知识库平台,后端采用 Java 21 + Spring Boot + MyBatis-Plus + PostgreSQL/pgvector,前端采用 Vue 3 + TypeScript + Element Plus,AI 层基于 Spring AI Alibaba(通义千问)+ ReactAgent 图引擎,以 MinIO + Elasticsearch 为存储与检索引擎。
At a glance
- What is it?
- Argus is an enterprise-grade RAG knowledge base platform built on Java 21, Spring Boot 3.5.0, and Spring AI Alibaba. It combines PGvector semantic search with Elasticsearch keyword retrieval using RRF fusion, evaluates the evidence quality at four levels before answering, and refuses to respond when the retrieved material is insufficient.
- Who is it for?
- Java engineering teams that already operate PostgreSQL, Elasticsearch, and MinIO and want to add a RAG-based knowledge base without adopting a Python-first stack will find Argus worth evaluating. The four-level evidence evaluation and the ReactAgent's refusal-to-answer behavior are specific design choices that reduce hallucinated responses at the cost of occasionally declining answerable questions.
- Can I use it commercially?
- Not without permission. GitHub finds no licence file in the repository, and without a licence all rights are reserved by default: you may read the code but not reuse it. Check the README, or ask the authors, before using it.
- Is it still maintained?
- Yes. The repository last received commits 44 days ago.
- What is it written in?
- Mainly Java, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 27, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What Argus Is and Who It Is For
Many enterprises upload PDF documents, DOCX files, and Markdown wikis to a shared storage system and then struggle to query them. A general-purpose LLM will answer questions about those documents confidently but incorrectly, because it has no access to the actual content. The standard RAG pattern addresses this by retrieving relevant chunks first, then using the LLM only to generate a response grounded in the retrieved text. The difficulty is in the details: how to split documents, how to combine vector and keyword retrieval, how to evaluate whether the retrieved evidence is actually sufficient, and how to manage conversation history in a multi-turn session.
Argus is a Java-native platform that implements the full RAG pipeline from document upload through answer generation, with group-based access control and a Vue 3 frontend. Its target environment is an enterprise Java team that wants to keep its AI knowledge base inside the existing JVM and Spring ecosystem rather than adopting a Python-based tool. The AI layer runs through Spring AI Alibaba, which connects to Tongyi Qianwen (DashScope) for both the chat model and embeddings.
Hybrid Retrieval Architecture: PGvector and Elasticsearch with RRF Fusion
Argus runs two retrieval paths in parallel and fuses their results. The semantic path uses PostgreSQL with pgvector: document chunks are stored as 512-dimensional embeddings in an HNSW index and queried using COSINE_DISTANCE. The keyword path uses Elasticsearch 8.x with the IK Chinese tokenizer and BM25 ranking, which handles precise technical term matches that semantic search tends to miss.
The results from both paths are merged using Reciprocal Rank Fusion (RRF), a rank-based fusion method that does not require calibrating scores across different retrieval systems. After RRF, the system applies cluster aggregation and neighbor window expansion to supplement the top chunks with adjacent context, avoiding the fragmented snippets that appear when a chunk boundary cuts through relevant content.
The retrieval process begins with LLM query planning: the system classifies the user's question as DIRECT (pass through), REWRITE (reformulate for retrieval), or DECOMPOSE (break into sub-questions), and runs at most three sub-queries in parallel. Before generating an answer, the system evaluates the retrieved evidence at four levels: NONE, WEAK, PARTIAL, or SUFFICIENT. When the evidence is NONE or WEAK, the system refuses to answer rather than generating a response it cannot support with document evidence.
Setting Up the Infrastructure
Argus requires four backend services: PostgreSQL 16+ with pgvector, Elasticsearch 8.x with the IK Chinese tokenizer plugin, MinIO for object storage, and a DashScope API key for the Tongyi Qianwen model. JDK 21 and Node.js 20.19 or higher are required for building.
First, enable the pgvector extension and run the schema script:
psql -h <host> -U <user> -d <database> -c "CREATE EXTENSION IF NOT EXISTS vector;"
psql -h <host> -U <user> -d <database> -f sql/schema.sqlStart MinIO with Docker:
docker run -d --name minio \
-p 9000:9000 -p 9001:9001 \
-e MINIO_ROOT_USER=minioadmin \
-e MINIO_ROOT_PASSWORD=minioadmin \
minio/minio server /data --console-address ":9001"Start Elasticsearch with the security features disabled for local development:
docker run -d --name elasticsearch \
-p 9200:9200 -p 9300:9300 \
-e "discovery.type=single-node" \
elasticsearch:8.xThe README documents seven required configuration parameters in the backend's application configuration file, including the DashScope API key, database URL, Elasticsearch address, and MinIO connection settings. The IK tokenizer plugin must be installed into the Elasticsearch container separately.
ReactAgent and Memory Compression in the AI Assistant
Argus includes an AI assistant separate from the knowledge base Q&A module. The assistant is built on the Spring AI Alibaba ReactAgent graph execution engine and operates in two switchable modes: CHAT for general conversation and KB_SEARCH for knowledge base retrieval. Both modes can be used within the same session.
The assistant uses a BEFORE_MODEL hook to inject context before each model call: it prepends the compact summary, then the session memory, then the most recent messages in order. Three-level memory compression manages long conversations within the context window limit. The L1 level accumulates incremental LLM summaries of the conversation, retaining key facts and decisions. The L2 level distills L1 into a more compressed summary by discarding redundant detail. The L3 level is a runtime truncation guard that activates when the total token count exceeds 50,000, keeping the tail end of the conversation history to fit within the window.
SSE streaming sends the model's reply as delta tokens to the frontend. The implementation includes delta deduplication and an AGENT_MODEL_FINISHED event to handle edge cases in the streaming protocol.
Document management supports PDF, DOCX, MD, and TXT formats. The upload protocol uses a three-phase process (init, chunk upload, complete) that supports resumable uploads and SHA-256 hash-based deduplication to avoid re-uploading identical files.
Limitations and Alternative Approaches
The DashScope API key is a hard requirement. Tongyi Qianwen is used for both the chat model and the text-embedding-v3 embedding model through DashScope. The README mentions OpenAI-compatible mode for the embedding model, but the chat model integration is through Spring AI Alibaba's DashScope native integration. Teams outside Alibaba Cloud's service area or without a DashScope account must verify compatibility with their model provider before using Argus.
The infrastructure footprint is significant: PostgreSQL with pgvector, Elasticsearch with a specific plugin, MinIO, and a running Spring Boot backend. A team without existing Elasticsearch infrastructure will need to set up and operate an Elasticsearch cluster as a dependency. The IK Chinese tokenizer is specific to Chinese-language content; teams working with non-Chinese documents may find the keyword retrieval less effective.
The repository lists no license. Using Argus in a production or commercial environment without a clear license is a legal risk. There are no GitHub releases or version tags.
Dify is a comparable RAG-based knowledge base platform that supports multiple AI providers (OpenAI, Anthropic, Hugging Face, and others) through its model provider marketplace. It ships as a self-hosted Docker Compose deployment and does not require Java. The trade-off is that Dify's architecture is fixed; it does not expose the Spring AI Alibaba ReactAgent graph for custom node development the way Argus does.
Technology Stack and Security Design
The backend runs Java 21 and uses virtual threads and Record syntax from modern Java. The ORM is MyBatis-Plus with lambda-style type-safe queries. The retry strategy uses Spring Retry's @Retryable and @Recover annotations. API documentation uses Knife4j with SpringDoc, accessible at /doc.html. The ETL pipeline uses Spring Events and @Async for asynchronous processing with automatic retry on failure.
Authentication uses JJWT with HMAC-SHA256 signed JWTs. Access tokens expire after 15 minutes; refresh tokens use httpOnly cookies with database-backed rotation to prevent token reuse. Passwords are hashed with BCrypt. The three-tier group role model (Admin, Group Owner, Member) applies group ID filters on both the vector retrieval and Elasticsearch queries, so a group member cannot retrieve documents from other groups even if the embedding similarity is high.
The last push to this repository was on 2026-08-16. The repository is not archived. No license file is present in the repository.
Editorial conclusion
Java engineering teams that already operate PostgreSQL, Elasticsearch, and MinIO and want to add a RAG-based knowledge base without adopting a Python-first stack will find Argus worth evaluating. The four-level evidence evaluation and the ReactAgent's refusal-to-answer behavior are specific design choices that reduce hallucinated responses at the cost of occasionally declining answerable questions. The DashScope API key requirement ties the AI layer to Tongyi Qianwen and Alibaba Cloud; teams outside China or without DashScope access must verify whether the OpenAI-compatible embedding path covers their model provider before committing. The repository lists no license, which is a material constraint before any commercial or production deployment. The last push was on 2026-08-16 and there are no releases.
Frequently asked questions
What AI models does Argus support?
Argus uses Spring AI Alibaba to connect to Tongyi Qianwen (DashScope) for both the chat model and the text-embedding-v3 embedding model. The README mentions OpenAI-compatible mode for the embedding layer, but the chat integration is DashScope-native.
What document formats does Argus accept?
The ETL pipeline accepts PDF, DOCX, MD, and TXT files. The README documents automatic encoding detection for text files. Uploads use a three-phase protocol with SHA-256 deduplication to prevent re-uploading identical documents.
What happens when Argus cannot find sufficient evidence to answer a question?
The system evaluates retrieved evidence at four levels: NONE, WEAK, PARTIAL, and SUFFICIENT. When evidence is NONE or WEAK, the system refuses to answer rather than generating a response unsupported by the documents. This is controlled by the query planning and evidence evaluation stages in the RAG pipeline.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/devyangjc-argus)