AutoRAG 2.0: a librarian agent for your local documents
AutoRAG: Now your agent can find anything in your computer. It gets smarter if you are using it frequently.
At a glance
- What is it?
- AutoRAG 2.0 is a TypeScript agent that searches configured directories in place, reads the source files itself and returns numbered knowledge units instead of file paths. It is not the Python RAG AutoML tool that used to carry the same name.
- Who is it for?
- Adopt AutoRAG 2.0 if your documents already live on disk in directories you control and you want curated answers rather than grep output, especially if you are willing to run Node 24 and let MinSync install a binary into the workspace. Do not adopt it if you need a hosted search API, a stable v1-compatible Python library, or per-file access control finer than configured search paths.
- Can I use it commercially?
- Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
- Is it still maintained?
- Yes. The repository last received commits 3 days ago.
- What is it written in?
- Mainly TypeScript, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 27, 2026, and from our analysis. They are not legal advice.
Editorial analysis
AutoRAG 2.0 is not the AutoRAG you may have read about
Search for AutoRAG and you will find the Python RAG AutoML tool, a paper, and a set of Cloudflare results that have nothing to do with this repository. The README addresses that head-on: the original Python project now lives in the legacy/ directory, and this repository hosts AutoRAG 2.0, described as "a self-evolving librarian agent for document collections." The legacy tool is not abandoned. According to the README it stays in maintenance mode with bug fixes, dependency updates and PyPI releases through pip install AutoRAG, while new feature development goes into 2.0. That split matters more than the version number. The two projects share a name, a topic list and almost nothing else: one optimizes RAG pipelines, the other is an agent that answers questions about files on your machine.
The package on npm is @autorag/librarian, version 2.4.2, and it ships a bin entry called autorag. The repository's primary language is TypeScript, the package declares "type": "module" and an engines field of node >=24.0.0. If your toolchain is pinned to Node 20 or 22, this is not a drop-in addition; you are changing runtimes first. The underlying agent loop comes from @earendil-works/pi-agent-core and @earendil-works/pi-coding-agent, both pinned at 0.85.1, so AutoRAG is best understood as a configured Pi agent rather than a retrieval library with a CLI bolted on.
Federated search instead of a central index
The README states the first design principle plainly: never migrate your data to search it. AutoRAG federates CLI-owned stores such as katok, discrawl, qmd, msgvault and rclone in place, and results carry source-native identities in the form kakao:<chat>/<sender>/<chunk> with scope-checked access. No ingestion step, no third-party server holding a copy of the corpus. For anyone who has watched a team spend a quarter building an index pipeline before answering a single question, that is the interesting part of the design.
Local retrieval is a different path. The README says MinSync indexes parsed markdown mirrors under .autorag for BM25, vector and hybrid retrieval, all sharing one CDC chunk lifecycle and wired through a RetrievalMethodRegistry. BM25, vector and hybrid are enabled by default whenever MinSync is enabled. You can turn local indexing off entirely with "minSync": false, or disable only lexical search with "bm25": false. The librarian invokes retrieval tools, then reads the underlying documents directly through its built-in bash tool, and curates one result set after ResultMerger normalizes scores and deduplicates. The agent opening the source file before answering is the detail that separates this from a vector store with a chat wrapper.
The retrieval method table in the README is worth reading as a statement of intent rather than a benchmark. Plain text and config files map to grep, research papers and dense prose to vector search, legal documents and specifications to BM25, mixed collections to hybrid. Nothing in the repository demonstrates that these pairings hold for your corpus. Treat them as defaults to test, not results.
Installing AutoRAG and running a first search
The package is published to npm as @autorag/librarian with the CLI exposed as autorag. The README does not print an install command, but package.json declares the bin entry and the engines constraint, so installation follows the standard global-install path. Node 24 or newer is required.
npm install -g @autorag/librarian
autorag --helpAfter that you should see the CLI's help output listing the available subcommands. The README references autorag init with --embedder-* flags for configuring minSync.embedder against a remote embedding endpoint, and a --minsync-max-chunk-size flag for minSync.maxChunkSize when a local embedder has a smaller context window. Those flag names are the ones the README gives; the full flag list is not reproduced here.
Bundled usage is shown as a constructor call. The README's example configures search paths and tunes the thin PDF extraction gate through parserOptions:
new AutoRAGAgent({
searchPaths: ["/path/to/documents"],
parserOptions: {
thinExtract: {
minPages: 3,
minChars: 800,
minCharsPerPage: 40,
timeoutMs: 30_000,
hybrid: "docling-fast",
hybridMode: "auto",
},
},
});The package exports both "." and "./core", with types at ./dist/index.d.ts and ./dist/core.d.ts, so the agent can be embedded in a Node application rather than only driven from the terminal. On first use, MinSync auto-installs a verified release into <workspace>/.autorag/bin when autoInstall is true, which is the default. Set "autoInstall": false only if you are managing that binary yourself. Answers come back as a structured SearchDocumentsResponse, with the real source file path or datasource id kept in the internal mapping for feedback and curation.
The PDF retry gate and its thresholds
The default PDF parser runs a cheap quality check on multi-page PDFs. When the local markdown is unusually sparse, defined as fewer than 800 characters or fewer than 40 characters per detected page across at least three pages, it retries through OpenDataLoader's docling-fast hybrid backend with hybridMode "auto" and a 30-second timeout. Dense PDFs are not retried. Hybrid is never used for single-page PDFs and never as the first path for images.
This is a reasonable design and also a source of surprise. A scanned report where the first three pages are cover art and a table of contents can trip the gate and send the whole document down the hybrid path with its 30-second timeout attached. The thresholds are exposed through trusted programmatic parserOptions, which means they are tunable, but the README frames them as parser-owned. If you have a corpus of image-heavy or oddly paginated PDFs, the numbers above are the first thing to measure against your own files. The README does not document what happens when the hybrid retry also fails or times out, which is a gap worth knowing about before you depend on it.
Self-evolving memory and what it actually learns
The README describes a self-evolving memory system: every search teaches the agent which retrieval methods work for which query types, which document areas are productive, and what the caller found useful through explicit feedback. A fresh install tries everything; a seasoned one knows where to look. The claim is that this is learned behavior from real usage rather than static configuration.
The mechanism is only partly visible in the repository. Feedback is explicit, and results carry their real source in an internal mapping for feedback and curation, which is the plumbing that would let the agent associate a retrieval strategy with a document region. What the README does not describe is where that memory is stored, how it is scoped, whether it survives a workspace move, or how you reset it when it has learned something wrong. For a feature presented as a core value, that is thin documentation. The practical consequence is that you should treat the memory as an opaque cache: useful, unversioned, and not something to reason about when an answer looks wrong. Delete the workspace state and the agent starts fresh, but the README does not say that in those words.
Where AutoRAG is the wrong tool
AutoRAG reads configured source directories directly through bash. That is what makes the answers grounded, and it is also the sharpest limitation. If your documents require per-file access control finer than the search paths you configure, this design gives you nothing. The README describes scope-checked access for federated datasource identities, but the local file path is opened by an agent with shell access to whatever you pointed it at. There is no documented per-document permission layer in the repository.
The second boundary is the name collision. If you arrived here looking for the Python AutoML tool that finds an optimal RAG pipeline, you want legacy/, not src/. The 2.0 CLI will not optimize your chunking strategy or sweep embedding models against a labeled QA set. It answers questions about files.
Third, the runtime constraint is real. Node >=24.0.0, an ESM-only package, and a first-use binary download into <workspace>/.autorag/bin. In an air-gapped environment or a locked-down CI image, the auto-install is a failure mode, not a convenience, and you would need autoInstall: false plus your own MinSync binary. The README gives the flag but not a procedure for supplying that binary.
AutoRAG 2.0 against the Python AutoRAG in legacy/
The most useful comparison is inside this repository. The legacy tool, installed with pip install AutoRAG, is a RAG AutoML system: you give it data and it searches for an optimal pipeline configuration. It is a build-time optimizer. AutoRAG 2.0 is a run-time agent: it retrieves, reads, judges and curates at query time, and its configuration surface is search paths, retrieval toggles and parser options rather than pipeline components.
The difference in approach shows up in what each one produces. The legacy tool's output is a pipeline you then deploy. AutoRAG 2.0's output is a SearchDocumentsResponse containing numbered knowledge units with page references, as in the README's Q3 report example where findings come back as three cited statements rather than a list of matches. If you already run the legacy tool in production, nothing here forces a migration, and the README explicitly says existing users can keep using it as before. The two can coexist: legacy to choose a pipeline, 2.0 to interrogate a document collection interactively.
The licence situation is worth a note. The repository's LICENSE file is present at the top level, and package.json declares "license": "MIT", but the repository metadata reports NOASSERTION, meaning GitHub could not match the file to a known licence. The npm package metadata is the clearer signal of intent. This is not legal advice; if the distinction matters to your organisation, read the LICENSE and NOTICE files directly.
Maintenance cost and what to check before upgrading
The last push to the default branch was on 2026-09-09, and the most recent releases are v2.4.1 on 2026-09-05, v2.4.0 on 2026-09-04 and v2.3.0 on 2026-08-24. Three releases in under two weeks, with package.json at 2.4.2, indicates a project moving quickly. Fast movement cuts both ways: fixes arrive soon, and so do behaviour changes. The thin-extract thresholds, the default-on retrieval methods and the auto-install path are all things a minor version could move.
The upgrade surface is mostly the workspace. MinSync installs a verified release into <workspace>/.autorag/bin, and the parsed markdown mirrors live under .autorag too. A version bump that changes the chunk lifecycle can invalidate that state, and the Makefile hints at the project's own answer: make e2e-live reuses clone-local state, while make e2e-live-cold deletes state and rebuilds. If you run AutoRAG in CI, mirroring that distinction is more honest than assuming an upgrade is transparent. The Makefile also documents E2E_ROOT and AUTORAG_LIVE_E2E_ROOT as the corpus root controls, and states that targets never create or bootstrap that root implicitly, which is a deliberate choice to keep test runs from writing into unexpected places.
The dependency footprint is larger than the feature list suggests: pi-agent-core, pi-ai, pi-coding-agent, pi-tui, @opendataloader/pdf, @rhwp/core, chardet, fast-xml-parser, iconv-lite, jszip, mailparser and slides-grab all appear in package.json. Several exist to parse formats beyond plain text and PDF. That is a wide surface to track for security updates, and the repository includes a supply-chain script and a Makefile target that runs license, NOTICE and local CycloneDX gates, which suggests the maintainers are aware of it.
Editorial conclusion
Adopt AutoRAG 2.0 if your documents already live on disk in directories you control and you want curated answers rather than grep output, especially if you are willing to run Node 24 and let MinSync install a binary into the workspace. Do not adopt it if you need a hosted search API, a stable v1-compatible Python library, or per-file access control finer than configured search paths. Before committing, verify that your Node version satisfies the engines field, that the MinSync auto-install is acceptable in your environment, and that the thinExtract thresholds match the PDFs you actually have.
Frequently asked questions
What is AutoRAG?
AutoRAG 2.0 is a self-evolving librarian agent for document collections, packaged on npm as @autorag/librarian. It searches configured directories, reads the source files itself and returns numbered knowledge units instead of raw matches.
What is Cloudflare AutoRAG?
The repository does not describe Cloudflare AutoRAG. This project is Marker-Inc-Korea/AutoRAG, a TypeScript librarian agent, and the README notes that the original Python RAG AutoML tool of the same name now lives in the legacy/ directory here.
Is RAG still relevant?
The README takes a position on this indirectly: AutoRAG 2.0 keeps retrieval (BM25, vector and hybrid over MinSync CDC chunks) but moves the judgment and curation into the agent, which reads source files before answering. Retrieval stays; the pipeline tuning does not.
Can you explain RAG to a beginner?
In this project's terms, RAG is retrieval plus an answer: AutoRAG retrieves candidate chunks through BM25, vector or hybrid search, then an agent opens the underlying files and curates one unified result set with page references instead of returning file paths and matching lines.
Is ChatGPT a RAG?
The repository does not discuss ChatGPT. It describes AutoRAG as a customized Pi agent whose model and provider come from the user's authenticated runtime, with retrieval running locally over MinSync CDC chunks.
What is RAG vs LLM?
AutoRAG's README separates the two by role: the LLM is the configured model that owns the retrieval, reading, judgment and curation loop, while RAG is the retrieval layer beneath it, here BM25, vector and hybrid search over parsed markdown mirrors under .autorag.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/marker-inc-korea-autorag)