Open-source project
chrisryugj/Docufinder avatar
chrisryugj/Docufinder

Anything (Docufinder): offline full-text search for Korean documents on Windows and macOS

Anything — 파일 이름이 아니라 본문으로 찾는 문서 검색기. HWP·PDF·Office 전문 검색을 오프라인으로. Windows·macOS | Offline content search across Korean documents (HWP, PDF, Office). Windows and macOS

655 stars92 forksRustNOASSERTION

At a glance

What is it?
Anything indexes the body text of HWP, HWPX, DOCX, XLSX, PDF, image and .eml files inside a local SQLite FTS5 index. It is built for Korean office work, and the trade-offs sit in its licence and its platform split.
Who is it for?
Adopt it if your documents are Korean office files and you cannot send their contents to a cloud indexer: the installer carries its own parsers and OCR, and the README states that a default install downloads nothing at runtime. Do not adopt it if you need a permissive licence, a Linux build, or a headless index you can script from a server.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Yes. The repository last received commits 1 day ago.
What is it written in?
Mainly Rust, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem: filenames are not how Korean offices remember documents

Most desktop search tools index paths, sizes and modification dates. That works when you remember what a file is called. It fails when you remember a sentence. In Korean offices the document you need is often a HWP or HWPX file named after a date, a department code or a template number, and the only thing you actually recall is a clause inside it.

Anything (the repository is chrisryugj/Docufinder, the product name in the README is Anything) targets that gap. The README describes it as a 100% local search engine that finds documents by their contents rather than their names, and lists HWP/HWPX, Word, Excel, PowerPoint, PDF, images and .eml mail as the indexed formats. The intended user is someone on a Windows or macOS machine with a folder of Korean business documents and a reason not to upload them anywhere: the README states that search, indexing and OCR all work from the installer alone, and that a default configuration fetches nothing while running.

The audience is narrower than "anyone with files". HWP3 support (the 1996 to 2002 format) and EUC-KR/CP949 detection are the parts that a general-purpose indexer will not have, and they are the reason this project exists rather than a shell script around ripgrep.

How Anything indexes: parsers, a kordoc sidecar and SQLite FTS5

The stack is Tauri 2 with a React front end and a Rust backend, packaged as a desktop installer. The README's feature list names FTS5 for full-text search, and the topics on the repository include fts5, onnx, hwpx and korean-nlp, which matches the described pipeline: parse each supported format into text, run Korean morphological analysis over it, and store the result in a local full-text index.

Two mechanisms are worth calling out because they determine what the app can and cannot do.

First, HWP parsing runs through a Node sidecar. The Lite build notes state that the Lite distribution still spawns node.exe (described as the kordoc sidecar, required for HWP parsing) and msedgewebview2.exe as child processes. So even the build with no network code is not a single-binary application; it is a Tauri shell plus a Node process. On a machine with a process-blocking security policy, both need exceptions.

Second, the page-number feature added in v3.8.1 is not a text-layer trick. The README says the page badge on HWP, HWPX and PDF results is reconstructed from the typesetting information in files saved by Hangul, that it therefore applies only to documents re-indexed after the feature landed, and that files generated by other programs, which lack that typesetting data, keep working without page numbers. That is an honest constraint and a good example of the project's general posture: features are described with their failure cases attached.

Search itself has two paths. Filename search reads from an in-memory cache and works before indexing finishes, which the README compares to Everything. Content search waits on the index. The optional AI question answering and summarisation features sit on top and call either Gemini or an OpenAI-compatible server, including self-hosted options such as vLLM, Ollama and LiteLLM.

Installing Anything on Windows and macOS

The README does not document a package manager install. It points to the GitHub Releases page, where you pick a build for your machine. For a personal Windows 10 (21H2 or later) or Windows 11 PC, the README says one file is enough: Anything_<version>_x64-setup.exe, roughly 380 MB.

If you are on a corporate PC without administrator rights, on Windows 10 LTSC, or you hit a WebView2 error on first launch, the README says to download two files and install the runtime first.

bash
# 1. Download both files from the Releases page
#    Anything_<version>_x64-setup.exe
#    MicrosoftEdgeWebView2RuntimeInstallerX64.exe

# 2. Run the WebView2 runtime installer first (administrator rights if available)
MicrosoftEdgeWebView2RuntimeInstallerX64.exe

# 3. Then run the app installer
Anything_<version>_x64-setup.exe

After installation, the first real use is registering a folder. The README describes the flow as: add a folder, let it index automatically, then type a keyword into the search box and get hits from the body text. Filename search responds immediately from the in-memory cache, so you can confirm the app is alive before the index is warm.

If you are moving the installer across an air-gapped or one-way transfer link, the README is explicit that the old combined installer was withdrawn because large single files arrived corrupted, and that the two files should be carried separately, one at a time, with the WebView2 runtime installed first.

There is also a documented offline switch. Setting the environment variable DOCUFINDER_OFFLINE=1 disables the layout-analysis feature (which the README calls online-only) while leaving OCR, search and indexing working.

bash
set DOCUFINDER_OFFLINE=1

For macOS, the README offers an Apple Silicon build. Note that the README's own requirements section states that the Lite build and macOS use manual updates, so the in-app update notification does not cover them.

The Lite build, and why it is the most interesting decision in the repository

Anything Lite is a separate Windows installer (Anything.Lite_<version>_x64-setup.exe, about 280 MB) introduced in v3.7.0. The stated reason is that internal-network PCs running endpoint security products such as AhnLab V3 or ZombieZERO were quarantining the app or blocking it from starting.

The response is not a bypass. The README says the detection surface was removed at compile time: the Lite build does not link an HTTP client at all, so no HTTP/TLS stack or external host strings end up in the binary, and the CI pipeline verifies this on every release. That is a claim an IT department can check with strings, which is a more useful form of assurance than a vendor statement.

The cost is a real feature split. Lite drops semantic and hybrid search, OCR, AI question answering and summarisation, automatic updates and automatic error reporting. It keeps full-text search with Korean morphological analysis, proximity search and search operators, filename search, HWP/HWPX/DOCX/XLSX/PDF and .eml parsing with preview, live folder watching and incremental indexing, document comparison, version lineage, tags and bookmarks, duplicate finding, and export.

Two details matter for anyone evaluating it. Settings and databases are stored separately from the standard build, so the two do not share state. And copying a settings.json from the standard build, or editing it by hand, will not re-enable semantic search, OCR or AI in Lite. If your requirement is "OCR on scanned PDFs", Lite is the wrong download and no configuration will fix that.

Where Anything is the wrong tool

The licence is the first thing to check, and the material is inconsistent about it. The repository metadata reports NOASSERTION. The README badge and package.json both declare BUSL-1.1, the Business Source License 1.1. Those are not the same signal, and anyone evaluating this for commercial deployment should read the LICENSE file directly rather than trusting either. BUSL is source-available, not open source in the OSI sense, and its terms typically include a change date and a usage restriction. The README does not summarise those terms.

Platform coverage is the second limit. The README's download table is Windows and macOS only, with the Lite build Windows-only. There is no Linux build and no documented headless or server mode, so you cannot run this as a shared index behind a web interface. It is a desktop application.

Third, the AI features are not offline in the way the search is. Semantic search downloads an embedding model (roughly 106 MB) the first time you enable it, and question answering and summarisation require an API key for Gemini or an OpenAI-compatible endpoint. The README is clear that search still works without a key, but if your reason for choosing this project is that nothing leaves the machine, the AI features are the part that does not fit that description unless you point them at a local server.

Fourth, the DRM fallback is a narrow, Windows-only escape hatch. In environments where a security product encrypts documents so that only whitelisted applications can read them, the README says all parsers fail and the files look corrupt. The v3.8 setting that reads files through installed MS Office is Windows-only, off by default, runs only on files where the normal parsers already failed, hides the Office window, and blocks automatic macro execution. It is a fallback, not a supported ingestion path, and it depends on the machine having licensed Office installed.

Alternatives and how their approach differs

The README itself names Everything as the reference point for filename search, and the package.json description calls Anything an "Everything 대항마" (a rival to Everything). The comparison is worth taking seriously because the two tools solve different problems. Everything indexes the NTFS master file table and returns filename matches in milliseconds across millions of files. It does not parse document bodies, so it cannot find a clause inside a HWP file. Anything keeps a filename cache for the same instant-lookup feel, but its reason to exist is the content index underneath, which is why it ships parsers, a Node sidecar for HWP, and optional OCR models. If your problem is "where did I save this file", Everything is faster and has no index build to wait for. If your problem is "which contract mentions this payment term", Everything will not help.

For a cloud-based document assistant, the difference is where the index lives. Anything's default install keeps parsing and indexing on the machine and, per the README, makes no runtime downloads; the AI layer is opt-in and points at a provider you choose. That is a different trust model from uploading a folder to a hosted service, and it is the reason the Lite build exists at all.

For a general-purpose desktop search tool that handles PDF and Office but not Korean formats, the gap is the parser layer: HWP3, HWP5, HWPX and CP949/EUC-KR detection are not things you get for free, and the README treats them as the core of the product rather than an add-on.

Maintenance, releases and what an upgrade costs you

The repository is not archived, and the last push was on 2026-09-05. The release history is dense: v3.8.2 on 2026-08-14, v3.8.3 on 2026-08-27, and v3.8.4 on 2026-09-05. The README's version badge still reads 3.8.3 while package.json reads 3.8.5, which suggests the documentation lags the code by a release or two. Treat the README as a description of intent and the release notes as the record of what shipped.

Upgrade cost is not uniform across builds. The README says the app notifies you inside the app when a new version appears, but that Lite and macOS are manual updates. If you deploy Lite across an internal network, you own the update process: download, transfer, reinstall, per machine.

There is one upgrade cost that is specific to this project and easy to miss. The page-number feature only applies to documents re-indexed after v3.8.1. If page badges matter to your workflow, an upgrade means a re-index of the affected folders, not just a new binary.

On licensing, the practical step is to read LICENSE and THIRD_PARTY_NOTICES.md in the repository root before you plan a deployment. The repository carries both files, and the third-party notices matter here because the app bundles a Node sidecar and OCR models. This is not legal advice; it is a pointer to the two files that answer the question.

Editorial conclusion

Adopt it if your documents are Korean office files and you cannot send their contents to a cloud indexer: the installer carries its own parsers and OCR, and the README states that a default install downloads nothing at runtime. Do not adopt it if you need a permissive licence, a Linux build, or a headless index you can script from a server. Before installing, check the release page for the WebView2 runtime file if you are on an LTSC or locked-down Windows image, and read the LICENSE file, because the repository metadata reports NOASSERTION while package.json declares BUSL-1.1.

Frequently asked questions

Does Anything (Docufinder) work without an internet connection?

Yes. The README states that a default install performs no runtime downloads and that search, indexing and OCR all work from the installer alone. The optional semantic search, formula recognition and layout analysis each download a model once when you enable them, and DOCUFINDER_OFFLINE=1 disables layout analysis entirely.

Which document formats can Anything search inside?

The README lists HWP and HWPX, DOCX, PPTX, XLSX and XLS, PDF, images (JPG, PNG, WEBP, BMP, TIFF via OCR), TXT and MD with EUC-KR/CP949 detection, and EML mail files. HWP3 files from 1996 to 2002 are converted through the kordoc engine.

Is Anything free to use in a company?

The licence question is not settled by the README. The repository metadata reports NOASSERTION, while the README badge and package.json both declare BUSL-1.1, the Business Source License 1.1. Read the LICENSE file in the repository root for the actual terms.

Why is there a separate Anything Lite installer?

The README says internal-network PCs running endpoint security products such as AhnLab V3 or ZombieZERO were quarantining the app or blocking startup. Lite is a build with the network code removed at compile time, at the cost of semantic search, OCR, AI features, automatic updates and error reporting.

Official sources

  1. chrisryugj/Docufinder on GitHub
  2. Issues
  3. README
  4. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/chrisryugj-docufinder.svg)](https://hysenlabs.com/projects/chrisryugj-docufinder)