Model or dataset
dtsola/xiaoyaosearch avatar
dtsola/xiaoyaosearch

XiaoyaoSearch (小遥搜索): a local desktop search tool that takes voice, text and images

小遥搜索,听懂你的话、看懂你的图,用AI找到本地任何文件。让搜索像聊天一样简单。XiaoyaoSearch: Understands your words, reads your images, finds any local file with AI. Making search as easy as chatting.

1,038 stars54 forksPythonNOASSERTION

At a glance

What is it?
XiaoyaoSearch is a cross-platform desktop application that indexes local files with BGE-M3, FasterWhisper and CN-CLIP and searches them through Faiss and Whoosh. It is aimed at knowledge workers and developers who want chat-style retrieval over their own disk, but the licence and the install path both need a close look before adoption.
Who is it for?
Adopt XiaoyaoSearch if you are a Windows user who wants semantic search over a personal document, audio and video collection without sending files to a cloud service, and you accept that the free tier is non-commercial. Do not adopt it if you need a permissively licensed component in a commercial product, or if you want a documented rollback and upgrade path: the README does not document either.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Yes. The repository last received commits 132 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The gap XiaoyaoSearch is trying to fill

Desktop search on most systems is filename matching. You remember that a file exists but not what it was called, and the search box is useless. XiaoyaoSearch targets that specific failure: the README describes it as a cross-platform local desktop application for knowledge workers, content creators and technical developers that converts voice, text and image input into a semantic query and runs it against local files.

The supported content types are broader than plain documents. According to the README, search covers video (mp4, avi), audio (mp3, wav) and documents (txt, markdown, office, pdf), and it applies to both file contents and filenames. That matters because the interesting material is often spoken or visual: a recorded meeting, a screenshot, a slide deck. A text-only indexer cannot reach any of it.

The audience is narrow on purpose. This is a single-machine tool for someone with a large personal corpus, not a team search service. The README lists no server deployment, no multi-user accounts and no shared index. If your problem is finding documents across a team, this is the wrong shape of tool.

How the hybrid retrieval pipeline is put together

The architecture is a two-process desktop app. The frontend is Electron with Vue 3 and TypeScript, using Ant Design Vue, Pinia and Vite. The backend is Python 3.10 with FastAPI and Uvicorn. The two are packaged together, which is why the install instructions talk about starting a backend service and a frontend service rather than launching a single binary.

Retrieval is hybrid, and the README names both halves: Faiss for vector search and Whoosh for full-text search. Each has its own index directory in the repository layout, data/indexes/faiss/ and data/indexes/whoosh/, with SQLite at data/database/ as the metadata store. Hybrid here means the two result sets have to be combined, and the README does not describe the fusion or ranking step, so how a keyword hit and a vector hit are ordered against each other is not something you can reason about from the documentation.

The AI layer is four models with distinct jobs. BGE-M3 produces embeddings for text and documents. FasterWhisper handles speech, which is what makes audio and video content searchable at all. CN-CLIP handles images, so an uploaded picture can be matched against indexed visual content. Ollama runs a local LLM, and the quick-start tells you to pull qwen2.5:1.5b. Since v1.3.0 the README also lists cloud model support through OpenAI, DeepSeek and Aliyun compatible APIs, and since v1.6.0 the same for embedding APIs, with switching between local and cloud.

That last point is the real design decision. Local models keep data on disk, but the README's own hardware guidance is 16GB of memory and an RTX3060 with 6GB, which is a real machine requirement, not a formality. The cloud option trades that away and the README frames it as a privacy and performance trade-off the user chooses. Note that the trade-off is not symmetric: once you point the embedding API at a cloud provider, the text being embedded leaves the machine, and the README does not document what is sent or whether it can be scoped.

Installing XiaoyaoSearch: the all-in-one package and the developer path

There are two documented routes and they are not equivalent. The all-in-one package is described as recommended for ordinary users, and the README states it supports Windows only. It is distributed through Baidu Netdisk, with a link and extraction code in the README, and the example archive name is XiaoyaoSearch-Windows-v1.1.1.zip. If you are on macOS or Linux, this route does not exist for you.

After extracting to a path without Chinese characters, the setup script prepares the runtime. The README notes that RTX 50 series owners should use a different script for CUDA 12.8 PyTorch.

bash
scripts/setup.bat

The script unpacks an embedded Python runtime, installs backend Python dependencies and frontend Node dependencies, generates configuration and creates the data directories. Then Ollama is installed separately from runtime\ollama\OllamaSetup.exe, and the README gives these two commands to run afterwards:

bash
ollama serve
ollama pull qwen2.5:1.5b

Model files are downloaded separately from a second Baidu Netdisk link and extracted into three directories: data\models\embedding\BAAI\bge-m3\ for embeddings, data\models\cn-clip\ for the visual model and data\models\faster-whisper\ for speech recognition. Finally scripts/startup.bat starts the backend and frontend. The README points to docs/部署文档/整合包部署指南.md for the detailed version.

The developer path is the one that works across platforms. It requires Python 3.10.11 or later and Node.js 21.x or later, and the README recommends 16GB of memory and an RTX3060 6GB or better.

bash
git clone https://github.com/dtsola/xiaoyaosearch.git
cd xiaoyaosearch
cd backend
pip install -r requirements.txt
pip install faster-whisper

CUDA is optional and version-dependent. The README warns that you must uninstall torch, torchaudio and torchvision first, then gives two pinned sets: CUDA 12.1 wheels for RTX 40 series and earlier, and CUDA 12.8 wheels for RTX 50 series. You also need ffmpeg and Ollama installed separately, from their own download pages. Configuration lives in backend/.env, and the README shows keys including FAISS_INDEX_PATH, WHOOSH_INDEX_PATH and DATABASE_PATH pointing at the directories under data/, plus API_HOST and an API port setting. The README excerpt does not show the full key list, so check the file itself before assuming a setting exists.

Where the design and the documentation leave you exposed

The first limitation is the licence. The repository reports NOASSERTION, and the README states that non-commercial use is free with modification and distribution allowed provided copyright and agreement notices are kept, while commercial use requires authorization under the separate 小遥搜索软件授权协议 in LICENSE (with LICENSE_EN alongside). That is not an OSI-approved open source licence, and the README itself calls the project non-commercial-free rather than open source. If you are evaluating this as a dependency inside a product, the licence is the blocking question, not the feature list.

The second is the install path. The recommended route for non-developers depends on two Baidu Netdisk downloads and a set of Windows batch scripts. There is no documented package manager install, no container image and no signed installer mentioned in the README. The README also does not document rollback, uninstall or upgrade procedures, and the only versioning evidence is the release list, which shows v2.0.0 on 2026-04-24, v1.9.0 on 2026-04-12 and tag-1.8.0 on 2026-04-08. Three releases inside a month suggests fast movement, and fast movement without a documented upgrade path means re-indexing is a real possibility you should plan for.

Third, the repository layout puts indexes and models under data/, and the README does not state whether an index built with one embedding model remains valid after you switch to a cloud embedding API. Since v1.6.0 advertises cloud embedding support and v1.3.0 advertises cloud model switching, mixing local and cloud embeddings against a single Faiss index is a plausible way to get bad results. The README is silent on this, so treat it as unverified rather than safe.

Finally, the performance numbers in the README (a 60 percent recall improvement for the terminology library in v1.9.0, a 67 percent visual comfort improvement in v2.0.0) are not accompanied by a described methodology. They are the project's own claims, not independently reproducible results.

How it compares with a general-purpose desktop indexer

The obvious alternative is a mature desktop content indexer such as Recoll, which builds full-text indexes over local files and exposes a query language and a GUI. The difference in approach is stark. Recoll is text extraction plus inverted index: it reads what is inside a file and matches query terms, with no embedding model and no semantic matching. XiaoyaoSearch adds a vector layer through Faiss and BGE-M3 on top of Whoosh, so a query can match a document that never contains your words. That is the entire reason to consider it.

The cost of that approach is weight. Recoll runs on modest hardware with no GPU and no model downloads. XiaoyaoSearch asks for 16GB of memory, a discrete GPU for reasonable performance, plus separate downloads of BGE-M3, CN-CLIP and FasterWhisper, and an Ollama runtime with a pulled model. If your searches are usually for a filename or a phrase you remember exactly, the vector layer earns nothing and you have paid for it in setup time and disk space.

The second comparison is against wiring a local LLM or an MCP client to a filesystem tool yourself. XiaoyaoSearch already ships an MCP server (v1.4.0) so Claude Desktop can query local files, and Agent Skills support (v1.5.0) for Claude Code, VS Code and Cursor. If your actual need is an AI assistant that can search your disk, those integrations may be the part you want, and the desktop GUI may be incidental. The README does not document the MCP tool surface in the excerpt, so the integration's real capability is something you would need to check in the source.

Maintenance, licensing and what an upgrade actually costs

The repository is not archived, and the last push was on 2026-05-22. That is four months before today, which is recent enough that the project is not dormant, but it is not a signal of continuous activity either. The release history is the more useful evidence: v2.0.0 on 2026-04-24, v1.9.0 on 2026-04-12 and tag-1.8.0 on 2026-04-08. The 2.0.0 release was framed around a Notion-style visual redesign with a design system, system font stack, Lucide Icons and layered shadows. A major version number spent on visual work, rather than on the retrieval pipeline, tells you where the maintainer's attention went in that cycle.

The upgrade cost is dominated by the models and the index. BGE-M3, CN-CLIP and FasterWhisper are downloaded manually into data/models/, and Ollama models are pulled separately with ollama pull. Any change to the embedding model or to the embedding provider invalidates a Faiss index built with the previous one, because vectors from different models are not comparable. The README does not document a re-index-on-model-change behaviour, so budget for rebuilding data/indexes/faiss/ yourself when you switch between local and cloud embeddings. The Whoosh index under data/indexes/whoosh/ is text-based and less exposed to that problem.

On licensing, the README states non-commercial use is free with modification and distribution permitted if copyright and agreement notices are retained, and that commercial use requires authorization under the agreement in LICENSE and LICENSE_EN. The repository metadata reports NOASSERTION because the licence is a custom agreement rather than a standard identifier. If you intend to redistribute a modified version, read the agreement text itself; this is a description of what the README says, not legal advice.

Editorial conclusion

Adopt XiaoyaoSearch if you are a Windows user who wants semantic search over a personal document, audio and video collection without sending files to a cloud service, and you accept that the free tier is non-commercial. Do not adopt it if you need a permissively licensed component in a commercial product, or if you want a documented rollback and upgrade path: the README does not document either. Verify three things first: the exact terms in LICENSE and LICENSE_EN, whether the model files at data/models/embedding/BAAI/bge-m3/, data/models/cn-clip/ and data/models/faster-whisper/ are present and complete, and whether your hardware matches the stated 16GB memory and RTX3060 6GB guidance. If your only machine is macOS or Linux, start from the developer deployment path, not the all-in-one package, which the README restricts to Windows.

Frequently asked questions

Does XiaoyaoSearch send my files or queries to the cloud?

By default it runs locally and the README states data is not uploaded. Since v1.3.0 and v1.6.0 it also supports cloud model and cloud embedding APIs from OpenAI, DeepSeek and Aliyun, which you switch to deliberately. The README frames this as a privacy and performance trade-off the user chooses.

Can I install XiaoyaoSearch on macOS or Linux?

The all-in-one package route is documented as Windows only, distributed through Baidu Netdisk. The developer deployment path supports Windows, macOS and Linux and requires Python 3.10.11 or later plus Node.js 21.x or later. CUDA wheels are optional and version-dependent.

What hardware does XiaoyaoSearch need?

The README recommends 16GB of memory or more and an RTX3060 with 6GB or better. It also gives separate PyTorch install commands for RTX 40 series and earlier on CUDA 12.1 and for RTX 50 series on CUDA 12.8. CPU-only installs are the default from requirements.txt.

Can I use XiaoyaoSearch commercially?

The README states non-commercial use is free with modification and distribution allowed if copyright and agreement notices are kept, and that commercial use requires authorization. The terms live in LICENSE and LICENSE_EN, and the repository reports the licence as NOASSERTION rather than a standard identifier.

Official sources

  1. dtsola/xiaoyaosearch on GitHub
  2. Issues
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/dtsola-xiaoyaosearch.svg)](https://hysenlabs.com/projects/dtsola-xiaoyaosearch)