YaCy: a search engine you run, either solo or as a P2P network
Distributed Peer-to-Peer Web Search Engine and Intranet Search Appliance
At a glance
- What is it?
- Java, Apache Ant and a crawler, an index and a web UI in one application, with a peer to peer mode that shares indexes across independent installations.
- Who is it for?
- YaCy is unusual in offering two products from one codebase: a private intranet or personal search appliance that behaves like conventional server software, and a P2P network where independent installations contribute index partitions to shared queries. Choosing between them is the real decision, because the default mode puts your node in a public network and a robotics bot on an internal network is a different deployment entirely.
- Can I use it commercially?
- Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
- Is it still maintained?
- Yes. The repository last received commits 16 days ago.
- What is it written in?
- Mainly Java, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 9, 2026, and from our analysis. They are not legal advice.
Editorial analysis
One application containing a server, a UI and a crawler
The README describes YaCy as a full search engine application containing a server hosting a search index, a web application providing a user front-end for searches and index creation, and a production-ready web crawler with a scheduler to keep that index fresh.
Those three parts in one process is the architectural fact that shapes everything else. There is no separate Solr to run, no message queue, and no Elasticsearch cluster to provision. A crawler, an index, a query engine and an HTTP interface live together, started by a shell script.
The positioning is deliberately broad. The README says YaCy search portals can be placed in an intranet environment, making it a replacement for commercial enterprise search solutions, and it mentions a network scanner for discovering available HTTP, FTP and SMB servers. It also notes you can use YaCy with a customized search page in your own web applications, which is the feature that turns it from a search box into a component of something else. Privacy is named as the founding motivation rather than a feature added later.
The default mode joins you to a public P2P search network
This is the detail that catches people out. Each YaCy peer can join a large search network where indexes are exchanged with other installations over a built-in peer to peer protocol, and the README describes this as the default operation that gives new users instant access to a large scale search cluster operated only by YaCy users.
It also describes the two ways out. You can opt out of cluster operation by choosing a different operation mode in the web interface, and you can opt out of the network on individual searches, which turns YaCy into a fully privacy-aware tool computing results from the local index only.
Those two levels of opt-out matter operationally. A standalone node behaves like any other intranet search appliance and can crawl internal hosts without ever asking a peer. A node left on the default setting will contribute its index to shared queries, which is exactly what makes the P2P mode worth using and exactly what you need to decide deliberately rather than by accepting a default.
The repository is a Java project whose topics describe it as decentralized, intranet search, p2p and privacy, with a home page at yacy.net and no license declared automatically.
Building from git with Apache Ant and Java 17
The README recommends compiling from git sources, and states plainly that pre-compiled packages exist but are not generated on a regular basis. On Debian or Ubuntu the prerequisites are one command:
sudo apt-get install openjdk-17-jdk-headless antThen clone and build:
git clone --depth 1 https://github.com/yacy/yacy_search_server.git
cd yacy_search_server
ant clean allJava 17 or later and ant are the stated requirements, and the repository supports that story with `build.xml`, an `ivy.xml` for dependency resolution, a `.classpath` and `.project` for Eclipse, and a Nix flake as `flake.nix` and `flake.lock`. The `--depth 1` in the clone is worth keeping, since the history of a long running search engine project is large.
Starting and stopping are script calls rather than service commands:
./startYACY.sh
./stopYACY.shWindows equivalents sit alongside as `startYACY.bat` and `stopYACY.bat`, plus `installYaCyWindowsService.bat` for registering it as a service and `getWin32MaxHeap.bat` for the heap size the JVM needs. The administration interface is on port 8090, with 8443 for TLS in the Docker image.
Default credentials and what happens on first boot
The README documents the first-run security posture in more detail than most projects do, and part of what it describes is a convenience you should close immediately.
The administration interface is at `http://localhost:8090`. Some web pages are protected and need an administration account. Those pages are usually also available without a password from localhost, but remote access needs a log-in. The default admin account name is `admin` and the default password is `yacy`, and the README instructs you to change it after installation using the `http://<server-address>:8090/ConfigAccounts_p.html` service.
That localhost exception is worth thinking about before you deploy. In a container or behind a reverse proxy, what counts as localhost changes, and a configuration page that trusts the loopback interface can be reachable in ways the operator did not intend. Changing the password via `ConfigAccounts_p.html` is the documented first step.
The Docker run line in the README is the packaged alternative, and it shows the shape of a real deployment:
docker run -d --name yacy_search_server -p 8090:8090 -p 8443:8443 -v yacy_search_server_data:/opt/yacy_search_server/DATA --restart unless-stopped --log-opt max-size=200m --log-opt max-file=2 yacy/yacy_search_server:latestThe named volume for `DATA`, the log rotation flags and the restart policy are all specified, which is more deployment detail than most READMEs provide.
Every web page is also an API endpoint
The interface story is unusual and elegant: YaCy exposes both HTTP/XML and HTTP/JSON versions of its own web pages. The README says you discover them by noticing the orange API icon in the upper right corner of some pages, clicking it to see the XML or JSON version of the page you are looking at.
That is worth internalizing before looking for API documentation, because it means the query interface, the index browser and the crawler controls are all scriptable through the same URLs a human uses. The repository reinforces it with a `bin/` directory of shell scripts that call the YaCy web interface, and the README suggests cloning those scripts as a starting point for building your own shell API access.
Windows installers are a separate build path using NSIS, which requires the release payload that Ant produces. The README gives the single command and the two step alternative:
ant copyMain4Dist
makensis RELEASE/WINDOWS/build.nsiThe two step form stages files into `RELEASE/MAIN` first and runs `makensis` against `RELEASE/WINDOWS/build.nsi`, producing `yacy_v<version>_*.exe` in `RELEASE/`.
Licensing that is not one license
The repository's license field carries the value NOASSERTION, which means nothing was detected automatically. The README tells a more nuanced story: the project is open source under the GPL 2.0 or later, with some elements licensed under the GNU Lesser General Public License, and it tells you to check individual files for accurate information about licenses and copyrights.
The tree backs that up rather than contradicting it. There is a `gpl.txt`, an `lgpl21.txt`, a `COPYRIGHT` file and a `LICENSES/` directory, which is a consistent layout for a codebase that genuinely mixes copyleft terms. The README also states that the GPLv2+ source code used to build YaCy is distributed with the package, in the `/source` and `/htroot` directories.
If you plan to redistribute or embed parts of it, that per-file instruction is the operative one, and the presence of LGPL elements means some components carry obligations the GPL alone does not imply.
Two deployment targets beyond the README are visible in the tree and may matter for a container platform: `kubernetes/` with charts in `charts/`, an `app.json` for Heroku, a `Procfile`, a `Heroku.md`, a `vagrant_yacy/` directory, and `snap/` packaging.
Editorial conclusion
YaCy is unusual in offering two products from one codebase: a private intranet or personal search appliance that behaves like conventional server software, and a P2P network where independent installations contribute index partitions to shared queries. Choosing between them is the real decision, because the default mode puts your node in a public network and a robotics bot on an internal network is a different deployment entirely. Build it from git with `ant clean all` on Java 17 rather than waiting for a release, since the README is candid that packaged builds are not produced regularly and the Docker image is the recommended production path. The HTTP and JSON interfaces are discoverable from the orange API icon in the web UI, which makes it scriptable without reading any documentation. The last push was on 2026-09-23, and the licensing is GPL 2.0 or later with some LGPL elements, which is worth checking per file if you plan to redistribute.
Frequently asked questions
How is YaCy different from Google?
Google is a global crawler you query over a network. YaCy is software you run yourself, containing its own crawler, index and web interface, so the same thing can serve as a private intranet search appliance. In its default mode it also joins a peer to peer network that exchanges indexes with other YaCy installations, which is the closest thing it has to a shared corpus. You can opt out of that network entirely and search the local index only.
What are the pros and cons of YaCy?
The strengths are that it needs no external search service, that the crawler, index and UI live in one Java application, and that every web page doubles as an XML or JSON endpoint. The costs are operational: a real crawl needs time and disk, indexing quality depends on configuration, and by default the node participates in a public P2P network unless you change the operation mode. Builds are also less polished than a packaged product, since the README says pre-compiled packages are not generated regularly.
How do I build and run YaCy from the git sources?
Install Java 17 or later plus ant, then clone the repository and build it. On Debian that is sudo apt-get install openjdk-17-jdk-headless ant, followed by git clone --depth 1 https://github.com/yacy/yacy_search_server.git, cd yacy_search_server and ant clean all. Start it with ./startYACY.sh and open http://localhost:8090. Stop it with ./stopYACY.sh.
What is the default admin password for YaCy?
The account name is admin and the password is yacy, and the README is explicit that you should change it after installation using the http://<server-address>:8090/ConfigAccounts_p.html service. It also notes that some protected pages are reachable without a password from localhost, while remote access requires a log-in, which matters when you deploy behind a proxy or in a container where the loopback interface is reachable more widely than expected.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/yacy-yacy-search-server)