Model or dataset
codelibs/fess avatar
codelibs/fess

Fess: a self-hosted search server that crawls your sites and stores

Open-source, self-hosted enterprise & site search server built on OpenSearch. Crawls web / file / DB / cloud sources, 20+ languages, REST API, and AI/RAG & semantic search. Apache-2.0.

1,138 stars175 forksJavaApache-2.0

At a glance

What is it?
Fess is an Apache-2.0 Java search server built on OpenSearch, with a browser admin UI and crawlers for web, file and data store sources. Here is what it installs like, how the crawler pipeline is wired, and where it stops being the right tool.
Who is it for?
Fess fits teams that need one search box over a web site, a file share and a database, and that can run an OpenSearch backend and a Java 21 runtime. It does not fit anyone who wants a hosted service with no infrastructure, or a search box for a single static site, where Fess Site Search is the lighter path.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository received new commits within the last day.
What is it written in?
Mainly Java, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem Fess solves: one index over sources that do not talk to each other

Most organizations do not have a search problem. They have a fragmentation problem. The wiki, the shared drive, the ticketing system and a handful of internal databases each have their own search, each with different syntax and different permissions, and none of them knows what the others contain. Fess targets exactly that gap: it is a self-hosted server that crawls web sites, file systems and data stores into a single OpenSearch index and exposes one query interface over the result.

The intended user is an operations or platform engineer who can run a Java service and an OpenSearch cluster, not an end user who wants to sign up for something. The README is explicit that prior OpenSearch knowledge is not required because configuration happens through a browser-based administration UI, but someone still has to install and supervise the backend. The feature list names role- and permission-based filtering of results, single sign-on through LDAP, OpenID Connect, SAML, SPNEGO and Microsoft Entra ID, and text analysis for 20 or more languages. Those are the features that separate it from a static site search box: access control and multilingual analysis are the parts that are expensive to build yourself.

How the crawler, index and query path fit together

The architecture is a Java application in front of OpenSearch. Fess itself owns crawling, document parsing and the query UI; OpenSearch owns the inverted index and scoring. That split matters when you plan capacity, because the search backend scales and fails independently of the crawler.

Crawling targets are not configured in a file. You register them on the Web, File or Data Store configuration pages in the Admin UI, then start the crawler from the Scheduler page. Each target type maps to a different collection path: web configuration for sites, file configuration for file systems, data configuration for databases, CSV files and the connectors listed in the README. Documents are parsed on the way in, with Microsoft Office, PDF and ZIP archives named as supported formats, and the extracted content is indexed into OpenSearch.

Data store support is where the repository layout becomes important. The connectors for Confluence and Jira, Box, CSV, databases, Dropbox, Elasticsearch, Git, Gitbucket, G Suite, JSON, Office 365, S3, Salesforce, SharePoint and Slack live in separate repositories under the codelibs org, each named fess-ds-something. They are not part of the core tree. If your source is not in that list, you are writing a connector, not configuring one. The same pattern applies to themes, ingest processors and script engines, which are also separate plugin repositories.

Fess installation: ZIP package, first crawl, first query

The README gives three download formats, DEB, RPM and ZIP, all from the Releases page, and states that Java 21 or later is required for the packages. The ZIP route is the one documented with commands. Unzip the release, change into the directory, and start the bundled launcher:

bash
unzip fess-<version>.zip
cd fess-<version>
./bin/fess

The server then serves two interfaces on port 8080: the search UI at http://localhost:8080/ and the admin UI at http://localhost:8080/admin/. The README states the default admin username and password are admin/admin. Change that before the service is reachable from anywhere but your own machine; the documentation does not describe a forced password change on first login.

With the server up, the first real use is registering a crawl target and running it. In the Admin UI, open the Web configuration page, add the site you want indexed, save, then go to the Scheduler page and start the crawler. When the crawl finishes, the target's content is queryable from the search UI. For scripted checks, the README documents a health endpoint that returns JSON when the server is ready:

bash
curl -s "http://localhost:8080/api/v1/health"

The README notes the server may take up to 60 seconds to become ready, so a failing health check immediately after startup is expected rather than a fault. If you prefer containers, Docker images are published on ghcr.io and a Compose file is available in the docker-fess repository; the README says the Docker images bundle OpenSearch, while other installation methods require you to set it up separately.

Where Fess is the wrong tool

The clearest limitation is the backend dependency. Fess is not a single binary that manages its own storage. OpenSearch is required, bundled only in the Docker images, and separate for every other installation method. That means a second service to size, back up and upgrade, plus the version compatibility question between Fess and OpenSearch, which the README defers to the Installation Guide rather than stating inline. If you wanted a search box and nothing else, this is more moving parts than the problem deserves.

Crawl-based indexing is the second constraint. Content becomes searchable after a crawl, not when it changes. The scheduler controls when that happens, and the README does not describe incremental change detection or a push-based indexing path for the web crawler. For a site that publishes continuously, the freshness of results is a scheduling decision you own.

Connector gaps are the third. The Data Store section is a list of external repositories, so support for a given SaaS product depends on whether that plugin exists and is maintained separately from the core server. The README does not document a fallback for unsupported sources beyond writing a plugin. Finally, the README does not document rollback for an upgrade, so plan the OpenSearch snapshot side yourself before moving versions.

Fess Site Search versus the full server

The most useful alternative is not another search server, it is the project's own smaller component. Fess Site Search is described in the README as a free alternative to Google Site Search that you embed in your own website, distributed through a JavaScript generator with its own manual. The difference in approach is fundamental: Fess Site Search puts a search interface on a site that already has content, while the Fess server crawls and indexes sources you do not control, applies role-based filtering and runs behind SSO.

If your only requirement is a search box over pages you publish, Fess Site Search removes the OpenSearch cluster, the Java runtime and the crawl schedule from the picture. If your requirement includes a file share, a database, or per-user result filtering, the full server is the only one of the two that addresses it. Choosing the heavier component for a single static site is the most common way to take on maintenance you never needed.

Maintenance, releases and the Apache-2.0 terms

The repository is not archived, and the last push was on 2026-09-10. The release cadence visible in the repository is roughly every six to eight weeks: fess-15.6.1 on 2026-05-02, fess-15.7.0 on 2026-06-25, fess-15.8.0 on 2026-08-20. That cadence is the upgrade cost you are signing up for, and it interacts with the OpenSearch version question, since a Fess upgrade may move the backend version with it. The README does not document a rollback procedure.

Building from source is a separate commitment. It requires Java 21 or later and Maven, plus an antrun step that downloads OpenSearch plugins into the plugins directory. The README also documents a code generation step involving dbflute and a license formatting goal, which tells you the generated layer is expected to be regenerated rather than hand-edited. Integration tests need a running Fess server, a running OpenSearch instance, and a clone of the fess-testdata repository before they will pass.

Fess is Apache-2.0. That is a permissive licence, and the practical consequence is that the plugin repositories under the same org are separate deliverables with their own terms; check each connector you depend on rather than assuming the core licence covers the whole deployment. Nothing here is legal advice, and licence compatibility with your own distribution model is a question for whoever handles that at your organization.

Editorial conclusion

Fess fits teams that need one search box over a web site, a file share and a database, and that can run an OpenSearch backend and a Java 21 runtime. It does not fit anyone who wants a hosted service with no infrastructure, or a search box for a single static site, where Fess Site Search is the lighter path. Before committing, verify the OpenSearch version your package expects in the Installation Guide, confirm the admin/admin default credentials are changed, and check that the data store connector you need exists as a separate plugin on the codelibs org, because the core repository ships the crawler framework and not every connector.

Frequently asked questions

What are the requirements to install Fess?

The README states Java 21 or later is required for the ZIP, RPM and DEB packages, and that OpenSearch is the search engine backend. The Docker images bundle OpenSearch, but for other installation methods you set it up separately.

How do I start the Fess crawler?

Register the crawling target on the Web, File or Data Store configuration pages in the Admin UI, then start the crawler from the Scheduler page. The README describes this as the standard flow after installation.

What is the default Fess admin login?

The README states the admin UI at http://localhost:8080/admin/ uses admin/admin as the default username and password. It does not describe a forced password change on first login.

Does Fess include connectors for Confluence, S3 or Slack?

The README lists those sources under Data Store, but each connector lives in its own repository under the codelibs org, such as fess-ds-atlassian and fess-ds-s3, rather than in the core tree. Support for a given source depends on that separate plugin.

Official sources

  1. codelibs/fess on GitHub
  2. License: Apache-2.0
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/codelibs-fess.svg)](https://hysenlabs.com/projects/codelibs-fess)