# swirlai/swirl-search: federated search and RAG that queries your sources in place

> SWIRL Community is an Apache-2.0 Django application that federates queries across more than a hundred connectors, re-ranks the results with cosine similarity, and can generate a cited answer with your own OpenAI key. It is for teams that will not stand up a vector database just to search their own tools.

**swirlai/swirl-search** — AI Search & RAG Without Moving Your Data. Get instant answers from your company's knowledge across 100+ apps while keeping data secure. Deploy in minutes, not months.

- Repository: https://github.com/swirlai/swirl-search
- Website: https://swirlaiconnect.com/
- Stars: 3,046 · Forks: 286
- Language: Python
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/swirlai-swirl-search

## The copy problem SWIRL Community refuses to accept

The conventional path to AI search starts with duplication. You stand up a vector database, build pipelines that pull content out of SharePoint, Confluence, Jira and the rest, embed it, and then maintain that second copy forever. Every permission change in the source system becomes a governance question about the copy. SWIRL takes the opposite position: the query goes to the sources, not the sources to the query. The README frames the contrast directly, describing the usual way as standing up a vector database, moving and duplicating data, and building ETL pipelines, against querying live and in place. The intended user is an engineering or platform team inside a company that already has knowledge spread across many SaaS tools and does not want a new store to secure and audit. It is not a general web search engine and it is not a scraping framework. The pitch is that permissions are enforced at the source, because the source is what answers.

## How federation, re-ranking and RAG fit together

SWIRL is a Django application, and the repository layout makes the shape clear: a swirl package holding the core, a swirl_server package for the ASGI application, and top-level scripts named swirl.py, swirl_load.py and swirl_sync.py for setup, loading and synchronisation. Connectors live under swirl/connectors, and the README describes them as extensible Connector and Mixer objects. Mixers combine result sets from different sources; a pipelined Processor architecture transforms queries, responses and results. Federation can run synchronously or asynchronously over a REST API. Ranking at the Community tier is cosine vector similarity using the spaCy large model plus NLTK, with duplicate detection and result mixers on top. Results are stored in SQLite or Postgres for post-processing and analytics. The RAG step is optional and uses the LLM of your choice; with an OpenAI key set, the server generates an answer with citations. Two environment variables in the compose file bound that step: SWIRL_RAG_TOK_MAX defaults to 4000 and SWIRL_RAG_MAX_TO_CONSIDER defaults to 10. Those defaults are worth reading as a statement about cost and context, not just configuration.

## Installing SWIRL Community with docker-compose

The README's quick start assumes the Docker app is installed and running. Fetch the compose file from the repository, then optionally export an OpenAI key so real-time RAG is enabled. The compose file passes OPENAI_API_KEY through to the app container, along with Azure equivalents and the SWIRL_RAG_MODEL, SWIRL_REWRITE_MODEL and SWIRL_QUERY_MODEL variables.

```bash
curl https://raw.githubusercontent.com/swirlai/swirl-search/main/docker-compose.yaml -o docker-compose.yaml
```

```bash
export MSAL_CB_PORT=8000
export MSAL_HOST=localhost
export OPENAI_API_KEY='<your-OpenAI-API-key>'
```

Pull and start the stack. The compose file defines two services: redis on port 6379 with a healthcheck, and app on port 8000, which waits for redis to report healthy before starting.

```bash
docker-compose pull && docker-compose up
```

Then open http://localhost:8000/galaxy and log in with admin / password. The README states SWIRL comes ready to search Arxiv, European PMC and Google News out of the box, so a first search works without configuring a connector. The app container's command runs python swirl.py setup, writes a default API config, starts a celery worker and celery beat, and serves the ASGI application with daphne on 0.0.0.0:8000. The Dockerfile installs the spaCy model en_core_web_lg and downloads NLTK resources at build time, which is why the image is not small. One warning sits directly under the quick start: the Docker version does not retain data or configuration when shut down. For anything you intend to keep, the README points to the Quick Start Guide for a persistent install.

## Where Community stops and Enterprise begins

The README is unusually explicit about the boundary, which makes the limitation easy to state. Community gives you federated search and RAG, the Galaxy UI, the connector set, RAG with your own LLM key, and cosine-similarity ranking. It does not give you the three-pass reranker, which the comparison table describes as BM25 first, then E5 embeddings with hybrid fusion, then a cross-encoder. It does not give you canonical answers or Pinned Results, a first-class MCP server for agents, a hallucination warning on generated answers, or the business console with AI-Yield analytics, semantic cache and dedup at scale. Those are listed as Enterprise capabilities only. That is a real ceiling, not a marketing footnote. If your queries are short and your corpus is large, cosine similarity over spaCy and NLTK may rank worse than the hybrid pipeline, and there is no configuration switch in Community that closes the gap. The other honest constraint is operational: the compose file ships with the admin / password default, and the README documents no rollback procedure for a bad upgrade. Plan around both.

## Alternatives and the difference in approach

The obvious alternative is the vector-database stack: pull content into something like a dedicated embedding store, embed it, and query the store. That approach buys you fast approximate nearest-neighbour retrieval and a single ranking model you control, at the cost of an ingestion pipeline and a governed copy of the data. SWIRL's difference is that the copy never exists. Its retrieval quality is bounded by what each source's API returns and by cosine similarity, but its freshness and permission story are inherited from the sources themselves. A second alternative is writing your own federation layer over the handful of APIs you actually use. That is genuinely tractable for two or three sources, and it avoids running Django, Celery, Redis and Daphne. It stops being tractable when you want the Galaxy UI, duplicate detection, result mixers and a processor pipeline, because at that point you are rebuilding SWIRL. The third comparison is internal: Community against SWIRL 5 Enterprise. Same federation model, better ranking and extra answer features. The README says plainly that plenty of teams run Community in production and never need more.

## Licence, upgrade cost and what the repository tells you

SWIRL Community is Apache-2.0, and the repository carries both a LICENSE and a NOTICE file, which is the usual arrangement for Apache-2.0 distribution: keep the notices intact when you redistribute. That is a description of the files present, not legal advice, and the Enterprise tier is a separate commercial product with terms the README does not reproduce. On upgrade cost, the repository shows the shape of the work. requirements.txt pins exact versions across a wide dependency surface, including Django 5.2.14, djangorestframework 3.17.1, celery 5.6.3, channels 4.3.2, elasticsearch 8.19.3, google-cloud-bigquery 3.41.0, litellm 1.83.14 and msal. The Dockerfile derives the spaCy version by grepping that file, installs the en_core_web_lg model, and runs download-nltk-resources.sh. Upgrading therefore means moving a pinned set, not floating on latest. The release history shows the cadence: v4.5.0.5, v4.5.0.6 and v4.5.0.7 landed between 2026-05-21 and 2026-06-22, and the last push to the default branch was on 2026-09-05. There is no documented rollback path in the README.

## Conclusion

Adopt SWIRL Community if your searchable content already lives in systems with APIs and you want ranked, cited answers without building an ETL pipeline or a second copy of the data. Do not adopt it if you need the three-pass reranker, canonical answers or an MCP server: the README places those in SWIRL 5 Enterprise, not here. Before you commit, verify three things on your own machine. First, whether every source you care about has a connector, since the connector list is not in the repository. Second, whether the Docker install's non-persistent storage matters to you, because the README states the Docker version does not retain data or configuration when shut down. Third, whether cosine similarity with spaCy and NLTK ranks your content well enough, because that is the ranking you get at this tier.

## FAQ

### Can you trust AI search?

SWIRL's answer to this is architectural rather than statistical: it returns sources you can click through to, and the README states data stays where it lives, with permissions enforced at the source. A hallucination warning on generated answers is listed as an Enterprise capability, not a Community one, so at this tier the citation links are the check you have.

### What is an AI search called?

SWIRL describes what it does as federated AI search and RAG, and the repository topics include federated-search, metasearch, unified-search and retrieval-augmented-generation. It queries sources live rather than indexing a copy, which is the distinction the README draws against the usual vector-database approach.

### What does the swirl-search repository contain?

It is a Python Django application with a swirl package for the core, a swirl_server package for the ASGI server, top-level scripts including swirl.py, swirl_load.py and swirl_sync.py, a docker-compose.yaml, a Dockerfile, and an Apache-2.0 LICENSE and NOTICE. Connectors sit under swirl/connectors and AI provider code under AIProviders/.

### Does SWIRL Community keep my data between restarts?

Not in the Docker quick start. The README states the Docker version does not retain data or configuration when shut down, and points to the Quick Start Guide for a persistent install.

### Which sources can SWIRL search out of the box?

The README states SWIRL comes ready to search Arxiv, European PMC and Google News out of the box, and that the full connector list lives at swirlaiconnect.com/connectors rather than in the repository.

## Sources

- [License: Apache-2.0](https://github.com/swirlai/swirl-search/blob/main/LICENSE)
- [Project website](https://swirlaiconnect.com/)
- [README](https://github.com/swirlai/swirl-search/blob/main/README.md)
- [Releases](https://github.com/swirlai/swirl-search/releases)
- [swirlai/swirl-search on GitHub](https://github.com/swirlai/swirl-search)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/swirlai-swirl-search
