Model or dataset
watercrawl/WaterCrawl avatar
watercrawl/WaterCrawl

WaterCrawl: Self-Hosted Web Crawling and Scraping for LLM-Ready Data

Transform Web Content into LLM-Ready Data

2,177 stars277 forksTypeScriptNOASSERTION

At a glance

What is it?
WaterCrawl is a self-hosted web application built on Python, Django, Scrapy, and Celery that crawls and scrapes websites, exposes a REST API with real-time Server-Sent Events progress, and integrates with Dify, N8N, and other AI automation platforms to turn web content into data ready for LLM pipelines.
Who is it for?
WaterCrawl suits teams that need a self-hosted, API-driven web crawler for feeding LLM pipelines and want Python, Node.js, Go, or PHP client SDKs and ready integrations for Dify and N8N. It is not the right tool for teams that need to crawl websites without self-managing infrastructure, or for one-off research scraping where a local script would do.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Yes. The repository last received commits 46 days ago.
What is it written in?
Mainly TypeScript, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 2, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What WaterCrawl Solves and Who It Is For

Feeding web content to an LLM pipeline requires crawling pages, extracting relevant text, and delivering that text in a format the LLM can process. Building this pipeline from scratch involves crawler logic, queue management, result storage, and an API layer. WaterCrawl provides all of these as a deployed web application.

The target users are developers and teams who want a self-hosted, API-accessible crawler rather than a cloud scraping service, with control over data retention, rate limiting, and integration with their own AI infrastructure. The README describes WaterCrawl as transforming web content into LLM-ready data, and the feature list includes multi-language support with country-specific targeting, asynchronous processing with real-time SSE progress, and a REST API with OpenAPI documentation.

WaterCrawl also includes a search capability: the README lists basic, advanced, and ultimate search depths, along with language and country targeting. This makes it useful not just as a site crawler but as a research and content collection tool for AI applications.

Getting Started: Docker Deployment

The README documents a Docker-based quick start:

bash
git clone https://github.com/watercrawl/watercrawl.git
cd watercrawl

Then build and run the containers:

bash
cd docker
cp .env.example .env
docker compose up -d

Access the application at http://localhost after the containers start. The README adds one important deployment note: if deploying on a domain or IP address other than localhost, three MinIO configuration values in the .env file must be updated to the actual domain:

bash
MINIO_EXTERNAL_ENDPOINT=your-domain.com
MINIO_BROWSER_REDIRECT_URL=http://your-domain.com/minio-console/
MINIO_SERVER_URL=http://your-domain.com/

The README marks this as important and states that failure to update these values results in broken file uploads and downloads. A DEPLOYMENT.md file in the repository root provides additional production deployment guidance. Development setup and contribution instructions are in CONTRIBUTING.md.

REST API, OpenAPI Documentation, and Real-Time Progress

WaterCrawl exposes a REST API with comprehensive OpenAPI documentation. The README lists the API Overview at docs.watercrawl.dev/intro as the reference. Crawl jobs and search requests are submitted via API, and progress is reported through Server-Sent Events (SSE), allowing clients to monitor crawl status in real time without polling.

The API design enables programmatic control over depth, speed, and content targeting. The advanced and ultimate search depth options are listed without further specification in the README; the API documentation at docs.watercrawl.dev covers the full parameter set. The OpenAPI spec also supports the client SDK generation that produced the four official client libraries.

Client SDKs and Integrations

WaterCrawl ships four official client SDKs: Python (full-featured, all endpoints), Node.js (complete JavaScript and TypeScript integration), Go (full-featured), and PHP (full-featured). A Rust client is listed as coming soon. Each SDK is documented at docs.watercrawl.dev with its own section.

The integration list at launch includes a Dify plugin (available on the Dify marketplace), an N8N workflow node (published on npm as @watercrawl/n8n-nodes-watercrawl), a Dify Knowledge Base connection, a WaterCrawl plugin, and an OpenAI plugin. Langflow integration is listed as a pull request not yet merged. Flowise is listed as coming soon.

These integrations position WaterCrawl as a data source component in no-code and low-code AI automation pipelines, not only as a tool for custom application development.

Repository Layout and Development Setup

The repository is a monorepo. The backend/ directory holds the Django and Scrapy code; the frontend/ directory holds the web interface. The docker/ directory contains Docker Compose configuration and environment templates. A Makefile at the root provides development targets:

bash
make install          # Install all dependencies
make lint             # Run all linters
make install-backend  # Install backend via Poetry
make install-frontend # Install frontend via pnpm

The backend uses Ruff for linting and formatting; the frontend uses ESLint. Pre-commit hooks are configured via .pre-commit-config.yaml. The README notes that frontend linting is handled separately from pre-commit due to SSL certificate issues. A DEV_SETUP.md file provides additional development environment instructions.

Limitations and Trade-offs

WaterCrawl is a self-hosted application, which means teams are responsible for infrastructure: server provisioning, Docker and MinIO management, database maintenance, and keeping the .env configuration correct across environments. Teams without DevOps capacity may find the operational overhead higher than using a managed scraping API.

The license is a custom WaterCrawl License, described in the README as "essentially MIT with a few additional restrictions." Teams that require a standard open-source license for compliance or procurement reasons should review the LICENSE file before committing. The README does not enumerate the specific additional restrictions.

The repository has no GitHub releases. The last push was on 2026-08-17. The Makefile, pre-commit setup, and CI workflows indicate an actively maintained codebase, but the absence of tagged releases means there is no stable version to pin to. Teams deploying WaterCrawl in production should document the commit hash they deploy from.

WaterCrawl Versus Scrapy as a Standalone Framework

Scrapy is the Python web scraping framework that WaterCrawl uses as its crawler engine. Scrapy alone is a library: it requires writing Python spiders, managing a Scrapy project, and building the surrounding infrastructure (storage, queuing, API access) yourself. It gives full control over crawl logic at the cost of writing and maintaining more code.

WaterCrawl wraps Scrapy in a deployed application with a database, a task queue (Celery), object storage (MinIO), a REST API, and a web interface. Teams that need a turnkey crawling service rather than a crawler framework will find WaterCrawl reduces the integration work substantially. Teams that need very custom crawling behavior, unusual spider logic, or tight integration with an existing Python codebase may find the framework layer easier to work with directly.

Editorial conclusion

WaterCrawl suits teams that need a self-hosted, API-driven web crawler for feeding LLM pipelines and want Python, Node.js, Go, or PHP client SDKs and ready integrations for Dify and N8N. It is not the right tool for teams that need to crawl websites without self-managing infrastructure, or for one-off research scraping where a local script would do. Before deploying on a domain or IP other than localhost, update MINIO_EXTERNAL_ENDPOINT in the .env file and the two MINIO_BROWSER_REDIRECT_URL and MINIO_SERVER_URL values; the README marks this as important because failure to do so breaks file uploads and downloads. The repository last received a push on 2026-08-17.

Frequently asked questions

What is WaterCrawl?

WaterCrawl is a self-hosted web application that crawls websites, extracts content, and delivers it via a REST API for use in LLM pipelines and AI automation platforms. It uses Python, Django, Scrapy, and Celery, and deploys via Docker.

What output format does WaterCrawl produce?

The README describes WaterCrawl as transforming web content into LLM-ready data. The REST API delivers results, and the SDKs (Python, Node.js, Go, PHP) provide typed access to crawl results. The API documentation at docs.watercrawl.dev covers the full response schema.

Can WaterCrawl be deployed on a custom domain?

Yes, but the README requires updating three MinIO configuration values in the .env file: MINIO_EXTERNAL_ENDPOINT, MINIO_BROWSER_REDIRECT_URL, and MINIO_SERVER_URL. Leaving them at their localhost defaults breaks file uploads and downloads on a custom domain or IP.

Official sources

  1. Issues
  2. Project website
  3. README
  4. watercrawl/WaterCrawl on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/watercrawl-watercrawl.svg)](https://hysenlabs.com/projects/watercrawl-watercrawl)