Self-hosted service
crawlseo/crawlseo avatar
crawlseo/crawlseo

CrawlSEO: a self-hosted SEO dashboard whose Docker docs run ahead of its releases

Open-source SEO monitoring. GSC + Site Crawler + Core Web Vitals in one dashboard.

619 stars91 forksTypeScriptMIT

At a glance

What is it?
CrawlSEO bundles Google Search Console analytics, a site crawler, Core Web Vitals and an MCP server into one Next.js dashboard you host yourself. The mechanics worth reading are the ones the marketing table leaves out: image tags documented for a 1.2.x line that has no release, two names for one secret, and keyword data that needs a paid key.
Who is it for?
CrawlSEO fits a founder or small team that wants Search Console numbers, a crawl audit and an agent-accessible API on its own infrastructure, and who is prepared to supply Google OAuth credentials and a PostgreSQL database. Before deploying, pin the image deliberately, because the documented tag scheme refers to 1.2.x releases while the only tagged release is 0.1.0, and read the environment file rather than the README table, since the secret appears under more than one name.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 2 days ago.
What is it written in?
Mainly TypeScript, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 5, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The docs pin images to 1.2.x, the only release is 0.1.0

Versioning is where this project contradicts itself first. The single GitHub release is v0.1.0, published on 2026-08-05, and the manifest still carries version 0.1.0 with `private: true`. The Docker documentation, by contrast, tells you that version tags are published as immutable image tags such as `1.2.3` and as minor tags such as `1.2`, and then tells you to pin one with `CRAWLSEO_IMAGE`. Those two schemes cannot both be current: there is no 1.2.x anywhere in the release history, and a 0.x manifest is not going to produce 1.2 image tags. The compose file resolves the ambiguity in the simplest way possible, defaulting to `ghcr.io/crawlseo/crawlseo:latest` when `CRAWLSEO_IMAGE` is unset. Practically, that default means an unpinned deploy tracks whatever the registry currently serves, so if you want reproducibility you have to set the variable yourself, and pick a tag that the project's own docs have not invented.

One secret, three variable names, two env file names

Configuration is described in two places that do not agree. The quick start is six commands, and the second one already picks the env filename:

bash
git clone https://github.com/crawlseo/crawlseo.git
cd crawlseo
cp .env.example .env.local
# Add your Google OAuth credentials to .env.local
docker compose up -d db
npm install
npx prisma migrate dev --name init
npm run dev

Then the README's environment table lists five variables, with `NEXTAUTH_SECRET` marked required and generated by `openssl rand -hex 32`. The example env file carries more: `APP_SECRET=change-me-to-random-string`, `NEXTAUTH_SECRET`, a commented `AUTH_SECRET` that the file notes works as an alias for Auth.js v5, plus `AUTH_TRUST_HOST=true` for running behind a reverse proxy. So the session secret has three possible names in one file, and only one of them is the name the README tells you to set. The env file name is inconsistent too. The quick start copies `.env.example` to `.env.local`, the manual install also copies to `.env.local`, and the Docker Compose section copies to `.env`. The compose file treats the missing `.env` as acceptable by marking it not required, so a quick start that skipped the copy will start Compose with default database credentials instead of yours.

Postgres is bound to loopback because the default password is changeme

The compose file ships a Postgres service whose password defaults to `changeme`, and the app's `DATABASE_URL` is assembled from the same three variables, so an unconfigured deployment runs with a known password. The file addresses this at the port mapping rather than at the credential: the published port is `127.0.0.1:5432:5432`, and an inline comment explains why, noting that a bare `5432:5432` would publish the database to the internet on a host with a public IP while the password stayed at its default. The mapping is kept on loopback only so host-side steps in CONTRIBUTING.md, such as `prisma migrate dev` and `psql`, keep working. Readiness is handled by health checks on both sides: the app checks `http://localhost:3000/api/health` every 30 seconds with three retries, and the database uses `pg_isready` every 10 seconds with five retries, with the app gated on the database being healthy.

The image stages the Prisma CLI instead of installing it twice

The Dockerfile is a four-stage build on node:20-alpine, and two of its comments record bugs rather than instructions. The builder runs `prisma generate` with an inline placeholder `DATABASE_URL` so the generated client is valid without connecting to anything, and the comment says it is inlined so the string does not persist as an image layer. Then `scripts/stage-prisma-cli.mjs` computes the Prisma CLI's full runtime dependency closure from the installed tree and copies it forward, because a hand-listed subset shipped an image that crashed at startup with MODULE_NOT_FOUND, logged as issue #27. The runner deliberately has no npm invocation, since a second install under QEMU arm64 emulation crashes intermittently with SIGILL, tracked as PR #23. The runner also creates a system user at uid 1001 and switches to it, and the command line invokes the CLI's real entry point rather than the bin symlink, because COPY flattens that symlink into a file and the CLI then resolves its WASM from the wrong directory.

Ten MCP tools in four groups, next to a rival with twenty-four

The agent-facing surface is a Model Context Protocol server run straight from TypeScript with no build step. The configuration block points at `npx tsx mcp/server.ts` with a working directory of your checkout, and the manifest repeats it as the `mcp` script:

json
{
  "mcpServers": {
    "crawlseo": {
      "command": "npx",
      "args": ["tsx", "mcp/server.ts"],
      "cwd": "/path/to/crawlseo"
    }
  }
}

Ten tools are exposed across four categories: two for sites, `list_sites` and `get_site_overview`; three for keywords and pages, `get_keywords`, `get_pages` and `get_traffic`; three for crawling and audit, `run_crawl`, `get_crawl_status` and `get_crawl_issues`; and two for performance, `get_vitals` and `get_opportunities`. Ten is fewer than half of what the comparison table attributes to the nearest self-hosted alternative, which is credited with 24 tools, so the agent integration is the feature where the free option is furthest behind. Full setup detail sits in `mcp/README.md` rather than in the main document.

Two rows of the comparison table need your own paid key

The feature table marks keyword research and backlinks with a check for CrawlSEO and for every commercial column, which reads as parity until the note underneath is parsed. Keyword research and backlink data run through DataForSEO on your own key, and the note marks both features as bring your own key, with Google Autocomplete as a free fallback for suggestions. Autocomplete gives you query strings, not volume, difficulty or CPC, and it gives you no backlink profile at all, so the two rows describe what is reachable rather than what is included. The same table marks Core Web Vitals as present for CrawlSEO and Semrush but absent for the other three competitors, which is a real difference, since vitals come from PageSpeed Insights and only need an optional API key for higher quota. The honest summary is that CrawlSEO covers the free halves of the category well and the paid halves as an integration.

Three plan documents and a PRD sit in the repository root

The root tells you how the project is run. `PLAN.md`, `PLAN-V2.md` and `PLAN-V3.md` sit next to `RESEARCH.md`, `TESTING.md`, `DEVELOPMENT.md`, `CONTRIBUTING.md` and `crawlseo-prd.md`, so three generations of planning document are committed rather than superseded in history. Alongside them are `AGENTS.md` and `CLAUDE.md`, and the application itself is split into `app/`, `components/`, `lib/`, `types/`, `prisma/`, `scripts/`, `mcp/` and `docs/`. Two deployment signals disagree in a useful way: the tech stack table names Docker Compose as the deployment route, while a `vercel.json` sits at the root and the example env file tells you to point `NEXTAUTH_URL` at a Vercel app URL for production. The auth layer is also worth a look before you commit, since the stack table says NextAuth.js v5 and the manifest depends on a 5.0.0 beta of next-auth.

Editorial conclusion

CrawlSEO fits a founder or small team that wants Search Console numbers, a crawl audit and an agent-accessible API on its own infrastructure, and who is prepared to supply Google OAuth credentials and a PostgreSQL database. Before deploying, pin the image deliberately, because the documented tag scheme refers to 1.2.x releases while the only tagged release is 0.1.0, and read the environment file rather than the README table, since the secret appears under more than one name. If your keyword and backlink data has to be free out of the box, note that those two features depend on a DataForSEO key with only autocomplete as a fallback.

Frequently asked questions

How many MCP tools does CrawlSEO expose?

Ten, in four categories: list_sites and get_site_overview; get_keywords, get_pages and get_traffic; run_crawl, get_crawl_status and get_crawl_issues; and get_vitals with get_opportunities. The server runs through npx tsx mcp/server.ts.

Does CrawlSEO need a paid API key for keyword research?

Keyword research and backlink data use DataForSEO with your own key, and Google Autocomplete is the free fallback for suggestions. Autocomplete supplies query strings rather than the volume, difficulty and CPC figures the feature table promises.

How do I pin a specific CrawlSEO release in Docker Compose?

Set CRAWLSEO_IMAGE in .env. Compose otherwise defaults to ghcr.io/crawlseo/crawlseo:latest, and the docs describe immutable tags such as 1.2.3 and minor tags such as 1.2, even though the only tagged release is v0.1.0.

Which environment variables does CrawlSEO require?

DATABASE_URL, NEXTAUTH_SECRET, GOOGLE_CLIENT_ID and GOOGLE_CLIENT_SECRET are listed as required, with NEXTAUTH_URL optional and auto-detected. The example env file also carries APP_SECRET, an AUTH_SECRET alias, AUTH_TRUST_HOST and optional cron schedules for crawl, GSC sync and vitals.

Official sources

  1. crawlseo/crawlseo on GitHub
  2. Issues
  3. License: MIT
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/crawlseo-crawlseo.svg)](https://hysenlabs.com/projects/crawlseo-crawlseo)