Self-hosted service
getmaxun/maxun avatar
getmaxun/maxun

Maxun: Open Source No-Code Platform for Web Scraping and AI Data Extraction

🔥 The open-source no-code platform for web scraping, crawling, search and AI data extraction • Turn websites into structured APIs in minutes 🔥

17,595 stars1,519 forksTypeScriptAGPL-3.0

At a glance

What is it?
Maxun is an open source TypeScript platform that lets teams extract structured data from websites without writing code, using a visual recorder, natural-language AI mode, or programmable API and SDK. It self-hosts via Docker Compose and requires PostgreSQL and MinIO alongside the application.
Who is it for?
Maxun fits teams that need a self-hosted, browser-based extraction platform with a visual UI, scheduling, and both API and SDK access, without writing custom scraping code for each site. Teams that need to embed scraping logic inside an existing Python or TypeScript application, rather than running a separate scraping service, will find a lightweight library a better fit.
Can I use it commercially?
Yes, with strict conditions. AGPL-3.0 is a network copyleft licence: if people use a modified version over a network, for example as a hosted service, you must offer them its source code under the same licence.
Is it still maintained?
Yes. The repository last received commits 1 day ago.
What is it written in?
Mainly TypeScript, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What Maxun Solves and Who It Is For

Most web scraping projects start with a one-off script and end with a maintenance burden. Sites update their DOM, add anti-bot measures, change login flows, and the scraper breaks. Maxun addresses this as a platform rather than a library: instead of writing code against each site, engineers record browser interactions and the platform replays them as reusable robots. The robots run on schedule, expose a structured API, and can be monitored for changes.

The README describes the project as turning "any website into a structured API" and lists six distinct operation types: Extract, Scrape, Crawl, Search, Monitor, and Document Extraction. The intended users are data engineering teams, product teams that need competitive data or price monitoring, and developers building data pipelines that require regularly refreshed web data. An alternative hosted version exists at app.maxun.dev, but the self-hosted path is the focus for teams with data residency or cost constraints.

Maxun is an application, not a library. It runs as a multi-container service with a frontend, backend, database, and object store. Teams that want to embed web extraction logic inside their own application code will need to use Maxun's SDK or API rather than the platform directly.

Recorder Mode and AI Mode: Two Extraction Approaches

The Extract operation in Maxun provides two distinct modes.

Recorder Mode records browser interactions and converts them into a reusable extraction robot. The engineer navigates a site in the recorder, Maxun captures the actions, and the resulting robot can replay that sequence to extract data on demand or on a schedule. This approach handles login flows, pagination, and multi-step navigation that declarative CSS-selector-based scrapers cannot express without code.

AI Mode accepts a natural-language description of what to extract and uses the AI backend to handle extraction. The package.json dependencies include `@anthropic-ai/sdk` for the Anthropic SDK and the docker-compose.yml includes an `OLLAMA_BASE_URL` environment variable (`http://host.docker.internal:11434` by default), meaning Maxun supports both cloud AI models and a locally running Ollama instance for extraction. Teams with data privacy requirements can route AI extraction through their own Ollama server rather than a third-party API.

The Scrape operation converts a full webpage to Markdown, HTML, or a screenshot. This is the simpler path for teams that need raw content rather than structured fields.

Crawl, Search, Monitor, and Document Extraction

Beyond the two extraction modes, Maxun provides four additional operation types.

Crawl traverses entire websites across pages, not just a single URL. This covers link-following, sitemap navigation, and content collection across a domain.

Search queries the web and extracts results with time-based filters. The documentation link in the README points to a dedicated search introduction, suggesting this is a distinct operation with its own configuration rather than a thin wrapper over a search API.

Monitoring tracks a website over time and detects changes. The v0.0.47 release, published on 2026-09-16, is titled "Website Monitoring," indicating this feature received focused development in the most recent release cycle. Teams that need alerts when pricing, availability, or content changes on a target site can use this as a scheduled check with notifications.

Document Extraction parses structured data from documents, covering the use case where the target content is a PDF, spreadsheet, or other document format rather than an HTML page.

All six operation types are accessible through the same Dashboard, API, SDK, and CLI interfaces, meaning teams can mix and match extraction methods within a single Maxun deployment.

Deploying Maxun: Docker Compose and Infrastructure Requirements

The recommended self-hosting path is Docker Compose. The `docker-compose.yml` in the repository defines four services: a PostgreSQL 13 database, a MinIO object store, a backend server, and an nginx-fronted frontend.

The backend container runs the Playwright-based browser automation. The docker-compose.yml sets `shm_size: '2gb'` and `mem_limit: 6g`, meaning the Docker host needs at least 6GB of memory available for the backend alone. The backend image is `getmaxun/maxun-backend:latest` and the environment includes `CHROMIUM_FLAGS: '--disable-gpu --no-sandbox --headless=new'` for running Chromium in a containerized environment. The `seccomp=unconfined` security option is set to allow Chromium's sandboxing mechanisms.

MinIO provides the object store for screenshots and other binary outputs. It listens on port 9000 for API access and port 9001 for the MinIO web console by default, both configurable through environment variables. PostgreSQL listens on port 5432 (configurable). The backend defaults to port 8080.

The environment variables are listed in the `ENVEXAMPLE` file at the repository root. The `SETUP.md` file provides additional local setup guidance. For teams that need to upgrade an existing Docker Compose installation, the documentation at docs.maxun.dev covers the upgrade path separately from the initial installation.

To start the stack:

bash
docker-compose up -d

Access Interfaces: Dashboard, API, SDK, CLI, and MCP

Maxun exposes five interfaces, each serving a different integration pattern.

The Dashboard at app.maxun.dev (or the self-hosted equivalent) is a visual web interface for building robots, triggering runs, and inspecting results without writing code.

The API provides programmatic access to all of Maxun's extraction and crawling capabilities. Teams can trigger robots, retrieve results, and manage configuration over HTTP, allowing Maxun to fit into an existing automation pipeline.

The SDK, with its Node.js package at `github.com/getmaxun/node-sdk`, provides a typed programmatic interface for the same capabilities. This covers use cases where the API would need custom retry logic or strongly typed result handling.

The CLI creates robots, triggers runs, and retrieves data from the terminal, making it straightforward to integrate Maxun into shell scripts or CI workflows.

The MCP interface connects Maxun to AI agents through the Model Context Protocol. Teams building AI agents that need to retrieve live web data can expose Maxun's robots as MCP tools, allowing an AI agent to trigger a scrape and receive structured results as part of its context.

The package.json includes `@anthropic-ai/sdk` as a runtime dependency and uses `@tanstack/react-query` for the frontend, `express` and `express-session` for the backend, `graphile-worker` for background job queuing, and `playwright` for browser automation.

AGPL-3.0 Licensing and Self-Hosting Constraints

Maxun is licensed under AGPL-3.0-or-later, as shown in the `package.json` license field and the `LICENSE` file. AGPL-3.0 includes a network use provision: if you run a modified version of Maxun as a service accessible to users over a network, you must make the modified source code available to those users under the same license. This is a stricter requirement than the MIT or Apache licenses that many infrastructure projects use.

Teams deploying an unmodified Maxun instance internally, or as an internal tool for their own engineers, can do so without triggering the modification and disclosure requirements. Teams that plan to build a commercial product on top of Maxun, modify its source, and offer that service to external customers need to review whether their modifications must be disclosed under AGPL-3.0.

The self-hosted path does not require any connection back to the Maxun cloud. The docker-compose.yml images pull from Docker Hub as `getmaxun/maxun-backend:latest` and `getmaxun/maxun-frontend:latest`, which requires network access to Docker Hub at deploy time, but the running application does not depend on Maxun's cloud services.

Maxun vs Crawl4AI: Application Platform vs Python Library

Crawl4AI is an open source Python library for LLM-friendly web crawling. It provides a programmatic Python API that development teams embed directly in their application code. Crawl4AI handles one task: crawl a URL and return content in a format optimized for LLM consumption. It does not include a visual dashboard, a built-in job scheduler, an object store for screenshots, or a multi-user authentication layer.

Maxun is a full application stack. It provides the browser infrastructure, storage, job scheduling, user management, and the visual interface as a single deployment. Teams that want to add web scraping capability to their organization without building that infrastructure from scratch get more from Maxun's approach. Teams that have existing Python infrastructure and want to call a scraping function as part of their own code will find Crawl4AI's library model a simpler integration.

The RELATED SEARCHES for Maxun include Crawl4AI and Langflow, suggesting people evaluate them in the same context. They solve overlapping problems but at different layers: Maxun is the service, Crawl4AI is the library. The choice depends on whether the team needs a self-contained scraping service or a library that fits into existing code.

Editorial conclusion

Maxun fits teams that need a self-hosted, browser-based extraction platform with a visual UI, scheduling, and both API and SDK access, without writing custom scraping code for each site. Teams that need to embed scraping logic inside an existing Python or TypeScript application, rather than running a separate scraping service, will find a lightweight library a better fit. Before deploying, verify that the Docker host has at least 6GB of memory available for the backend container, review the AGPL-3.0 terms if you plan to build a commercial service on top of Maxun, and work through the environment variables documented in `ENVEXAMPLE`.

Frequently asked questions

What is Maxun and what does it do?

Maxun is an open source no-code platform for web scraping, crawling, monitoring, and AI-based data extraction. It turns websites into structured APIs by recording browser interactions or using AI to describe what to extract, without writing custom scraping code.

How do I self-host Maxun?

The recommended path is Docker Compose using the `docker-compose.yml` in the repository. It starts a PostgreSQL 13 database, a MinIO object store, the backend (which runs Playwright for browser automation), and an nginx-fronted frontend. The host needs at least 6GB of memory for the backend container.

What does the AGPL-3.0 license mean for using Maxun?

AGPL-3.0 allows free use and self-hosting of an unmodified Maxun instance. If you modify Maxun and run that modified version as a network service accessible to external users, you must make your source changes available under AGPL-3.0. Unmodified internal deployments do not trigger this requirement.

Official sources

  1. Official documentation
  2. Official README
  3. Project repository
  4. Release notes
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/getmaxun-maxun.svg)](https://hysenlabs.com/projects/getmaxun-maxun)