Model or dataset
pycaret/pycaret avatar
pycaret/pycaret

PyCaret 4.0: the sklearn-native engine plus a self-hosted control plane

Open-source, low-code AutoML platform for Python. PyCaret 4.0: sklearn-native engine + React control plane.

9,849 stars1,842 forksPythonNOASSERTION

At a glance

What is it?
PyCaret's 4.0 line splits the project into a scikit-learn based AutoML library and a FastAPI plus React platform you run with docker compose. Here is what the repository actually ships today, and where the gaps are.
Who is it for?
Adopt PyCaret 4.0 if you want a self-hosted, point-and-click AutoML platform on a laptop or a small-team server, and you accept that the 4.0 line is an alpha with the single-binary shape. Stay on the frozen 3.x line if you need the mature Python API today, and skip it entirely if you need a supported release with a published support window.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Yes. The repository last received commits 69 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What PyCaret 4.0 solves, and who it is aimed at

PyCaret's own description calls it an open-source, low-code AutoML platform for Python. The 4.0 line changes the shape of that offer. Instead of a library you import in a notebook and drive with function calls, the repository now describes a self-hosted ML platform with three parts: an engine, a control plane, and a UI. The README states the goal as running the whole thing on a laptop in five minutes.

The audience is narrower than the older library's. The README's golden path is a person who wants to upload a CSV, pick a task type (classification, regression, clustering, anomaly detection or time series), and let a run train roughly twelve algorithms without writing orchestration code. The control plane adds workspaces, projects, experiments, runs, a model registry, deployments, approvals, monitoring, drift, lineage, webhooks, schedules and LLM-advisory copilots. That is a platform for a small team that wants experiment tracking and promotion workflow inside its own infrastructure rather than in a hosted service.

If you only want a Python function that compares classifiers, the platform is more machinery than you asked for. The engine package is the part that matters to you, and it is published separately.

The engine, the control plane, and the single-binary compromise

The repository is a uv workspace, not a single package. The root pyproject.toml declares two Python members, packages/engine and services/api, and states plainly that there is no monolithic package at the repo root anymore. The name pycaret is published from packages/engine/pyproject.toml, and pycaret-server from the API service. Non-Python members (the React UI, a future Electron app, infra modules) sit outside the workspace declaration.

At runtime, compose.yml defines two services. pycaret-api is a FastAPI process that also runs the in-process APScheduler worker and the in-process compute, and it writes to a SQLite file. pycaret-web is an nginx-served React bundle that proxies /api and /ws to the backend. Persistence is a named volume, pycaret-data, mounted at /data inside the API container, holding the SQLite database, uploaded CSVs, fitted .pkl files and the Fernet key.

The README is unusually candid about this being a compromise. It describes the two-container shape as deliberately compact and compares it to the pattern used by Plausible, Vaultwarden and n8n for self-hosters. The stated intent is that every external dependency (storage, database, secrets, auth, queue, compute) should sit behind a Protocol with multiple implementations, so the same codebase can later target separate api, worker and runtime containers on RDS, S3, SecretsManager, SQS and Fargate. The README also says most of those Protocol extractions are still ahead, and points to PLATFORM_ARCHITECTURE.md for the gap matrix.

That is the honest reading of the 4.0 architecture: the split is designed, not built. Treat the current compose file as the deployment target, not a stepping stone you can rely on being finished.

Installing PyCaret 4.0 with docker compose

The README gives a one-command local install. Prerequisites are Docker Desktop 4.27+ or Docker Engine 25.0+ with Compose v2.24+. The README states the first build takes about five minutes (Python dependencies, npm install, Vite build) and later starts take under thirty seconds.

bash
git clone https://github.com/pycaret/pycaret.git
cd pycaret
docker compose up --build

When the API logs show the startup-complete line and the web container starts nginx, open http://localhost:3020. The setup screen asks you to create the first admin account and a workspace. The web UI listens on 3020 and the API on 8020; nginx proxies /api and /ws from the former to the latter.

Configuration goes through a .env file at the repo root, which is gitignored. The README and .env.example both point at .env.example as the template.

bash
cp .env.example .env
# edit .env: set PYCARET_JWT_SECRET to something real, etc.
docker compose up

The knobs the README calls out are PYCARET_SECRETS_KEY (the Fernet key for encrypting LLM API keys and connection passwords at rest, auto-generated on first run and persisted to the volume), PYCARET_JWT_SECRET (must be a strong random value in any production deploy), PYCARET_DATABASE_URL (swap SQLite for Postgres with a postgresql+psycopg:// URL), and PYCARET_STORAGE_BACKEND=s3 with bucket credentials.

If port 3020 or 8020 is taken, the README's troubleshooting table gives the override:

bash
PYCARET_WEB_PORT=3030 PYCARET_API_PORT=8030 docker compose up

For the first real use, follow the golden path: create the admin account, optionally paste an Anthropic or OpenAI key under Settings, then go to Build, Datasets and click Browse samples to load a bundled CSV (the README names juice, bank and iris). Create a project, start an experiment, choose the dataset, task type and target column, and the run trains about twelve algorithms. The run detail page ranks the trials on a leaderboard, and promoting one moves it into the registry.

Where the 4.0 line will bite you

The README carries a warning at the top: 4.0 is work in progress, and the main branch is the 4.0 line. The latest release listed is v4.0.0a1 from 2026-04-24, an alpha, preceded by v4.0.0a0 the day before. The previous stable release, 3.3.2, dates from 2024-04-28. The README states that the 3.x line is frozen on PyPI as pycaret 3.4.0 with no further commits. Anyone expecting a supported, stable 4.0 today is looking at the wrong artefacts.

The single-binary shape has operational consequences the README does not hide. The scheduler and compute run inside the API process, so a long training run competes with request handling in the same container. The database is SQLite by default. The README calls the shape great for laptops and fine for a small-team production deploy, which is a fair description of the ceiling.

There is a data-loss trap worth reading twice. The README's troubleshooting table says the database is encrypted and a forgotten admin password cannot be recovered; the fix is docker compose down -v, which wipes the volume, the database and the Fernet key. If you set PYCARET_SECRETS_KEY explicitly you can keep the key across instances, but the accounts and artefacts still go. The compose file's own comments spell out the same distinction: down preserves the volume, down -v wipes everything.

A smaller failure mode: on the very first up, the UI may load while every /api call returns 404 because the API container is still running Alembic migrations. The README says to wait about thirty seconds for the health check. It is cosmetic, but it will make you think the install failed.

PyCaret versus calling scikit-learn yourself

The engine is described as sklearn-1.7-based, and the root pyproject.toml targets Python 3.13 with ruff linting enforced only on the new 4.0 surface. Legacy paths under pycaret/internal, pycaret/containers, the per-task oop.py modules, datasets.py and utils are excluded from linting and annotated as scheduled for a Phase 5 drain. That tells you how the project sees its own relationship to plain scikit-learn: the old wrapper layer is on the way out, and the new engine is meant to sit closer to the estimator API.

The real alternative for most readers is not another AutoML brand, it is scikit-learn plus your own loop. With scikit-learn you write the cross-validation, the preprocessing pipeline and the model comparison yourself, and you own every decision. With PyCaret's engine you get the comparison and the leaderboard without writing it, and you accept the project's choices about preprocessing and which algorithms are in the twelve. That trade is the same one every AutoML wrapper makes, and the 4.0 architecture does not change it. What 4.0 adds is the surrounding platform: registry, approvals, drift and lineage, which scikit-learn has no opinion about and which you would otherwise assemble from separate tools.

One caveat on the comparison. Because the 3.x line is frozen and 4.0 is alpha, the honest alternative to PyCaret 4.0 for a production project today is PyCaret 3.4.0 from PyPI, not the main branch. The two are different codebases, and the repository says so.

Licence, maintenance and the cost of upgrading

The repository's licence field is NOASSERTION, and the README's badge links to a LICENSE file at the repo root. GitHub could not classify it automatically. If your organisation has licence review, that file is the thing to read before you plan a deployment, and I am not going to guess at its contents. The README does not discuss licence implications for commercial use.

Maintenance signals are mixed but legible. The last push to main was on 2026-07-23. The repository is not archived. The most recent release is the 4.0.0a1 alpha from 2026-04-24, and the README states that 3.x is frozen at pycaret 3.4.0 with no further commits. So the active line is 4.0, and it is pre-release.

Upgrade cost depends on which side you are on. Moving from 3.x to 4.0 is not a version bump: the workspace layout means the engine now lives at packages/engine, the linting exclusions mark whole 3.x modules for deletion, and the platform is a new deployment surface with its own database, secrets and volume. A team on 3.x in a notebook has a migration, not an upgrade. A team starting fresh on 4.0 should plan for the alpha's churn and read docs/revamp/ROADMAP.md and docs/revamp/STATUS.md before pinning anything. The repository also ships OPERATIONS.md, INSTALL.md and TEST_PLAN.md, which is more operational documentation than most projects at this stage carry.

Editorial conclusion

Adopt PyCaret 4.0 if you want a self-hosted, point-and-click AutoML platform on a laptop or a small-team server, and you accept that the 4.0 line is an alpha with the single-binary shape. Stay on the frozen 3.x line if you need the mature Python API today, and skip it entirely if you need a supported release with a published support window. Before committing, read docs/revamp/STATUS.md, check that packages/engine/pyproject.toml still publishes the pycaret name, and confirm which of the storage, database and queue Protocol extractions your deployment depends on are actually implemented.

Frequently asked questions

What is PyCaret used for?

PyCaret is described as an open-source, low-code AutoML platform for Python. In the 4.0 line it ships an sklearn-1.7-based engine plus a self-hosted control plane with workspaces, experiments, runs, a model registry, deployments, monitoring and drift.

How to install PyCaret 4.0?

The README's local install is git clone of the repository, cd into it, then docker compose up --build, which needs Docker Desktop 4.27+ or Docker Engine 25.0+ with Compose v2.24+. The UI then runs at http://localhost:3020 and the API at http://localhost:8020.

How to use PyCaret?

The README's golden path is to create the first admin account and workspace, optionally add an LLM API key under Settings, upload a dataset or load a bundled sample such as juice, bank or iris, create a project, then start an experiment by choosing the dataset, task type and target column. The run trains about twelve algorithms and ranks them on a leaderboard where you can promote a winner.

What is PyCaret in Python?

It is a Python AutoML project published from packages/engine inside a uv workspace, with the pycaret name on PyPI and a separate pycaret-server package for the FastAPI backend. The repository states there is no longer a single monolithic package at the repo root.

How to install PyCaret in Python?

The README documents the platform install through docker compose rather than a pip command for the 4.0 line, and states that the 3.x line is frozen on PyPI as pycaret 3.4.0. The 4.0 engine is published from packages/engine/pyproject.toml, but the README gives no pip install instructions for it.

Official sources

  1. Issues
  2. Project website
  3. pycaret/pycaret on GitHub
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/pycaret-pycaret.svg)](https://hysenlabs.com/projects/pycaret-pycaret)