# AutoClip: an open source AI pipeline that turns long videos into highlight clips

> AutoClip is a self-hosted Python and React tool that downloads a YouTube or Bilibili video, scores segments with a large language model, and cuts the highlights into clips and collections. It is MIT licensed, and its desktop build is where most of the recent work has gone.

**zhouxiaoka/autoclip** — AutoClip : AI-powered video clipping and highlight generation · 一款智能高光提取与剪辑的二创工具

- Repository: https://github.com/zhouxiaoka/autoclip
- Website: https://zhouxiaoka.github.io/autoclip_intro/
- Stars: 9,055 · Forks: 1,670
- Language: Python
- License: MIT
- Published: 2026-09-09 · Updated: 2026-09-09 · Language: en
- Canonical page: https://hysenlabs.com/projects/zhouxiaoka-autoclip

## What AutoClip automates, and for whom

Cutting a two-hour stream into short clips is mostly a search problem. Someone has to watch the whole thing, note the timestamps worth keeping, then cut and title each one. AutoClip moves that search to a language model. The README describes a workflow that starts with a YouTube or Bilibili URL, or a local file upload, and ends with a set of generated clips and AI-recommended collections.

The intended user is a creator or editor who republishes long-form video as short segments, and who is willing to run a server for it. This is not a browser extension or a cloud service. It is a FastAPI backend, a React front end, a Celery worker and Redis, and the README recommends Docker with at least 4GB of memory and 10GB of disk. If you only clip one video a month, the setup cost is hard to justify. If you process a channel's backlog, the pipeline is the point.

One thing to note up front: the README lists Bilibili upload, subtitle editing, mobile support and account management as in development. The core download, analysis, scoring and cutting path is the part that is described as working.

## The pipeline: outline, timeline, scoring, then the cut

The repository layout makes the data flow fairly explicit. Under backend/pipeline there are numbered steps: step1_outline.py, step2_timeline.py, step3_scoring.py and step6_video.py. The README's processing description matches that order. The system prepares the video and subtitle files, asks the model to extract an outline, then identifies topic time ranges, then scores each candidate segment for how interesting it is, then generates clip titles, then recommends collections, and finally renders the video files.

The model call is not hardcoded to one vendor. backend/core/llm_manager.py is described as the AI model manager, and the environment variables expose several providers: API_DASHSCOPE_API_KEY with API_MODEL_NAME defaulting to qwen-plus, plus API_OPENAI_API_KEY, OPENAI_BASE_URL, API_GEMINI_API_KEY and API_SILICONFLOW_API_KEY. The default path is Alibaba's DashScope with Qwen. The requirements file pins openai, google-genai and dashscope, and its comments say these were previously installed by a separate script, so the portable desktop build shipped without them and any provider call failed. That is a useful piece of history: the provider layer is the part most likely to break on an upgrade.

Long jobs are handed to Celery with Redis as the broker, and progress is pushed to the browser over WebSocket. That is why Redis is not optional even in the local install path.

## Installing AutoClip with Docker and running a first project

The README gives two supported routes. Docker is the recommended one, and the compose file defines a redis service, an autoclip app service exposing ports 8000 (backend API) and 3000 (front end), and a celery-worker service. Start by cloning and launching:

```bash
git clone https://github.com/zhouxiaoka/autoclip.git
cd autoclip
./docker-start.sh
```

Before the first start you need an LLM key, because the compose file passes the host environment through to the container. The README's env.example is copied to .env, and the relevant keys are shown there as:

```bash
API_DASHSCOPE_API_KEY=your_dashscope_api_key
API_MODEL_NAME=qwen-plus
```

The compose file uses ${API_DASHSCOPE_API_KEY:-} style substitution and comments that without that block the key a user fills in never reaches the container. So put the key in the project-root .env, not only in your shell.

Once it is up, check status rather than assuming:

```bash
./docker-status.sh
```

What you should see is the app answering its health endpoint at /api/v1/health/ and the worker running. If the worker is down, downloads will appear to hang, because the heavy steps run as Celery tasks.

Then open the front end on port 3000, click the new-project button, paste a YouTube or Bilibili URL, and start the download. The README's usage guide describes the rest: the system downloads the video and subtitles, extracts an outline, identifies topic times, scores segments, generates titles, recommends collections and renders clips. You watch progress in the UI and the finished clips appear on the project detail page.

For a local install instead of Docker, the manual steps are a virtualenv, pip install -r requirements.txt, npm install inside frontend/, plus Redis and FFmpeg from your package manager, then cp env.example .env. The pyproject file also declares a CLI entry point, autoclip, and an MCP server entry point, autoclip-mcp, though the README does not document their usage.

## Where AutoClip will frustrate you

The dependency on a hosted LLM is the first constraint. Every scoring and titling step is a paid API call to DashScope or whichever provider you configure. The README does not state token costs or rate limits, and there is no documented offline mode. If your key expires mid-run, the pipeline steps that depend on it fail while the download step has already succeeded, which leaves partial projects.

Memory is the second. The stated minimum is 4GB, recommended 8GB, and that is before FFmpeg starts encoding multiple clips. A machine sized for browsing will struggle.

Subtitle quality is the third, and it is the quiet one. The pipeline reads subtitles to build the outline and timeline. A video with no captions, or with auto-generated captions in the wrong language, gives the model poor input. The compose file exposes AUTOCLIP_YT_SUBTITLE_LANGS, which suggests language selection is configurable, but the README does not explain how to handle a source with no usable subtitles at all.

Finally, the wrong-tool case: if you want a hosted web app where you paste a link and get clips with no infrastructure, AutoClip is not that. It is a self-hosted system, and the setup is the price of keeping your source files and your API key on your own machine.

## AutoClip compared with yt-dlp plus a manual edit

The honest alternative is not another AI clipper. It is yt-dlp and a timeline in your editor. AutoClip pins yt-dlp==2026.8.19 in requirements.txt, so it is already using that downloader underneath, and the difference is entirely in what happens after the file lands.

With yt-dlp alone you get the video and subtitles and you decide the cuts. That is deterministic, free, and as good as your judgement. AutoClip adds a model in the middle that reads the transcript, proposes timestamps and scores them. The output is faster to produce and less predictable. For a talking-head video with a clear transcript, the scoring step has real signal. For content where the interesting part is visual and unspoken, a transcript-based model has nothing to work with, and yt-dlp plus your own eyes wins.

The other difference is state. AutoClip stores projects, clips and collections in SQLite through SQLAlchemy, with Alembic migrations present in the tree, and the README notes SQLite can be upgraded to PostgreSQL. A manual yt-dlp workflow has no database to migrate. That is a feature if you are running this repeatedly; it is overhead if you are not.

## Maintenance, releases and what the MIT licence leaves you

The repository is not archived, and the last push was on 2026-09-08. The most recent release, v1.2.1, is dated 2026-09-06, and its title names the desktop build. The two releases before it, v1.2.0 and v1.1.0, are both from June 2026. So the release cadence is irregular, and the changelog is the place to check what a version actually changes before upgrading.

Upgrade cost is concentrated in two places. The requirements file pins direct dependencies with == and explains that this is deliberate: an unpinned yt-dlp or openai upgrade is described as the cause of works-on-my-machine failures between CI, Docker and the desktop bundle. That pinning makes upgrades predictable but also means you are on someone else's schedule for security fixes in those libraries. The second place is the database. Alembic is present, so schema changes should come as migrations, but the README does not document a rollback path, and you should back up the SQLite file under ./data before pulling a new version.

The licence is MIT, which permits commercial use and modification. That covers AutoClip's own code. It does not cover the model you call or the videos you download. Bilibili and YouTube both have terms governing automated download, and the README does not discuss them. Treat the licence as answering the question about the software only.

## Conclusion

Adopt AutoClip if you already run Redis and FFmpeg and you want the highlight-selection step automated on your own machine, with a DashScope key as the only paid dependency. Skip it if you need a hosted service, a Windows-native install, or anything the README still marks as in development, including Bilibili upload and subtitle editing. Before you commit, run ./docker-status.sh after the first start and confirm the celery-worker container is healthy, then open one project end to end and check that the clip files land under your mounted ./data volume.

## FAQ

### How does AutoClip work?

It downloads a video and its subtitles, then runs a numbered pipeline: outline extraction, timeline analysis, interest scoring, title generation, collection recommendation and finally video rendering. The model calls go through a configurable provider, with DashScope and qwen-plus as the default, and the heavy steps run as Celery tasks with progress pushed over WebSocket.

### How do I use AutoClip?

Start it with ./docker-start.sh after putting an API key in the project-root .env, then open the front end on port 3000 and create a project from a YouTube or Bilibili URL or a local upload. The README's usage guide describes choosing the link type, pasting the URL and starting the download, after which processing runs automatically.

### Is AutoClip AI free?

The software is MIT licensed, so there is no fee for AutoClip itself. The AI analysis is not free: it calls a hosted model, and the README's configuration uses API_DASHSCOPE_API_KEY with API_MODEL_NAME set to qwen-plus, which is a paid API. The README does not state pricing.

### What is the AutoClip cache?

The README does not document a cache feature under that name. What it does document is Redis, which docker-compose.yml runs as a service and configures through REDIS_URL, used as the Celery message broker and for task state.

## Sources

- [License: MIT](https://github.com/zhouxiaoka/autoclip/blob/main/LICENSE)
- [Project website](https://zhouxiaoka.github.io/autoclip_intro/)
- [README](https://github.com/zhouxiaoka/autoclip/blob/main/README.md)
- [Releases](https://github.com/zhouxiaoka/autoclip/releases)
- [zhouxiaoka/autoclip on GitHub](https://github.com/zhouxiaoka/autoclip)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/zhouxiaoka-autoclip
