# AI Video Transcriber: self-hosted transcription that checks for subtitles first

> A Python and FastAPI service that pulls transcripts from YouTube, TikTok, Bilibili, Apple Podcasts, SoundCloud and 30+ other platforms, falling back to Faster-Whisper only when native subtitles are missing. The subtitle-first path is the interesting design decision; the model dependency is the cost.

**wendy7756/AI-Video-Transcriber** —  Transcribe and summarize videos and podcasts using AI. Open-source, multi-platform, and supports multiple languages.

- Repository: https://github.com/wendy7756/AI-Video-Transcriber
- Website: https://sipsip.ai
- Stars: 3,317 · Forks: 425
- Language: Python
- License: Apache-2.0
- Published: 2026-09-24 · Updated: 2026-09-24 · Language: en
- Canonical page: https://hysenlabs.com/projects/wendy7756-ai-video-transcriber

## What AI Video Transcriber actually solves

The project targets a narrow but real annoyance: getting text out of a video or podcast without uploading the media to someone else's service. You paste a URL from YouTube, TikTok, Bilibili, Apple Podcasts, SoundCloud or one of the 30+ platforms yt-dlp covers, or you drop a local file, and the service returns a transcript, an AI summary, and a translation when the summary language differs from the source language. The README lists accepted upload formats as `.txt`, `.mp3`, `.mp4`, `.m4a`, `.wav`, `.webm`, `.mkv`, `.ogg` and `.flac`.

Who it is for: someone who already has an OpenAI-compatible API key, is comfortable running a Python service or a Docker container, and wants the intermediate artifacts (transcript, summary, optionally the source video) on their own disk. The repository is Apache-2.0 and the last push was on 2026-09-15, so it is not an abandoned snapshot. It is also not a hosted product you can just sign up for. The homepage points at sipsip.ai, but the repository itself is the deliverable.

## Subtitle-first: the one design choice worth understanding

Most transcription tools download the audio and run a speech model on it unconditionally. This one branches. For platforms with native subtitles, the README says transcripts are extracted instantly and no audio download happens; Whisper is described as a fallback, which the project says makes the pipeline dramatically faster. The UI surfaces which branch ran through a badge: green for the subtitle path, cyan for Whisper.

The rest of the pipeline is shared. Text goes through optimization (typo correction, sentence completion, paragraphing), then optional translation, then summarization in one of 11 languages. Local uploads skip the branch entirely: media is normalized with FFmpeg and sent to Whisper, while a `.txt` file skips download and Whisper and enters the text pipeline directly. That last detail is useful if you already have transcripts and only want the summarization half.

Progress is streamed over server-sent events, which is why the README recommends `python3 start.py --prod` for long jobs: hot-reload is disabled so the SSE connection survives 30 to 60+ minute tasks. That is a specific, honest constraint rather than a marketing line.

## Installing AI Video Transcriber and running a first job

Three install paths exist. The automatic one runs a shell script; Docker Compose is the fastest if you already have Docker; manual installation is the one to read if you want to understand what is on your disk. Python 3.8+ and FFmpeg are the stated prerequisites, and FFmpeg is required for yt-dlp audio extraction, merging downloaded video, and normalizing uploads.

Start by cloning and letting the installer do the work:

```bash
git clone https://github.com/wendy7756/AI-Video-Transcriber.git
cd AI-Video-Transcriber

chmod +x install.sh
./install.sh
```

If you prefer containers, the README gives this sequence. Copy the example env file first, since the compose file reads `OPENAI_API_KEY` and `OPENAI_BASE_URL` from it:

```bash
cp .env.example .env
docker-compose up -d
```

The compose service maps port 8000 and sets `WHISPER_MODEL_SIZE` to `base`, `UPLOAD_MAX_MB` to 200, and `VIDEO_MAX_HEIGHT` to 720 by default. Transcripts and downloaded videos live in `/app/temp` inside the container; the `volumes` block that would persist them on the host is commented out, so you have to uncomment it yourself.

For a manual run on macOS, the README recommends a virtualenv because of PEP 668, then FFmpeg via Homebrew:

```bash
python3 -m venv venv
source venv/bin/activate
python -m pip install --upgrade pip
pip install -r requirements.txt
brew install ffmpeg
```

Then start the service and open `http://localhost:8000`:

```bash
python3 start.py
```

What you should see: a dark, responsive UI with a URL field and a dashed upload area. Paste a YouTube link, pick a summary language, leave Keep original video on if you want the source file, and click Transcribe. Watch the badge. If it turns green, subtitles were found and Whisper never ran. If it turns cyan, the audio was downloaded and transcribed, which is where the time goes.

## Bring-your-own-model is flexible, and also the failure mode

The AI Settings panel takes an API Base URL and API Key, and a Fetch button auto-discovers available models from that endpoint. Credentials are stored in the browser's `localStorage` and, per the README, sent only to the provider you chose. Server-side defaults in `.env` are optional.

That design has a sharp edge. The transcription half can run with no external service, but optimization, translation and summarization all need an OpenAI-compatible endpoint. Point the base URL at a local LLM and the whole pipeline stays on your network; point it at nothing and you get a transcript and no summary. The README does not document what happens when the endpoint returns an error mid-job, nor does it describe rollback or retry behavior for a partially completed run. Treat a failed summarization step as an unknown until you test it.

The second cost is Whisper itself. `WHISPER_MODEL_SIZE` accepts tiny, base, small, medium or large, and the default is base. Larger models mean better accuracy on hard audio and more memory and time per minute of media. The Docker Compose file caps the container at 2G of memory, which is a reasonable ceiling for base and a question mark for larger sizes. Nothing in the README states which model fits which hardware, so size it yourself against your own machine.

## Where it is the wrong tool

If you need a hosted service with an SLA, this is not it. You are running FastAPI under uvicorn on your own machine or container, and the README's own production advice (`--prod` to keep SSE stable) tells you long jobs are sensitive to the process staying up.

If your media has no subtitles and you have no GPU or spare CPU, the Whisper path is the whole job, and the default `base` model is a compromise between speed and accuracy that the documentation does not quantify. If your platform is not among the 30+ that yt-dlp supports, the URL path simply will not resolve. And if you need diarization (who spoke when), nothing in the README mentions it. The output is a transcript and a summary, not a speaker-labeled meeting record.

One more boundary: uploads are capped by `UPLOAD_MAX_MB`, 200 by default. A two-hour video file can exceed that, and the README does not describe chunked upload, so raise the limit deliberately or use a URL instead.

## How it compares to running Whisper yourself

The obvious alternative is a bare Faster-Whisper setup: download the audio with yt-dlp, run the model, write the text to a file. That approach gives you total control over the model, the chunking and the output format, and it has no dependency on an LLM endpoint at all.

The difference in approach is exactly the subtitle-first branch. A bare Whisper script transcribes audio that already had a perfectly good subtitle track, which is wasted compute on YouTube and similar platforms. AI Video Transcriber checks first. It also adds the post-processing layer (optimization, translation, summary) and a web UI with live progress, which a script does not have. If your need is batch transcription of local files with no summarization, a script is less machinery for the same transcript. If your need is links plus summaries plus a browsable result, the extra layer is the point. There is also an `mcp_server.py` in the repository with an `mcp>=2.0.0` dependency, described as being for agent and tool integration, so the same pipeline can be driven from Claude, Codex or scripts rather than only the UI.

## Licence and the real upgrade cost

The repository is Apache-2.0, with a LICENSE file at the top level. That permits commercial use and modification under its terms, but it says nothing about the models or APIs you connect it to: your OpenAI-compatible provider's terms, and the licence of any local LLM you point it at, are separate questions. This is not legal advice; read the LICENSE file and your provider's terms.

Upgrade cost is dominated by two moving parts. `requirements.txt` pins lower bounds, not exact versions (`fastapi>=0.110.0`, `yt-dlp>=2024.12.13`, `faster-whisper>=1.1.0`, `openai>=1.51.0`, `pydantic>=2.7.0`, `aiofiles>=24.1.0`, `python-multipart>=0.0.9`, `uvicorn[standard]>=0.27.0`), so a fresh `pip install -r requirements.txt` can pull newer releases than the ones the project was built against. yt-dlp in particular breaks when platforms change their pages, which is precisely why the loose pin matters here. The Dockerfile uses `python:3.12-slim-bookworm` and installs the same requirements file, so container and local installs track each other.

There are no retrieved releases, so there is no changelog to read before upgrading. Budget for re-testing one URL job and one upload job after any dependency bump, and keep a working container image around as your rollback.

## Conclusion

Adopt it if you want transcripts and summaries to stay on your own machine and you already have an OpenAI-compatible endpoint, or if you mostly process YouTube links where the subtitle path skips Whisper entirely. Do not adopt it if you have no GPU or CPU budget for Faster-Whisper and no API key for a summarization model, because the pipeline stops at the transcript without one. Before committing, check the `WHISPER_MODEL_SIZE` value in `.env` against your available memory, run one short YouTube URL and one local `.mp3` through the same instance to confirm both paths work, and read the Apache-2.0 LICENSE file for the terms that apply to your deployment.

## FAQ

### Can AI Video Transcriber transcribe a video for free?

The software itself is Apache-2.0 and there is no licence fee, but the pipeline still needs an OpenAI-compatible endpoint for optimization, translation and summarization, and that provider may charge. You can point the API Base URL at a local LLM to avoid a paid provider, at the cost of running the model yourself.

### How much does transcription with AI Video Transcriber cost?

The repository does not state any pricing. Costs come from whatever you connect: the OpenAI-compatible provider you configure in AI Settings, plus your own machine time for Faster-Whisper when no subtitles exist. The README does not publish token or per-minute estimates.

### Is AI Video Transcriber legal to use?

The project is released under Apache-2.0, which governs the code. Whether you may download or transcribe a given video depends on the platform's terms and your local law, and the README does not address that question at all.

### What is the best AI model for transcribing videos in AI Video Transcriber?

The README does not rank models. It exposes `WHISPER_MODEL_SIZE` with options tiny, base, small, medium and large, defaulting to base, and for summarization you pick any model your OpenAI-compatible endpoint advertises via the Fetch button.

## Sources

- [Issues](https://github.com/wendy7756/AI-Video-Transcriber/issues)
- [License: Apache-2.0](https://github.com/wendy7756/AI-Video-Transcriber/blob/main/LICENSE)
- [Project website](https://sipsip.ai)
- [README](https://github.com/wendy7756/AI-Video-Transcriber/blob/main/README.md)
- [wendy7756/AI-Video-Transcriber on GitHub](https://github.com/wendy7756/AI-Video-Transcriber)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/wendy7756-ai-video-transcriber
