# VidBee: a desktop downloader that transcribes and summarizes what it saves

> VidBee is a TypeScript desktop app that downloads video and audio from 1000+ sites, transcribes them locally with Whisper, SenseVoice, Parakeet or Qwen3-ASR, and sends the transcript to the AI provider you pick. The local-only promise stops at speech recognition.

**nexmoe/VidBee** — Download video and audio from  YouTube ,  TikTok ,  Twitter ,  Instagram ,  Facebook ,  Twitch ,  Bilibili , and 1000+ sites—or import local media. Create searchable transcripts on your computer, then summarize, translate, or ask questions with your preferred AI provider.

- Repository: https://github.com/nexmoe/VidBee
- Website: https://vidbee.org
- Stars: 10,703 · Forks: 831
- Language: TypeScript
- License: MIT
- Published: 2026-09-21 · Updated: 2026-09-21 · Language: en
- Canonical page: https://hysenlabs.com/projects/nexmoe-vidbee

## The gap VidBee fills: a download folder that is actually searchable

Most downloaders stop when the file lands on disk. You end up with a directory of MP4 files whose names tell you nothing about what is inside them, and finding one sentence you half remember means opening files one by one. VidBee's answer is to treat transcription as part of the download pipeline rather than a separate step. The README describes the app as turning video and audio into an organized, searchable library, and the feature list backs that up: transcripts carry speaker labels and timestamps, and you can search spoken text or speaker names, click a timestamp, and resume playback from that exact moment. The intended user is someone who accumulates long-form media (interviews, lectures, conference talks, podcasts) and wants to query it later. The AI prompts are the second half of that idea: built-in prompts produce bullet summaries, FAQs, statistics or translations from the transcript you already have. Note the split the README itself flags. Speech recognition runs on your computer, so the media is not uploaded to a transcription service. The AI step is different: the prompt and transcript content go to whichever provider you selected. That distinction is the single most important thing to understand before installing.

## How the pipeline runs: yt-dlp-style downloads, local ASR, provider-controlled prompts

The architecture visible in the repository is a pnpm workspace with several apps. There is apps/desktop, which is the Electron application the README points users to; apps/api, a server with its own Dockerfile and a health endpoint; apps/web, a Vite-built front end; and apps/extension. The root package.json exposes scripts that fan out to those packages, including dev:desktop, dev:api, dev:web and dev:extension, plus build:win, build:mac and build:linux for packaged desktop releases. The download layer accepts a single link or a batch and tracks active, completed and failed tasks with live progress, speed, file size and the format being saved; tasks can be paused, resumed, retried or removed without rebuilding the queue. Transcription can start automatically after a download completes, and the number of concurrent transcript jobs is configurable. The model side is deliberately plural: Whisper, SenseVoice, Parakeet and Qwen3-ASR families, with a recommendation for your machine and language, downloadable and switchable from Settings. Speaker handling is a two-part mechanism worth noting. Detection can be automatic, or you can fix a speaker count and re-label the conversation without changing the recognized words. That separation means a bad diarization result does not force you to re-run recognition. On the AI side, providers are pluggable: OpenAI, Anthropic, Google, DeepSeek, Groq, Azure, Hugging Face, OpenRouter, xAI, Ollama and LM Studio, plus a custom provider defined by name, Base URL, model ID and API key with a connection test. API keys are stored on your computer.

## Installing VidBee and running a first transcription

The README does not give a command-line install for end users. It points to a download page at vidbee.org/download and documentation at vidbee.org/docs, and the release history shows packaged versions such as v2.1.0. So the normal path is: fetch the build for your platform, install it, and configure models in Settings. If you want to build from source, the repository is a pnpm workspace, and the root package.json pins the package manager. The commands below come from that file. The first installs dependencies and the second starts the desktop app in development mode.

```bash
pnpm install
pnpm dev:desktop
```

The root scripts also include setup, which delegates to the desktop package, and per-platform build targets. A typical from-source build for Linux looks like this:

```bash
pnpm setup
pnpm build:linux
```

For a first real use, the README's flow is: paste one link or add a batch, let the download finish, and let transcription start automatically if you enabled it. In Settings you pick a model family the app recommends for your hardware and language, then download it. After the transcript exists, you can search spoken text or speaker names, click a timestamp to jump playback, copy the transcript, or export plain text or Markdown. To keep the AI step local, choose Ollama or LM Studio as the provider instead of a hosted API. The repository also ships a docker-compose.yml for the server side, which exposes the API on port 3100 by default and the web app on port 3000, with data under /data/vidbee and downloads under /data/downloads. That compose file is for the API and web components, not the desktop app.

```yaml
services:
  api:
    environment:
      VIDBEE_API_PORT: ${VIDBEE_API_PORT:-3100}
      VIDBEE_DOWNLOAD_DIR: /data/downloads
      VIDBEE_DATA_DIR: /data/vidbee
```

## Where VidBee stops being the right tool

The privacy story has a seam. Local transcription means your media never leaves the machine, but the moment you run a built-in prompt against a hosted provider, the transcript text goes out. For a confidential interview or an internal meeting recording, that is a meaningful difference from a fully offline workflow. The README states the workaround plainly: choose Ollama, LM Studio or another local endpoint to keep the AI step on your computer too. If you will not run a local model, treat the AI features as a cloud feature with a local front end. The second constraint is form factor. VidBee is a desktop app. The workspace does contain an API and a web app with a Dockerfile and a compose file, but the README's getting-started path is the desktop download, and the feature descriptions (download queue, settings, model switching, prompt editing) are written as GUI interactions. If your goal is a shell script that pulls a channel on a cron job and writes subtitles into a directory, this is more machinery than you need. Third, the project describes itself as under active development and invites feedback for any issue encountered, which is a fair signal about stability expectations. Finally, the README does not document rollback behaviour for a failed or partial download, and it does not document what happens to queued RSS items when a feed changes shape. Those are gaps to test yourself rather than assume.

## VidBee versus yt-dlp: a GUI library versus a download engine

The honest comparison is with yt-dlp, because VidBee's download layer covers the same territory: 1000+ sites, format selection, playlists and channels. The difference is what surrounds the download. yt-dlp is a command-line program designed to be scripted and composed; you get a file and whatever post-processing flags you pass. VidBee is a desktop application with a persistent queue, an RSS subscription layer with per-feed rules (keyword filters, automatic tags, a download-only-the-latest option, custom directory and filename templates), metadata controls such as embedding the source title and artist, adding the thumbnail as cover art and keeping chapter markers, and then transcription and AI prompts on top. The trade is control and portability for a managed workflow. If you already have yt-dlp wired into automation, VidBee does not replace that; it replaces the folder of untitled files at the end of it. A second comparison worth making is against transcription services that accept an upload. VidBee's local speech recognition avoids that upload for the media itself, at the cost of running a model on your own hardware and managing model downloads through Settings. That is a real cost in disk space and compute, and the README's model recommendation exists precisely because the choice depends on your machine.

## Maintenance, licence and what upgrading costs you

VidBee is MIT licensed, which permits commercial and private use, modification and redistribution provided the copyright notice and licence text are retained. That is the standard permissive arrangement; it is not legal advice, and if you redistribute a modified build you should read LICENSE in the repository rather than rely on a summary. The repository is not archived, and the last push was on 2026-09-13, so it is being changed. The release cadence visible in the repository is close together: v2.0.1 and v2.0.2 both landed on 2026-08-23, and v2.1.0 followed on 2026-08-30. That kind of spacing suggests fixes ship quickly, and it also means you should expect to update rather than pin and forget. The upgrade cost is concentrated in two places. Models are downloaded separately from the app and switched in Settings, so a model change is a user action, not an app update. The workspace uses database migrations (db:generate and db:migrate in the root scripts), which implies a local store that can change shape between versions; the compose file puts that store at /data/vidbee/vidbee.db. If you run the API and web containers rather than the desktop app, pin your image tags and back up the vidbee-data volume before pulling a new build. The README does not describe a migration rollback path.

## Conclusion

Adopt VidBee if you want one desktop queue that downloads, transcribes and then lets you summarize or translate without uploading the media to a transcription service. Skip it if you need a headless CLI on a server with no GUI, or if you expect the AI step to be local by default: only Ollama, LM Studio or another local endpoint keeps prompts on your machine. Before committing, verify the model family your hardware can run in Settings, and check whether your target site is covered by the 1000+ supported list rather than assuming it is.

## FAQ

### Is VidBee safe to use?

The README states that speech recognition runs on your computer, so media is not uploaded to a transcription service, and that API keys are stored on your computer. It also states that when you run an AI prompt, the prompt and transcript content go to the provider you selected.

### Is a video downloader legal?

The README does not address the legality of downloading. It describes VidBee as a free, open-source desktop app for downloading media from 1000+ supported sites and importing local files, and it is licensed under MIT.

### Which YouTube downloader is most trusted?

The README does not rank downloaders or make trust claims about other tools. It states that VidBee is open source under the MIT licence, that speech recognition runs on your computer, and that AI prompts go to the provider you select.

## Sources

- [License: MIT](https://github.com/nexmoe/VidBee/blob/main/LICENSE)
- [nexmoe/VidBee on GitHub](https://github.com/nexmoe/VidBee)
- [Project website](https://vidbee.org)
- [README](https://github.com/nexmoe/VidBee/blob/main/README.md)
- [Releases](https://github.com/nexmoe/VidBee/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/nexmoe-vidbee
