BiliNote: self-hosted AI notes from Bilibili, YouTube and Douyin videos
AI 视频笔记生成工具 让 AI 为你的视频做笔记
At a glance
- What is it?
- BiliNote turns a video link into a structured Markdown note using a local Whisper transcription plus an LLM you configure yourself. The Docker path is the supported one, and the trade-offs sit in model size, memory and where your API keys live.
- Who is it for?
- Adopt BiliNote if you want video notes to stay on your own machine and you are willing to run Docker, pick a Whisper size and enter your own LLM key. Do not adopt it if you want a managed service with no moving parts, or if your host cannot spare roughly 4 GB of RAM for the backend container.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 29 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 28, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What BiliNote actually produces, and who it is for
BiliNote takes a video URL and returns a Markdown note. The README lists Bilibili, YouTube, Douyin, Kuaishou and local video files as inputs, and the output is a structured note with optional auto-captured screenshots, optional jump links back to the original video, a video cover banner at the top and a source link. There is also a RAG-based question and answer panel over the generated note, with Function Calling so the model can query the original transcript data.
The target user is someone who watches long Chinese-language video and wants a written record without paying a hosted service. The project ships both a web frontend and desktop clients for Windows and macOS, plus a browser extension for Chrome, Edge and Firefox MV3. The README is explicit that the online version at www.bilinote.app exists for people who do not want to install dependencies, configure a proxy or download models, which tells you the maintainer expects self-hosting to be the harder path.
One thing worth noting for anyone outside China: the default transcription stack and the mirror endpoints are chosen for mainland network conditions, not for a European or American VPS. That is a design centre, not a defect, but it shapes what you have to change.
The pipeline: subtitle first, Whisper second, LLM last
The architecture is a FastAPI backend with a React 19 frontend, joined by nginx in the Docker setup. The backend directory holds the database, config and static assets; the frontend is built separately and served through nginx, which proxies to the backend over an internal port.
The interesting part is how the transcript is obtained. For Bilibili, the backend has a BilibiliSubtitleFetcher that goes through the player API to pull subtitles, used as a fallback behind yt-dlp. The browser extension goes further: it grabs subtitles directly in the user's browser using the local logged-in cookies, which skips backend audio transcription entirely. For YouTube, the changelog for v2.0.0 states that subtitles are preferred and audio download is skipped when subtitles exist. Only when no subtitle track is available does the audio path run, which means downloading the audio and running one of the configured transcribers.
Transcriber options listed in .env.example are fast-whisper, bcut, kuaishou, mlx-whisper (Apple Silicon only) and groq. The default is fast-whisper with WHISPER_MODEL_SIZE=tiny, which the file describes as roughly 75 MB and fast on first start. The transcription text then goes to whichever LLM provider you configured (OpenAI, DeepSeek, Qwen and others are named), and the model produces the Markdown note. The provider keys are not stored in .env; the README says they are entered on the model provider page and saved into the SQLite database at ./backend/bili_note.db.
Installing BiliNote with Docker and generating a first note
The README recommends Docker and gives a prebuilt image on ghcr.io. Four named volumes matter: data holds the SQLite database and generated notes, config holds the LLM provider configuration and cookies, static holds the screenshots referenced by notes, and models caches Whisper weights so they are not re-downloaded. The README warns against mounting a single volume at /app/backend, because a named volume freezes the image contents from first start and later docker pull upgrades get masked by the old code.
docker pull ghcr.io/jefferyhcool/bilinote:latest
docker run -d -p 80:80 \
-v bilinote-data:/app/backend/data \
-v bilinote-config:/app/backend/config \
-v bilinote-static:/app/backend/static \
-v bilinote-models:/app/backend/models \
--name bilinote \
ghcr.io/jefferyhcool/bilinote:latestAfter the container starts, the README says to open http://localhost. If you prefer compose, copy the environment file first, because the README states that an empty .env leaves BACKEND_PORT and APP_PORT unset and the stack fails to start.
cp .env.example .env
docker-compose up --build -dThe compose file also exposes a GPU variant for NVIDIA hardware. Note the defaults: BACKEND_PORT=8483, APP_PORT=3015, and the backend health endpoint is /api/sys_health, polled by the container healthcheck every 30 seconds. The backend container is capped at mem_limit: 4g, and the compose comments say to raise that to 8g or more if you move WHISPER_MODEL_SIZE to medium or above, or the host OOM killer may take the container down during first model load.
Once the UI is up, the first real task is entering an LLM provider and key on the model provider page, then pasting a video link. Expect the first run to be slower than later ones if the Whisper model has not been cached yet; the changelog for v2.3.0 mentions a transcription readiness gate that blocks video tasks when the local engine model is not downloaded, instead of silently hanging on the first download.
Where BiliNote gets awkward: memory, model size and Chinese network assumptions
The default Whisper model is tiny, and the changelog for v2.2.0 records that it was changed from medium (about 1.5 GB) to tiny (about 75 MB), with an explicit confirmation prompt when you switch to a larger one. That is a sensible default for first-run experience, but it is a real accuracy trade-off. If your videos are noisy, accented or code-switch between languages, tiny will produce a transcript that the LLM then summarises into a confident-looking note built on a shaky base. You can move to base, small, medium or large from the audio transcription settings page, but the memory ceiling moves with it.
Network assumptions are the second friction point. The .env.example sets HF_ENDPOINT to https://hf-mirror.com by default and notes that host VPN or proxy settings do not automatically enter the container. The changelog for v2.3.0 added a global proxy setting that applies to model APIs, transcription endpoints such as Groq, and YouTube downloads, with HTTP_PROXY as a fallback. If you are outside China and your container cannot reach hf-mirror.com, you override HF_ENDPOINT to the official source. If you are inside China and docker.io is unreachable during a local build, the Dockerfile accepts a BASE_REGISTRY build argument and the compose file wires it to a variable, so you can point builds at a mirror.
A third boundary is the desktop client on Windows: the README states plainly that Windows users must run it in a path without Chinese characters. That is a concrete environment constraint, not a general warning.
BiliNote versus BibiGPT and other hosted summarisers
BibiGPT is the obvious comparison point, and the difference is architectural rather than feature-level. BibiGPT is a hosted product: you give it a link, it returns a summary, and the transcription, model choice and infrastructure are someone else's problem. BiliNote inverts that. You run the backend, you supply the LLM key, you choose the transcriber, and the notes live in a SQLite file on your own volume. The cost of that inversion is everything in the previous section: model downloads, memory limits, proxy configuration, and the fact that a broken container is yours to debug.
The project also offers its own hosted option at www.bilinote.app, which is worth treating as the honest alternative to self-hosting rather than a separate product. The README frames it exactly that way, as the path for people who do not want to deal with dependencies, proxies and model downloads. So the real decision is not BiliNote versus BibiGPT; it is BiliNote self-hosted versus BiliNote hosted versus BibiGPT, and the deciding factor is whether your video links and generated notes can leave your machine. If they can, the hosted routes are less work. If they cannot, only the self-hosted path is available to you.
Licence, maintenance and what an upgrade actually costs
BiliNote is MIT licensed, which permits commercial use and modification provided the copyright notice and permission notice are retained. The repository also advertises a paid one-on-one setup service and enterprise consulting in the README, which is separate from the licence and does not restrict the open source code. Nothing here is legal advice; if you plan to redistribute a modified build, read the LICENSE file in the repository root rather than the badge.
The last push to the default branch was on 2026-09-01, and the most recent release is v2.4.5 from 2026-08-25. The repository is not archived. Releases have been frequent through 2026, and the changelog records both features and fixes, including CI repairs around pnpm and Node version pinning. That pattern suggests the project is still being worked on, but the changelog also shows how much of the recent work is deployment resilience rather than core capability, which is a signal about where the rough edges have been.
Upgrade cost is dominated by the volume layout. The README's warning is the whole story: mount the four data subdirectories, not /app/backend, then docker pull and recreate the container and your configuration and history survive. There is a second trap in .env.example: VITE_* variables are build-time and get baked into the frontend bundle, so changing them requires docker-compose build frontend followed by docker-compose up -d, not a restart. Backend variables are runtime and only need up -d. If you skip that distinction, you will change a frontend URL and see no effect.
Editorial conclusion
Adopt BiliNote if you want video notes to stay on your own machine and you are willing to run Docker, pick a Whisper size and enter your own LLM key. Do not adopt it if you want a managed service with no moving parts, or if your host cannot spare roughly 4 GB of RAM for the backend container. Before committing, run the ghcr.io image with the four named volumes, check /api/sys_health through nginx, and confirm which transcriber you actually want, because the default is fast-whisper with the tiny model.
Frequently asked questions
How do I install BiliNote without building from source?
The README recommends pulling the prebuilt image from ghcr.io and running it with four named volumes for data, config, static and models. That avoids a local build entirely, which the README notes is also the easier path when docker.io is unreachable.
Where does BiliNote store my LLM API key?
Not in the .env file. The .env.example states that LLM API keys should be entered on the model provider page after deployment, and that they are saved into the SQLite database at ./backend/bili_note.db and persist with the container volume.
Which transcription engines can BiliNote use?
The .env.example lists fast-whisper, bcut, kuaishou, mlx-whisper (Apple Silicon only) and groq under TRANSCRIBER_TYPE, with fast-whisper as the default. The default Whisper size is tiny, about 75 MB, and larger sizes can be selected from the audio transcription settings page.
Does BiliNote work with YouTube as well as Bilibili?
Yes. The feature list names Bilibili, YouTube, Douyin, Kuaishou and local video files. The v2.0.0 changelog states that for YouTube, subtitles are fetched first and the audio download is skipped when subtitles are available.
Why does the BiliNote container keep restarting or showing unhealthy?
The Docker FAQ in the README suggests checking the backend logs first. The compose file also caps the backend at 4 GB of memory and notes that medium Whisper models or larger need 8 GB or more, otherwise the host OOM killer can terminate the container during first model load.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/jefferyhcool-bilinote)