BiliSum: local-first AI summaries for Bilibili, YouTube and local video
为 Bilibili、YouTube 及本地视频提供 AI 视频摘要和知识库.AI video summarizer and knowledge base for Bilibili, YouTube and local videos.
At a glance
- What is it?
- BiliSum is a Python and Electron application that transcribes a video, turns it into notes, mind maps and a searchable knowledge base, and keeps the data on your machine. It is a good fit if you watch Bilibili and want the output in Obsidian; it is the wrong tool if you need a hosted service or a Linux desktop.
- Who is it for?
- Adopt BiliSum if you watch Bilibili regularly, want transcripts and notes exported as Markdown or Obsidian, and are willing to configure one transcription provider plus one LLM before the first run. Skip it if you need a Linux desktop build or a hosted multi-user service: the desktop packages are Windows and macOS only, and the Docker image is a single-container deployment with one access token.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 20 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What BiliSum solves, and for whom
Watching a two-hour Bilibili video to find one explanation leaves you with nothing you can search later. BiliSum takes a video URL or a local file, produces a transcript, and builds structured artifacts on top of it: a structured summary, a chapter timeline, text notes, illustrated notes, a mind map, and a knowledge base that can answer questions across videos. The README describes the pipeline as Bilibili, YouTube and local video into transcription, then text notes, then illustrated notes, then mind map, then knowledge base.
The intended user is someone who already takes notes and wants video to feed the same habit. The export targets give this away: Markdown and Obsidian, with notes and screenshots packaged together. If your notes live in Obsidian, the output format is the reason to look at this project rather than a browser extension that shows a summary in a sidebar.
The secondary audience is agents. BiliSum ships a CLI and a skill named bilisum-video-understanding so Codex, Claude or Cursor can request a full summary, a brief, a transcript or a task status. That is a different use pattern from the desktop app: the CLI prefers to talk to a running desktop instance and share its tasks, configuration and knowledge base, and falls back to its own environment when no desktop app is running.
The pipeline, and where the local-first claim actually holds
The stack table in the README names the components: an Electron, React, TypeScript and Vite desktop client, a FastAPI and SQLite backend service, yt-dlp plus ffmpeg for video handling, OpenAI-compatible, Anthropic and local model endpoints for inference, embedding retrieval plus an LLM agent for the knowledge base, and PyInstaller, electron-builder and Docker for distribution.
The local-first claim is about storage, not about inference. Notes, indexes and configuration are saved on the machine, and the README states that local Embedding and local LLM options exist. But the default environment template points ASR at SiliconFlow and the LLM at a hosted endpoint, so a stock install sends audio or text to a third party unless you change the provider. Read the local-first claim as "your knowledge base is a SQLite file you own", not as "no data leaves the machine".
Transcription has four routes according to the README: SiliconFlow ASR, multimodal ASR, local Whisper, and FunASR (QwenASR). The environment template shows the Whisper path concretely with VIDEO_SUM_WHISPER_MODEL, VIDEO_SUM_WHISPER_DEVICE and VIDEO_SUM_WHISPER_COMPUTE_TYPE, defaulting to tiny on CPU with int8. That default is a speed choice, not an accuracy choice, and a tiny model on CPU is the first thing to change if transcripts come out poor.
Illustrated notes are the part with a real design decision behind them. A vision model reads both the original notes and extracted video frames, then reorganizes the article so images sit next to the paragraph they illustrate rather than being dumped at the end. The README says the visual model is configured separately from the text model and supports OpenAI, Anthropic, compatible interfaces and custom endpoints. Separating the two is sensible: you can run a cheap text model for summaries and a stronger vision model only for the frames.
Installing BiliSum: desktop download, CLI, or Docker on port 3838
The desktop route is a download. The README says to get the Windows or macOS installer from GitHub Releases, then configure one transcription service and one LLM in the settings page on first launch. It also recommends the built-in QR login for Bilibili when you hit rate limiting, which is a hint that cookie-based access to Bilibili is fragile.
The CLI is published to npm and installs globally:
npm install -g bilisum
bilisum summarize "https://www.bilibili.com/video/BV1xxxx"
bilisum brief "https://www.bilibili.com/video/BV1xxxx"
bilisum transcribe ./demo.mp4 --output transcript.txtsummarize produces the full summary, brief produces the short one, and transcribe writes a transcript to the path given with --output. The README notes the CLI prefers a running desktop instance and shares its tasks, configuration and knowledge base with it.
If no desktop app is installed, initialize a separate environment. This needs Python 3.12:
bilisum env setup
bilisum --settingbilisum env setup creates the standalone environment and bilisum --setting opens the settings page where the ASR and LLM credentials go.
The Docker path is a single container exposing port 3838 with a data volume:
docker pull lycohana/bilisum:latest
docker run --rm -p 3838:3838 -v bilisum-data:/data \
-e VIDEO_SUM_ACCESS_TOKEN=your-token \
lycohana/bilisum:latestAfter that, the README says to open http://127.0.0.1:3838. The image sets VIDEO_SUM_DOCKER=1, binds VIDEO_SUM_HOST to 0.0.0.0, stores everything under /data, and keeps the SQLite database at /data/video_sum.db. Because the host binding is 0.0.0.0 inside the container, the access token is the only thing standing between the port and anyone who can reach it, so do not publish 3838 to a network without it.
For a source checkout the README lists Python 3.12, Node.js 20+ and uv:
uv sync --python 3.12 --all-packages
npm install --prefix .\apps\desktop
npm run devThe workspace layout matches this: a uv workspace with apps/service, packages/core and packages/infra as members, and a root package.json whose dev script delegates to apps/desktop.
Configuration keys you will actually touch
The .env.example file is the clearest picture of what is configurable. Transcription is selected with VIDEO_SUM_TRANSCRIPTION_PROVIDER, which defaults to siliconflow, and the SiliconFlow path uses VIDEO_SUM_SILICONFLOW_ASR_BASE_URL, VIDEO_SUM_SILICONFLOW_ASR_MODEL and VIDEO_SUM_SILICONFLOW_ASR_API_KEY, with TeleAI/TeleSpeechASR as the default model. The LLM side is VIDEO_SUM_LLM_ENABLED, VIDEO_SUM_LLM_BASE_URL, VIDEO_SUM_LLM_MODEL and VIDEO_SUM_LLM_API_KEY, and the template ships with a placeholder key that must be replaced.
Two keys deserve attention. VIDEO_SUM_YTDLP_COOKIES_FILE exists because YouTube increasingly requires a signed-in session for some videos; pointing it at a cookies file is the documented escape hatch. VIDEO_SUM_TWELVELABS_SUMMARY_ENABLED is off by default and, per the comment in the template, applies only to local videos, adding visual information that subtitles do not carry through the Twelve Labs Pegasus model.
The README says configuration can be set through the settings page in both the desktop app and the CLI standalone environment, and points to docs/configuration.md for environment variables, Docker deployment and configuration precedence. Precedence is the part worth reading before you deploy, because a Docker container and a desktop settings page can disagree about which value wins.
Where BiliSum is the wrong tool
Platform coverage is the first limit. The README badge and the packaging scripts both say Windows and macOS; there is no Linux desktop package, and the packaging scripts in package.json are package:win and package:mac variants only. On Linux the realistic option is the Docker image or the CLI, which means no desktop UI, no built-in QR login for Bilibili, and no in-app update.
The second limit is that this is a personal tool, not a multi-tenant service. The Docker deployment exposes one access token, one SQLite database and one data directory. There is no mention of user accounts, quotas or per-user isolation. Putting it in front of a team means sharing one knowledge base and one set of API keys.
Third, the quality ceiling is set by components BiliSum does not control. yt-dlp breaks when a site changes its player, and Bilibili rate limiting is common enough that the README calls it out and recommends the QR login as the workaround. If your workflow depends on summarizing a hundred videos a day, the failure mode is not the summarizer, it is the downloader and the ASR provider's rate limits.
Finally, a local model is not a free lunch. Running Whisper locally on CPU with the default tiny model trades accuracy for speed, and running a local LLM and local embeddings for the knowledge base shifts the cost from an API bill to your own hardware. The README presents local models as an option, not as the default path, and the default path is the hosted one.
How it compares with a transcript-first tool
The obvious alternative is a transcript-first tool such as a Whisper front end that produces a subtitle file and stops there, or a hosted summarizer that returns a paragraph in a web page. The difference is where the structure lives.
A transcript tool gives you text and leaves the organization to you. BiliSum treats the transcript as an intermediate artifact and keeps going: chapter timeline, structured summary, text notes, illustrated notes, mind map, then an embedding index over the notes so you can ask questions across videos. That extra layer is the product. It is also the reason setup is heavier, because illustrated notes need a vision model and the knowledge base needs embeddings plus an LLM agent.
Against a hosted summarizer, the trade runs the other way. A hosted service asks for a URL and returns a summary with no installation. BiliSum asks you to install an Electron app or run a container, then configure at least two providers, and in exchange the notes, indexes and configuration stay in files you control and can be exported to Obsidian. If you never open the notes again, the hosted summarizer is the better deal. If the notes are the point, the local store is.
The CLI and skill change the comparison once more. A transcript tool usually has no agent interface; BiliSum publishes the bilisum-video-understanding skill so an agent can pick between a full summary, a brief, a transcript or a status check, and the CLI can reuse a running desktop instance's knowledge base rather than building a second one.
Maintenance, licence and upgrade cost
The repository is not archived, and the last push was on 2026-09-10, the same day as the v1.21.1 release. The release history shows v1.21.0-alpha.1 on 2026-08-22, v1.21.0 on 2026-08-23 and v1.21.1 on 2026-09-10, so the project is moving at a pace where a tagged release can land within weeks of the previous one. For an adopter that means reading release notes before upgrading rather than pinning once and forgetting.
Upgrade cost depends on the install route. The desktop app has in-app updating according to the README. The CLI is an npm global package, so it follows npm versioning. The Docker route is a tag pull, and because all state lives under /data, upgrading the image does not touch the knowledge base, the cache, the tasks directory or the SQLite file. That separation is the strongest argument for the container deployment: the upgrade is a pull and a restart, and the data volume survives it.
The licence is MIT, which permits commercial use, modification and redistribution provided the copyright notice and permission notice are included. That is a permissive licence, but it covers BiliSum's code only. The models and APIs you plug in carry their own terms, and the .env.example defaults point at SiliconFlow and a DashScope-compatible endpoint, whose terms are separate from this repository. If you plan to summarize third-party video at scale, the licensing question is about the video source and the model provider, not about BiliSum's MIT grant. This is not legal advice.
Editorial conclusion
Adopt BiliSum if you watch Bilibili regularly, want transcripts and notes exported as Markdown or Obsidian, and are willing to configure one transcription provider plus one LLM before the first run. Skip it if you need a Linux desktop build or a hosted multi-user service: the desktop packages are Windows and macOS only, and the Docker image is a single-container deployment with one access token. Verify first that the model you intend to use is reachable from the machine that runs the summarization, because the default configuration points at SiliconFlow for ASR and at a DashScope-compatible endpoint for the LLM, and neither will work without a key.
Frequently asked questions
Does BiliSum work with YouTube as well as Bilibili?
Yes. The repository description and README both cover Bilibili, YouTube and local videos, and the import options list Bilibili, YouTube and local files in mp4, mkv, mov and webm. The .env.example includes VIDEO_SUM_YTDLP_COOKIES_FILE for supplying cookies, which matters when YouTube requires a signed-in session.
Can I run BiliSum on Linux?
There is no Linux desktop package. The README badges and the packaging scripts cover Windows and macOS only, so on Linux the options are the Docker image on port 3838 or the npm CLI with bilisum env setup, which requires Python 3.12.
Do I need an API key to use BiliSum?
The default configuration does. The environment template sets VIDEO_SUM_TRANSCRIPTION_PROVIDER to siliconflow with an ASR key field and enables the LLM with VIDEO_SUM_LLM_API_KEY, which ships as a placeholder that must be replaced. The README also lists local Whisper and local LLM options, so a fully local setup is possible if you configure it that way.
Where does BiliSum store my notes and knowledge base?
In the Docker deployment everything lives under /data, with the database at /data/video_sum.db and separate cache and tasks directories. The README states that notes, indexes and configuration are saved on the local machine, and that notes can be exported as Markdown or Obsidian format with screenshots included.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/lycohana-bilisum)