Chubby Skills: 14 Agent Skills for Ingesting Chinese Content into a Local Knowledge Base
把中文全渠道内容(抖音 / B站 / 小红书 / 公众号 / X / 播客)采集进个人知识库的 13 个 AI Skill:图文存图、视频转文字稿、字幕优先免 GPU,附带知识库 MCP server。 | Ingest Chinese content into your personal knowledge base — image/video routing, subtitle-first transcription, and a KB MCP server.
At a glance
- What is it?
- Chubby Skills is a MIT-licensed collection of 14 Python-based agent skills and command-line tools that pull content from Chinese platforms including Bilibili, Douyin, WeChat, Xiaohongshu, and X into a local Markdown knowledge base. It converts videos, podcasts, articles, and local documents into searchable, source-cited Markdown files, and includes a knowledge base MCP server for agent access.
- Who is it for?
- Chubby Skills is well suited for Chinese content creators and researchers who collect material from multiple Chinese platforms and want a local, searchable, source-cited knowledge base that agents can query. It is not a general-purpose web scraper: its platform support is explicitly limited and dependent on cookies, subtitles, and platform access conditions that can change.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 13 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What Chubby Skills Does and the Workflow It Enables
The README describes a common problem for content creators who research using Chinese video platforms, podcasts, and article feeds: the content is watched or read and then forgotten, because saving platform URLs is not the same as preserving the actual text and source context.
Chubby Skills solves this by converting platform content into Markdown files stored in a local vault directory. Each Markdown file retains its source URL or file URI, and attachments referenced in the document are copied locally. The output is a vault that a keyword search or semantic search can query, and that can generate an evidence brief (a Markdown or JSON file containing quoted excerpts, line numbers, source URLs, and SHA-256 hashes) for use in writing or agent workflows.
The full pipeline as described in the README is: a link or local document enters the system, becomes a Markdown file with local attachments, gets indexed, can be searched by keyword, and can be exported as a cited brief that includes verbatim quotes and source metadata. The brief command runs locally and does not call a cloud model to judge the content; it only extracts and organises what is already in the vault.
The 14 Skills and Their Coverage
The project ships 14 distinct skills, each responsible for a specific platform or processing type. They are grouped by function in the README.
The video skills cover Bilibili, YouTube, Douyin (TikTok China), TikTok, Weibo, and Zhihu. For Bilibili and YouTube, the ingestion path is subtitle-first: if a subtitle track exists, it is used directly without any local model inference. When no subtitle is available, the skill falls back to local video transcription, which requires additional dependencies.
The podcast skill handles single episodes, RSS feeds, and local audio files. It uses faster-whisper for local transcription with the small model as default. An experimental cloud transcription backend via Atlas Cloud or MuAPI was added in v0.13.0. Cloud transcription is opt-in and requires configuring credentials; it may incur charges.
The image and text skills cover WeChat public account articles (with optional PDF handling), Xiaohongshu image and video posts, and X (Twitter) posts including images and videos.
The content-enrich skill is an optional post-processing step that adds summaries, key points, and tags using a DeepSeek API key. It is separate from ingestion and can be skipped with the --no-enrich flag.
The knowledge-base-management skill covers document import, index management, evidence brief generation, archiving, and the MCP server.
The two workflow skills handle industry intelligence scanning (industry-intelligence-radar) and learning note automation (learning-notes-automation).
Initial Setup: Clone, Initialise, Import, and Search
The README recommends Python 3.11 or 3.12 on macOS or Linux. The quick start path requires no third-party packages beyond what Python already includes, making it useful for a first test without a full dependency install.
git clone https://github.com/chubbyguan/chubbyskills.git
cd chubbyskills
python3 -m venv .venv
source .venv/bin/activate
python3 tools/chubby.py init --vault "$PWD/creator-vault"The init command creates a chubby.yaml configuration file and the vault directory structure. Once initialised, you can import a local Markdown file:
python3 tools/chubby.py import demo-input/notes.md --no-enrichAfter importing, search and export a brief:
python3 tools/chubby.py search "内容复用"
python3 tools/chubby.py brief --topic "内容复用" \
--output "$PWD/creator-vault/30_Output/brief.md"
python3 tools/chubby.py status --latestThe search command queries the vault index. The brief command exports a Markdown and JSON file containing excerpts with line numbers and source references. The status command shows recent task results.
For PDF documents, the importer extracts the text layer only; scanned PDFs without a text layer require OCR preprocessing by a separate tool before import. The README documents PDF import as:
python3 tools/chubby.py import "/path/to/report.pdf" \
--source-url "https://example.org/original-report" --no-enrichPlatform Ingestion and the Dependency Table
Platform ingestion requires additional dependencies beyond the Python standard library. The README provides a processing path table:
- Markdown, TXT import, keyword search, and evidence briefs: Python standard library only. - X and Xiaohongshu image posts: Python standard library; login state and platform restrictions may affect ingestion. - Bilibili and YouTube subtitles: `python3 -m pip install yt-dlp`. - Local video transcription: `bash setup.sh video`; requires ffmpeg and downloads a local model on first run. - Local podcast transcription: `bash setup.sh podcast`; uses faster-whisper with the small model by default. - WeChat public account articles and PDF handling: `bash setup.sh wechat`. - General PDF text layer import: `python3 -m pip install 'pymupdf>=1.24'`.
The YouTube subtitle path as shown in the README:
bash setup.sh light
python3 -m pip install yt-dlp
python3 tools/chubby.py doctor --platform youtube
python3 tools/chubby.py ingest "<your-link>" --no-enrichThe doctor command checks whether the platform's current ingestion dependencies are satisfied. The README notes that platform fetching is affected by cookies, subtitle availability, regional restrictions, and page changes. Ingestion failures should be investigated using the platform-status and platform-fallbacks documentation in the docs/ directory.
Batch Ingestion, Retry, and Deduplication
For ingesting multiple links, the README provides a queue-based workflow. You place links line by line in a text file and run:
python3 tools/chubby.py run --queue inbox/links.txt --no-enrich
python3 tools/chubby.py status --failed
python3 tools/chubby.py retry --all-failedThe run command processes each link and writes task state to .chubby/runs.jsonl. The status command shows which tasks succeeded and which failed. The retry command re-attempts all failed tasks.
Deduplication is built in: identical source, processing parameters, and destination will reuse an existing valid output rather than re-ingesting. The README notes that changing the source file or referenced attachments triggers re-import. Passing --refresh forces re-processing and keeps the old version.
The README also notes that task IDs and completion results for cloud transcription are persisted, so an interrupted polling process or a process crash can be recovered. The --resubmit flag creates a new cloud transcription task, which may incur a new charge.
Installing Skills into an Agent and the MCP Server
Individual skills can be installed into a Codex, Claude Code, OpenCode, or compatible agent client using the install_skill.py script:
python3 tools/install_skill.py knowledge-base-management --dest ~/.codex/skillsTo see the full list of skills available for installation:
python3 tools/install_skill.py --listThe install script bundles repository-internal dependencies into a self-contained skill directory so it can be moved without the full repository. The README warns against downloading individual skill directories without the shared chubby_common/ module, as this will miss the common modules.
The MCP server provides six tools: search, semantic search, read original, recent notes, re-index, and statistics. To start it:
python3 -m pip install -r knowledge-base-management/requirements-mcp.txt
python3 tools/mcp_smoke.py --jsonThe smoke test verifies that the server starts and responds to tool calls against a temporary vault. Connecting it to your agent requires configuring the server command and the VAULT_DIR environment variable.
Limitations and Data Boundaries
The README documents the data flow and network boundaries explicitly.
Document import, keyword search, semantic-lite search, and brief export run locally without any network calls. Platform ingestion accesses the source platform (Bilibili, YouTube, WeChat, etc.) to fetch content. Local video and podcast transcription runs inference on-device; the first run may download a model. Content enrichment with summaries or tags sends content to the DeepSeek API if configured. OpenAI-based vector search sends content to OpenAI's embedding API if configured.
Platform access depends on factors outside the project's control. The README names four specific constraints: cookie state, subtitle availability, regional restrictions, and page structure changes. The docs/ directory contains live-verification.md and platform-status.md, which are maintained to reflect current ingestion reliability per platform.
PDF import extracts only the text layer; OCR is not included. Scanned documents without a text layer cannot be imported directly. Local file attachments referenced by a document are copied into the vault, but missing or out-of-bounds local attachments cause an error rather than a silent skip.
Cloud transcription for podcasts is experimental as of v0.13.0. The README states that real-world cloud transcription results have not yet been fully validated. This makes the local transcription path (faster-whisper) the more reliable default for production use.
The podcast automatic download path only accepts public HTTP or HTTPS direct URLs. It does not follow redirects and does not use a proxy. Podcast feeds or episodes that require redirect chains or proxy access must be downloaded manually first and passed as local files to the ingest command.
The project's README is written in Chinese, with a parallel English README (README.en.md) available at the top level. Documentation in the docs/ directory is in Chinese. Engineers who do not read Chinese will need to rely on the English README and may find that some operational details documented in the Chinese docs are not fully covered in the English version.
Editorial conclusion
Chubby Skills is well suited for Chinese content creators and researchers who collect material from multiple Chinese platforms and want a local, searchable, source-cited knowledge base that agents can query. It is not a general-purpose web scraper: its platform support is explicitly limited and dependent on cookies, subtitles, and platform access conditions that can change. Before building a workflow around it, run the doctor command against each platform you intend to use and verify current ingestion status in the platform-status documentation.
Frequently asked questions
What platforms does Chubby Skills support for content ingestion?
Chubby Skills includes skills for Bilibili, YouTube, Douyin, TikTok, Weibo, Zhihu, WeChat public accounts, Xiaohongshu, and X (Twitter). Each platform skill has its own dependency requirements; Bilibili and YouTube use subtitle-first paths with yt-dlp, while video and podcast transcription paths require additional setup via setup.sh.
Does Chubby Skills send my content to external APIs?
By default, document import, keyword search, and brief export run entirely locally. Content enrichment with summaries is optional and requires configuring a DEEPSEEK_API_KEY. OpenAI-based semantic search is also optional and requires an OPENAI_API_KEY. Platform ingestion accesses the source platform to fetch content but does not send it to third-party services.
How do I install a single Chubby Skills skill into my agent?
Run python3 tools/install_skill.py <skill-name> --dest <your-agent-skills-dir> from the chubbyskills repository root. The installer bundles repository-internal dependencies into a self-contained directory. Use python3 tools/install_skill.py --list to see all available skill names.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/chubbyguan-chubbyskills)