# x-reader: one CLI for YouTube, Bilibili, X, WeChat and RSS content

> runesleo/x-reader turns a URL into structured text through a Python CLI, an optional MCP server and optional Claude Code skills. The public-first X chain and the Whisper path are the parts worth judging before you install it.

**runesleo/x-reader** — Universal content reader MCP Server for 10+ platforms

- Repository: https://github.com/runesleo/x-reader
- Website: https://leolabs.me
- Stars: 963 · Forks: 94
- Language: Python
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/runesleo-x-reader

## The problem x-reader solves, and who actually has it

Reading a link is easy. Reading a link from ten different places, in ten different shapes, is the part that eats an afternoon. A YouTube video gives you a page with no transcript in the HTML. A Bilibili video gives you something else. WeChat articles resist plain HTTP clients, Xiaohongshu resists them harder, X splits into oEmbed, articles and login-gated pages, and RSS is the one format that behaves. Each source has its own extraction trick, its own failure mode, and its own idea of what a document looks like.

x-reader's answer is a single entry point that returns a unified object. The README describes it as a universal content reader that fetches, transcribes and digests content from any platform, and the pitch is that you hand it a URL and get back structured content regardless of where that URL points. The intended audience is narrow and specific: people building AI agents or note-taking pipelines who need text, not people who want to browse. If you are a person who clicks links, a browser tab is better than this. If you are a script that needs the text of a link, the platform-specific branching is exactly the work you do not want to write twice.

The project is MIT licensed, written in Python, requires Python 3.10 or newer, and the last push to the repository was on 2026-08-29. It is not archived.

## How the dispatch chain works, platform by platform

The mechanism is a dispatcher plus per-platform fallback chains. The README's diagram shows the shape: any URL goes through platform detection, then content fetching, then a unified output. Detection is exposed as its own MCP tool, detect_platform(url), which tells you how much the rest of the pipeline will trust the URL before it tries.

Text extraction leans on Jina Reader as the general fallback, with API paths where a platform offers one. Bilibili has an API path. Telegram goes through Telethon. RSS goes through feedparser, which is a hard dependency in pyproject.toml alongside requests, python-dotenv, loguru and idna. That last one is not incidental: the dependency comment says idna is there for IDN and homograph attack protection in URL validation, and one of the example files is named 03_url_validator_blocks_localhost.py, so URL validation is treated as a security boundary rather than a formatting step.

X is the most layered chain, and the README spells it out in order: X oEmbed for fast public tweet text, FxTwitter for structured public fallback, Jina Reader for public Articles and long-form pages, generic Jina Reader for profiles and non-status pages, and finally Playwright with a saved session for login-required content. Public-first is a deliberate cost decision. The cheap paths handle most tweets, and the expensive browser path only runs when the earlier four fail.

Video and audio are separate. YouTube subtitles come through yt-dlp, with a Groq Whisper fallback when subtitles are not available. Full Whisper transcription for video and podcast, plus AI analysis, lives in the optional Claude Code skills rather than in the Python package. That split is the architectural fact that matters most: the pip-installed library is a fetcher, and the transcription depth lives in files you have to copy out of the repository.

## Installing x-reader and reading your first URL

The README recommends installing from GitHub rather than PyPI. The base install pulls only the core dependencies, so start there and add extras when a platform demands them.

```bash
pip install git+https://github.com/runesleo/x-reader.git
```

That gives you the x-reader command, wired through the project.scripts entry point in pyproject.toml. The first real use is a single URL. The README's own example fetches a WeChat article:

```bash
x-reader https://mp.weixin.qq.com/s/abc123
```

You should get structured content back rather than raw HTML. Multiple URLs in one call are supported: x-reader https://url1.com https://url2.com. To see what has already been fetched, run x-reader list, which reads the inbox file.

If your target platform needs a browser session, add the browser extra and install the Chromium binary Playwright drives:

```bash
pip install "x-reader[browser] @ git+https://github.com/runesleo/x-reader.git"
playwright install chromium
```

Then perform the one-time login the README documents for Xiaohongshu, and the equivalent for gated X pages:

```bash
x-reader login xhs
x-reader login twitter
```

For Telegram, install the telegram extra and fill TG_API_ID and TG_API_HASH from my.telegram.org. For Whisper transcription, export GROQ_API_KEY, which the README says you can obtain free from the Groq console. Configuration otherwise lives in .env, copied from .env.example, and the documented keys are TG_API_ID, TG_API_HASH, GROQ_API_KEY, GEMINI_API_KEY, OUTPUT_DIR, INBOX_FILE and OBSIDIAN_VAULT. OUTPUT_DIR defaults to disabled, so Markdown output does not appear until you set it. OBSIDIAN_VAULT, when set, writes to 01-收集箱/x-reader-inbox.md.

## Using it as a Python library or an MCP server

The library surface is small. The README's example imports UniversalReader from x_reader.reader, constructs it, and awaits read() on a URL, then prints content.title and the first 200 characters of content.content.

```python
import asyncio
from x_reader.reader import UniversalReader

async def main():
    reader = UniversalReader()
    content = await reader.read("https://mp.weixin.qq.com/s/abc123")
    print(content.title)
    print(content.content[:200])

asyncio.run(main())
```

That is the whole adoption story for library users, and it is the right size. The unified schema in x_reader/schema.py is what makes the output portable across platforms.

The MCP path is different, and the README is explicit that mcp_server.py is not included in pip install. You clone the repository, install the mcp extra, and run the server directly:

```bash
pip install -e ".[mcp]"
python mcp_server.py
```

The server exposes read_url(url), read_batch(urls), list_inbox() and detect_platform(url). read_batch fetches multiple URLs concurrently, which is the reason to choose MCP over shelling out to the CLI in a loop. Client configuration is a standard MCP server block pointing at the script path:

```json
{
    "mcpServers": {
        "x-reader": {
            "command": "python",
            "args": ["/path/to/x-reader/mcp_server.py"]
        }
    }
}
```

The version constraint here deserves attention. The README states the server currently targets FastMCP 1.x, and that the mcp and all extras pin mcp<2. It also states plainly that moving to MCP 2.x requires a server migration rather than removing the version cap. If you are standardising on MCP 2.x today, this project is behind you, and the README does not promise a date.

## Where x-reader breaks, and what it will not do

The honest limitation is that the deepest capabilities are not in the pip package. Full Whisper transcription for video and podcast, and AI-powered content analysis, live in the Claude Code skills directory. The README says so directly: the Python layer handles text fetching and YouTube subtitle extraction, while the skills add the rest. Installing from pip gets you a fetcher, not a transcriber, and the skills require cloning the repository and copying directories into a skills path you define yourself.

Second, several platforms are only reachable through a browser session you establish manually. Xiaohongshu needs x-reader login xhs once. Login-required X pages and Articles need x-reader login twitter. That session is state on your machine, and the README notes that local X cookies stay local by default. If you want Jina to use your saved X session for gated Articles, you have to opt in explicitly with X_READER_ALLOW_EXTERNAL_SESSION_COOKIES=1. That is a sensible default, but it also means the gated-Article path does not work out of the box in a fresh container.

Third, the X chain is a chain. Five fallbacks means five things that can change without warning, because oEmbed endpoints, third-party mirrors and anti-scraping behaviour are not under this project's control. The same applies to WeChat and Xiaohongshu, where Playwright is the fallback precisely because plain fetching is blocked. Expect to maintain credentials and re-run logins rather than set and forget.

Finally, the README does not document rollback, rate limits, retry policy or how the inbox deduplicates entries, even though example files reference dedup and clearing old entries. If those behaviours matter to you, read x_reader/ and the examples directory rather than the README.

## Alternatives: x-reader against a general web extractor

The obvious comparison is a general-purpose reader such as Jina Reader used directly. The difference is not quality of extraction; it is the layer above it. Jina Reader takes a URL and returns readable text, and x-reader actually uses it as the fallback for generic pages and for X Articles. Calling Jina directly means no platform detection, no per-platform chain, no unified schema, no inbox, and no MCP tools. It also means you write the branching yourself for the platforms where Jina is not enough, which is exactly the Bilibili API path, the Telethon path, the yt-dlp subtitle path and the Playwright sessions.

A second comparison is yt-dlp for video and a dedicated scraper per site. yt-dlp is more capable than x-reader will ever be at video metadata and format selection, and x-reader uses it for subtitles rather than reimplementing it. Where x-reader differs is that the video result lands in the same object as a WeChat article, so downstream code does not branch on source. If your pipeline only ever touches YouTube, yt-dlp plus a transcript tool is less machinery. If it touches YouTube, Bilibili, X, WeChat, Telegram and RSS, the unified output is the entire value proposition.

The trade-off is dependency weight and failure surface. A single Jina call has one failure mode. x-reader has one per platform, plus a browser binary, plus optional API keys for Telegram and Groq. That is more to install and more to debug, and the README's three-layer structure exists because the author knows not everyone wants all of it.

## Maintenance cost and licence implications

The repository is not archived and the last push was on 2026-08-29. Two releases are listed: v0.2.1 on 2026-03-25, described as an ARM64 compatibility fix, and v0.2.2 on 2026-05-25. The pyproject.toml still declares version 0.2.0, so if version strings matter to your tooling, trust the release tags over the package metadata.

The upgrade cost concentrates in three places. The MCP pin, mcp<2, means an MCP 2.x migration is a server rewrite rather than a constraint bump, and the README says as much. The browser extra pulls Playwright, which means a Chromium binary and its updates ride along with your dependency updates. The per-platform chains mean a platform change can break extraction without any change in this repository, so an upgrade is not the only event that can break your pipeline.

On licensing: the project is MIT, declared in both the LICENSE file and pyproject.toml. MIT is permissive and places few conditions on reuse. What the licence does not cover is the content you fetch. Scraping WeChat, Xiaohongshu or X is governed by those platforms' terms and by local law, and the README does not discuss this. That is a question for your own counsel, not for this review.

## Conclusion

Adopt x-reader if you already pipe URLs into an agent or a notes vault and want one Python entry point instead of five scrapers, and if you accept that the strong paths (X gated pages, XHS, Whisper transcription) depend on a saved browser session or a Groq key. Do not adopt it if you need a hosted service, a stable MCP 2.x server, or any guarantee about platforms whose anti-scraping behaviour changes without notice. Before committing, run x-reader detect_platform on the exact URLs you care about, check whether your platform needs x-reader login first, and confirm that the pip package alone covers your case, because mcp_server.py and skills/ only ship in the cloned repository.

## FAQ

### What is x-reader and what does it read?

x-reader is a Python content reader that takes a URL, detects the platform, fetches the content and returns it in a unified schema. The README lists YouTube, Bilibili, X/Twitter, WeChat, Xiaohongshu, Telegram, RSS, Xiaoyuzhou, Apple Podcasts and generic web pages.

### How do I install x-reader?

The README recommends installing from GitHub with pip install git+https://github.com/runesleo/x-reader.git. Optional extras cover Telegram, browser fallback via Playwright and all optional dependencies combined.

### Does x-reader work as an MCP server?

Yes, but mcp_server.py is not included in the pip install, so you have to clone the repository, install the mcp extra and run python mcp_server.py. The README states the server currently targets FastMCP 1.x and that the extras pin mcp<2.

### Why does Xiaohongshu or a gated X page fail to read?

Those platforms need a saved browser session. The README documents x-reader login xhs for Xiaohongshu and x-reader login twitter for X Articles and login-required pages, which requires the browser extra and playwright install chromium.

## Sources

- [License: MIT](https://github.com/runesleo/x-reader/blob/main/LICENSE)
- [Project website](https://leolabs.me)
- [README](https://github.com/runesleo/x-reader/blob/main/README.md)
- [Releases](https://github.com/runesleo/x-reader/releases)
- [runesleo/x-reader on GitHub](https://github.com/runesleo/x-reader)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/runesleo-x-reader
