Douyin_TikTok_Download_API: a self-hosted scraper for Douyin and TikTok
🚀 Self-hosted TikTok & Douyin scraper and no-watermark video downloader — async REST API, MCP server, CLI and web console for posts, profiles, comments and playlists. Self-healing identity pool, PostgreSQL archive, one docker compose up. 抖音、TikTok 数据采集与无水印视频下载 API,自托管,支持 MCP 调用与 Docker 一键部署。
At a glance
- What is it?
- Evil0ctal's project bundles a REST API, MCP server, CLI and web console around a self-maintaining identity pool and a PostgreSQL archive. The Docker path is the one that works; the Python path needs a secret you must not lose.
- Who is it for?
- Adopt it if you need Douyin or TikTok posts, profiles, comments and playlists behind your own API, and you can run PostgreSQL, Redis and a container host. Skip it if you need a managed SLA, or if you cannot legally store the media you pull.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 7 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 26, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The gap this fills between a scraper script and a platform API
Douyin and TikTok do not publish a general read API for posts, authors, comments and search. The usual substitute is a throwaway Python script that signs one request, works for a week, and dies when the signature changes. Douyin_TikTok_Download_API takes the opposite position: the signature logic, the cookie material and the request scheduling live in a long-running service you host, and the surface you call is an HTTP API rather than a function in someone's notebook.
The README frames the pitch as "no signup, no quota, nobody else in the path." That is the honest summary of who it is for. If you are building a research dataset, an archive, a moderation queue or an internal dashboard over Douyin or TikTok content, and you would rather not send every request through a third-party data vendor, this is the shape of tool you want. It is not aimed at someone who wants to paste one link into a website and get one MP4. The demo instance at demo.douyin.wtf covers that case.
The project also fetches image albums and video without a watermark, and the README is specific about how: it "picks the clean stream the platform already publishes rather than stripping anything." That distinction matters if you care about not re-encoding media, and it also means the availability of a clean stream is the platform's decision, not the project's.
Identity pool, scheduler and PostgreSQL: the moving parts
The architecture visible in the repository is a FastAPI application (fastapi, uvicorn, pydantic) with SQLAlchemy over asyncpg for PostgreSQL and Redis for the parts that need to be fast and shared. Alembic sits alongside for schema migrations, which tells you the archive is a real schema with versions, not a JSON dump. structlog handles logging, and the dependency list includes wreq and httpx for outbound requests and cryptography plus argon2-cffi for secrets and password hashing.
The component that gives the project its character is the identity pool. The README calls it "an identity pool that maintains itself," and the .env.example explains what it holds: every stored cookie jar and proxy URL is encrypted with a single secret. The scheduler, the identities, the playground, the library, the downloads, the API docs and the MCP server all appear in the same web console, which suggests the pool and the scheduler are first-class resources with their own state rather than configuration baked into the process at boot.
The .env.example is explicit that bootstrap settings are the exception, not the rule: "Everything else is seeded into the database on first init and edited from the console afterwards." So the operational model is that you start the stack with a small environment, then tune identities and scheduling from the UI. If you expect to drive everything from environment variables in a GitOps pipeline, the configuration documentation is the thing to read before you commit, because the file itself points at docs/design/10-configuration.md for the rest.
Installing Douyin_TikTok_Download_API with docker compose
The README leads with one command: `docker compose up`. The repository's Makefile wraps that in a named compose project so stray containers do not accumulate, and that wrapper is the safer entry point because it also handles the wait-for-healthy behaviour.
Copy the example environment file to the repository root first, since the Makefile's compose invocation reads it there:
cp .env.example .env
openssl rand -base64 48The second command generates the value for the secret described in .env.example. The process refuses to start if that secret is missing or shorter than 32 characters, so this is not optional. Put the output in .env before you bring anything up.
make upThis runs `docker compose -p dtk -f docker/compose.yml up -d --wait` and returns once the services report healthy. To watch what happened, `make logs` follows the stack. `make down` stops it, and `make clean` removes every dtk container, network and volume.
There is a trap in .env.example that is worth reading twice. The file has two halves. Values above the marked line reach the containers through `env_file: ../.env`, resolved relative to docker/compose.yml. Values below it are interpolated into compose.yml itself, and that substitution reads `docker/.env` or the shell, not the repository root. The file states the failure mode plainly: "an instance built without the browser pin comes up healthy and mints nothing." If your stack starts and the identity pool stays empty, that split is the first thing to check. The documented workarounds are to pass the file explicitly with `COMPOSE_ENV_FILES=.env docker compose -p dtk -f docker/compose.yml up -d`, or to set a single value inline such as `DTK_IMAGE_TAG=5.0.3 docker compose -p dtk -f docker/compose.yml up -d`.
The Python path exists too. pyproject.toml requires Python 3.12 or newer and below 3.14, the Makefile's `install` target runs `uv sync --all-extras`, and the console script is `dtk`, mapped to `dtk.cli.main:app`. That gives you a CLI, but it does not remove the PostgreSQL and Redis requirement, which is why the compose route is the one to start with.
What breaks, and when this is the wrong tool
The central risk is not in this repository. The project's job is to reproduce what the platforms expect, and the platforms change. The pyproject.toml test markers acknowledge this directly: there is a `live` marker described as tests that hit the real platform and "never run in CI gating." Read that as the maintainers telling you the upstream contract is not covered by the pipeline that gates releases. A green CI badge on this repository does not mean the Douyin endpoint you depend on still returns what it returned last month.
Second, the identity pool is a dependency you now operate. It needs cookie material and, per .env.example, proxy URLs. Those are encrypted with one secret, and the file is blunt about the consequence of losing it: "Losing it does not lock you out of the console - it makes the identity pool unreadable." Back it up with the database, not separately. An operator who rotates that value without a plan has an archive that still opens and an identity pool that does not.
Third, this is the wrong tool when the answer needs to be a contract. If you need a vendor to answer a ticket, guarantee a request rate, or indemnify you, a self-hosted scraper does none of that. It is also the wrong tool for one-off downloads. The demo instance exists precisely so you do not deploy a database and a scheduler to fetch a single video.
Finally, the README's own rate limit note for the public demo (30 requests per 10 seconds) is a reminder that the unthrottled version is the one you host, and the throttling you get is the one you build.
How it differs from calling a social data API vendor
The obvious alternative is a commercial data API such as the one advertised in the project's own sponsor block, TikHub.io, which offers APIs across Douyin, Xiaohongshu, TikTok, Instagram, YouTube and Twitter. The difference is not features, it is where the failure lands. With a vendor, you get a documented endpoint, a quota and a support channel, and when the platform changes its signature the vendor absorbs the breakage. With Douyin_TikTok_Download_API, you get the same class of data without a quota or a per-request price, and you absorb the breakage yourself.
That trade is worth naming precisely. The vendor approach costs money per call and puts a third party in the request path, which the README treats as the thing it is avoiding. The self-hosted approach costs you a server, a PostgreSQL instance, a Redis instance, and the attention required to notice when the identity pool stops minting. Neither is strictly better; they fail in different places. If your workload is bursty and small, the vendor is cheaper in engineering time. If your workload is large, continuous, or sensitive about where the data flows, the self-hosted path is the one that scales without a meter running.
A narrower alternative is writing your own signer against the platform. The project's own README points at a forum for reverse engineering, reer.dev, and says the project "exists because people wrote down how a signature was built." That is an accurate description of what you would be signing up for: the same work, without the scheduler, the archive, the console and the MCP server that this repository already ships.
Licence, upgrades and what maintenance actually costs
The licence is Apache-2.0, declared in both the repository metadata and pyproject.toml. That is a permissive licence with an explicit patent grant, and it does not impose a copyleft obligation on your own code. It also says nothing about the content you retrieve. The licence covers the software; the videos, images and comments you archive are governed by the platforms' terms and by whatever law applies to you, and the repository does not attempt to resolve that. Treat the licence question and the content question as separate, and get your own answer on the second.
On maintenance, the release history is dense: v5.0.2 and v5.0.3 both landed on 2026-09-11, and v5.1.0 followed on 2026-09-15. The last push to the default branch was on 2026-09-15. Three releases inside a week is the pattern of a project reacting to upstream changes, not one coasting. Plan your upgrade budget accordingly: pin the image tag rather than tracking `latest`, because the .env.example shows DTK_IMAGE_TAG is a supported interpolation, and read the release notes before moving, since a version bump in this project can mean the platform contract moved underneath it.
Schema changes are the other upgrade cost. Alembic is in the dependency list and alembic.ini is at the repository root, so the archive has migrations. A major version bump may require running them against your PostgreSQL before the new application code will start. Back up the database and the secret together before any upgrade, per the guidance in .env.example.
Editorial conclusion
Adopt it if you need Douyin or TikTok posts, profiles, comments and playlists behind your own API, and you can run PostgreSQL, Redis and a container host. Skip it if you need a managed SLA, or if you cannot legally store the media you pull. Before committing, verify the platform side still answers: run the smoke target against the current build, confirm the identity pool mints cookies, and check that the DTK_SECRET_KEY you generated is backed up with the database.
Frequently asked questions
Is Douyin_TikTok_Download_API free to use?
The software is open source under Apache-2.0 and the README states there is no signup and no quota. What you pay for is the infrastructure: a host, PostgreSQL, Redis and the proxy material the identity pool needs.
How do I download videos from Douyin with Douyin_TikTok_Download_API?
Run the stack, open the console, and use the playground to paste a link and send it. The README describes the result as the normalised result returned through the identity pool and scheduler, and it notes the download is watermark-free because the clean stream the platform publishes is selected.
Can I try Douyin_TikTok_Download_API without installing it?
Yes. The README points at demo.douyin.wtf, a live read-only instance where the login page fills in the demo account and the API keys page shows a key in plaintext. Demo requests are not written to the request log or the archive, and the demo is rate limited to 30 requests per 10 seconds.
What does Douyin_TikTok_Download_API need besides Docker?
The stack expects PostgreSQL and Redis, both of which the compose file brings up, plus a secret of at least 32 characters that the process refuses to start without. pyproject.toml requires Python 3.12 or newer and below 3.14 if you run it outside containers.
Why does Douyin_TikTok_Download_API start healthy but never mint identities?
The .env.example warns that values below the marked line are interpolated into compose.yml from docker/.env or the shell, not from the repository root, so a value written there is silently ignored. The file names the browser pin as the specific case where the instance comes up healthy and mints nothing.
Can Douyin_TikTok_Download_API be called from an AI agent?
The project ships an MCP server alongside the REST API, CLI and web console, and the README links to documents/en/12-mcp.md for it. The MCP surface is listed as one of the console's sections.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/evil0ctal-douyin-tiktok-download-api)