CLI tool
vladkens/twscrape avatar
vladkens/twscrape

twscrape: the account pool waits forever unless you tell it not to

Python library and CLI for X/Twitter scraping with multi-account rotation and built-in rate-limit handling.

2,829 stars322 forksPythonMIT

At a glance

What is it?
An async Python library and CLI for X/Twitter Search and GraphQL endpoints, built around a local SQLite pool of your own account sessions with automatic rotation on rate limits. The defaults are where the judgement calls live: indefinite waiting when every account is locked, TLS fingerprinting as an install extra, and an account identifier the tool never checks against the username in the cookies.
Who is it for?
twscrape is a well shaped client for people who already own the accounts it uses, and its design is honest about being an account pool rather than an API client: sessions live in a local database, rotation happens automatically when an endpoint rate limits, and every parsed method has a raw twin for when the models drift from the response. Two defaults deserve an explicit decision before it runs in anything unattended.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 13 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 4, 2026, and from our analysis. They are not legal advice.

Editorial analysis

With no configuration the pool waits while every account is locked

The behaviour when no account is free is configurable, and the default is the one that never gives up. The constructor takes three switches:

python
api = API(raise_when_no_account=True, wait_timeout=30, wait_interval=1)

`wait_timeout` bounds how long to wait for a locked account, `wait_interval` sets how often the pool looks again, and `raise_when_no_account` raises `NoAccountError` instead of ending the operation. Left alone, the documentation says twscrape waits indefinitely while active accounts are locked, so an unattended script whose accounts are all held does not fail fast, it simply stops making progress. Locks have a matching requirement on the way out. Breaking out of an async generator early leaves the account locked until the generator is closed, which the documentation handles by wrapping it in `contextlib.aclosing`. Both halves describe the same resource model: an account is checked out for the duration of an iteration, so an abandoned loop holds a slot.

TLS fingerprinting is an install extra chosen by an environment variable

The default HTTP backend is `httpx`, and the more browser shaped option is opt in. Installing the extra and pointing the tool at it is two steps:

bash
pip install "twscrape[curl]"
TWS_HTTP_BACKEND=curl twscrape user_by_login xdevelopers

The manifest matches: `curl = ["curl-cffi>=0.7.0"]` under optional dependencies, with the same package repeated in the development group. The stated reason is browser-like TLS fingerprinting, which is a different layer from the headers, and that layer is handled unconditionally because `fake-useragent>=1.4.0` is a required runtime dependency rather than an extra. So a plain `pip install twscrape` gets randomised user agents and httpx, and only the TLS handshake changes when the extra is added. The choice is per invocation through an environment variable rather than a config file, which means two processes on one machine can use different backends without any shared state.

The cookies are the account, and the local name is never checked

Accounts are created from browser cookies, and the two values that matter are named in the documentation: `auth_token` and `ct0`. The recommended path exports them from a browser profile and pipes them straight in:

bash
unjar x.com -f header | twscrape add_cookie my_account
twscrape accounts
twscrape search "from:xdevelopers lang:en" --limit=20

The alternative is `twscrape add_cookie my_account` with no pipe, which prompts for cookies copied out of the browser's developer tools. Accounts holding both cookies are activated immediately, so there is no separate login step to run. Two details of the model are worth stating plainly. `my_account` is only a local identifier, and the tool does not verify that it matches the X username inside the cookies, so a mismatched name will not be caught here. Running the same command again replaces the saved session while preserving credentials, statistics, locks and proxy settings, which means re-adding a cookie pair is how you refresh a session rather than creating a second account.

Rotation is automatic, and there is no per-target boundary on the pool

The feature list describes automatic account switching across rate limited operations, and saved sessions with per account proxies. Nothing in that model asks whether switching is appropriate for the operation being performed: the pool rotates when an endpoint answers with a limit, and the caller is not consulted. That is a scope property of the design rather than a bug, and it is the part to think about before using the library in a shared or unattended setting, because the number of accounts in the database is the number of identities the tool will present. The project's own text is direct about the boundary: it requires authorized X/Twitter accounts, states that X's Terms of Service discourage using multiple accounts, and tells users to act responsibly at their own discretion. Two adjacent notes complete the picture. Ready-made cookie accounts and proxies are offered through links the README labels as referral links, and the sponsor section names a residential proxy vendor with a claimed address count and a discount code.

GraphQL operation ids are refreshed by a script, and the SQLite matrix has holes in it

The build tooling shows how the endpoints are kept current. A single target does the refreshing:

code
uv run scripts/update-gql-ops.py
uv run scripts/update-mocked-data.py refresh

So the GraphQL operation identifiers live in the repository as data and are replaced by a script rather than by hand, and the mocked responses that the tests replay are refreshed alongside them. The test suite depends on `pytest-httpx` for that replay and `pytest-asyncio`, with `asyncio_mode = "auto"` and a session scoped fixture loop, and deprecation warnings are filtered out rather than surfaced. The matrix targets are thinner than they look. The Python matrix is complete, running 3.10 through 3.14 against a build argument, but the SQLite matrix keeps only its endpoints: a 2018 build and a 2026 build are active while the 2019 through 2024 entries sit commented out under a link to the SQLite release chronology.

Dependency policy excludes the last seven days, and two checkers do the work

The manifest carries a supply chain setting that is easy to miss: `exclude-newer = "7 days"` under the uv section, so resolution refuses packages published in the last week. Alongside it sit two static analysers with different jobs. Ruff is configured at line length 100 targeting Python 3.10, selecting the pycodestyle, pyflakes, isort, pyupgrade, comprehensions and simplify rules while ignoring long lines, an old typing import and the contextlib suggestion. `ty` is the type checker, configured for Python 3.10 across all platforms. The Makefile splits them the way you would expect: `lint` rewrites files, running ruff with the import rule selected and `--fix` plus the formatter, while `check` only verifies, running the formatter check, the full ruff check and `ty check`. `prepare` is the combination of the two, so the pre-commit path is a rewrite followed by a verification.

Every parsed method has a raw twin, and the model list is cut off

The surface is wide and grouped by object: search, search_user and search_trend; tweet_details, tweet_replies, tweet_thread, retweeters and bookmarks; user lookups by id and login, user_about, following, followers, verified_followers, subscriptions, user_tweets and user_tweets_and_replies; list timelines and members; community info, members, moderators and tweets; and trends by category. Search defaults to the Latest tab and takes a `kv` mapping to switch product, `Top` or `Media`. The escape hatch is systematic: each parsed method has a `_raw` version returning the original wrapper, so a script can read `rep.status_code` and `rep.json()` when the parsed model stops matching the response. The last line of the visible documentation introduces the parsed models, naming `Tweet`, `User`, `Community` and the trend objects, and ends in the middle of that sentence. Package metadata around them: version 0.20.1, Beta status, Python 3.10 to 3.14, a committed lock file, and lowercase `readme.md` and `changelog.md`.

Editorial conclusion

twscrape is a well shaped client for people who already own the accounts it uses, and its design is honest about being an account pool rather than an API client: sessions live in a local database, rotation happens automatically when an endpoint rate limits, and every parsed method has a raw twin for when the models drift from the response. Two defaults deserve an explicit decision before it runs in anything unattended. With no configuration the pool waits indefinitely while accounts are locked, so a stalled scrape looks like a slow one, and a generator you break out of must be closed with aclosing or the lock is held on. And the browser-like TLS fingerprinting is an extra you install and switch on with an environment variable, which means the more evasive posture is opt in rather than the default. Where it does not fit is a context where using several accounts on one service is prohibited, since the project's own text says X's Terms of Service discourage it.

Frequently asked questions

How do I use twscrape?

Add an account from browser cookies that contain auth_token and ct0, for example with unjar x.com -f header piped into twscrape add_cookie my_account, confirm it with twscrape accounts, then query with twscrape search "from:xdevelopers lang:en" --limit=20. Accounts holding both cookies are activated immediately, with no separate login step.

Is scraping Twitter with twscrape legal?

The project states that it requires authorized X/Twitter accounts, that X's Terms of Service discourage using multiple accounts, and that you should use it responsibly at your own discretion. It offers no legal analysis, so that question is not settled by the project itself.

Do I need a Twitter API key to use twscrape?

No API key path is documented. Accounts are added from browser cookies containing auth_token and ct0, sessions are kept in a local SQLite database, and requests go to X's Search and GraphQL endpoints using the account you supplied.

Which HTTP backend does twscrape use?

httpx is the default. Installing the curl extra with pip install "twscrape[curl]" adds curl-cffi for browser-like TLS fingerprinting, and the backend is selected per invocation with TWS_HTTP_BACKEND=curl. The fake-useragent package is a required dependency either way.

Official sources

  1. License: MIT
  2. Project website
  3. README
  4. Releases
  5. vladkens/twscrape on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/vladkens-twscrape.svg)](https://hysenlabs.com/projects/vladkens-twscrape)