CLI tool
brightdata/cli avatar
brightdata/cli

Bright Data CLI: Scrape, Search and Extract Web Data from the Terminal

Official Bright Data CLI - scrape, search, and extract structured web data directly from your terminal.

6,288 stars84 forksTypeScriptMIT

At a glance

What is it?
The official @brightdata/cli wraps Bright Data's Unlocker, SERP, Web Scraper and Browser APIs behind one `brightdata` command. It is a thin client for a paid service, and the free tier is what decides whether it fits.
Who is it for?
Adopt it if you already pay for Bright Data and want scrape, search, pipelines and browser control in one scriptable command, or if you want to try the 5,000 monthly free credits without writing HTTP code. Skip it if you need a self-hosted scraper with no vendor account, or if you expect `brightdata browser` to run on the free tier: the README states Browser API is excluded from monthly credits and gets only a one-time $2 trial.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 14 days ago.
What is it written in?
Mainly TypeScript, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What Bright Data CLI actually wraps

This is not a scraper. It is a command line client for Bright Data's hosted products, and every command maps to a billed API call on their side. The README's command table is the clearest statement of scope: `brightdata scrape` goes through the Web Unlocker, `brightdata search` through the SERP API, `brightdata pipelines` through the Web Scraper API for 40+ named platforms, and `brightdata browser` through the Scraping Browser. The CLI adds argument parsing, output formatting and a config file on top of those endpoints.

The audience follows from that. If you already have a Bright Data account and you currently write throwaway Python or curl scripts to hit their endpoints, this replaces that glue. It is also aimed at people who want a scraper inside a shell pipeline, or who want to hand an agent a single command rather than an SDK. The `brightdata skill` and `brightdata add mcp` commands exist specifically for that second case.

What it is not: a way to avoid paying Bright Data. There is no local scraping engine in the repository, and the dependency list (`commander`, `@clack/prompts`, `csv-parse`, `playwright-core`) is CLI plumbing, not a crawler.

How a scrape request flows through the CLI

The mechanism is a conventional TypeScript CLI built with `commander`, compiled by `tsc` into `dist/index.js`. The `package.json` declares two binaries, `brightdata` and `bdata`, both pointing at the same entry file, so the shorthand is not a separate program.

A scrape call takes a URL and options, then sends the request to Bright Data with the credentials the CLI has stored. Format conversion happens server side: `-f markdown`, `html`, `screenshot` or `json` selects what comes back, with `markdown` as the default. Geo-targeting is a parameter (`--country us`), not a local proxy configuration, which is the point of the wrapper: the proxy layer is invisible to you.

Authentication state is the other half of the flow. `brightdata login` opens a browser and saves a key; `brightdata login --api-key <key>` skips the browser; or you export `BRIGHTDATA_API_KEY` and skip login entirely. On first login the CLI checks for the zones it needs, `cli_unlocker` and `cli_browser`, and creates them if they are missing. That auto-creation is convenient and slightly surprising: a login command that provisions cloud resources is doing more than the name suggests.

One asymmetry is worth flagging. The README states that `brightdata add mcp` uses the key stored by `brightdata login` and does not currently read `BRIGHTDATA_API_KEY` or the global `--api-key` flag. If you run everything else through an environment variable, that one command still requires an interactive login first.

Installing @brightdata/cli and running a first scrape

Node.js 20 or later is required. On macOS and Linux the README gives a shell installer:

bash
curl -fsSL https://cli.brightdata.com/install.sh | sh

On Windows, or on any platform if you prefer npm, install the package globally. The same command works for a manual install everywhere:

bash
npm install -g @brightdata/cli

If you do not want a global install, the README documents an npx form that runs the CLI without installing it:

bash
npx --yes --package @brightdata/cli brightdata <command>

After installing, the recommended first step is the interactive wizard, which detects your API key, lets you pick a zone, sets a default output format and prints quick-start examples:

bash
brightdata init

With credentials in place, the smallest useful call is a scrape to markdown. Running it against a plain page should print the page body as markdown in your terminal:

bash
brightdata scrape https://example.com

Search is the same shape, and returns structured JSON rather than a rendered page:

bash
brightdata search "web scraping best practices"

For known platforms, `pipelines` skips the parsing work. The README's own example extracts a LinkedIn profile by dataset name and URL:

bash
brightdata pipelines linkedin_person_profile "https://linkedin.com/in/username"

Finally, check what you have spent. `brightdata budget` shows account balance plus per-zone cost and bandwidth, which is the command you will run most often once anything is automated:

bash
brightdata budget

The free tier is the real adoption constraint

Every new account gets 5,000 credits per month, renewing on the 1st, with unused credits not rolling over. The README equates that to roughly $7.50. Four products draw on it: Unlocker (`scrape`), SERP (`search`), Web Scraper (`pipelines`) and Scraper Studio (`scraper create`, `run`, `heal`), each at 1 credit per request, except Scraper Studio which is 1 credit per page load from a shared pool rather than per record.

Two exclusions matter more than the headline number. Proxy products (Datacenter, ISP, Residential) and the Browser API behind `brightdata browser` are not covered by the monthly credits. They get a separate one-time $2 trial valid for 7 days, plus a $5 bonus valid for 30 days if you add a payment method. So the browser automation command, which is arguably the most interesting part of the tool, is the part the free tier does not sustain.

The README also notes that accounts on custom pay-as-you-go or pre-commit subscription plans do not receive the recurring monthly free credits. If your organization already has a Bright Data contract, the free tier described here may not apply to you at all, and the CLI becomes purely a convenience wrapper over endpoints you are already paying for. Adding a card is described as a verification step; the README states you are not charged unless credits are exhausted and funds are deposited.

Where the CLI model breaks down

The mismatch is between a stateless command and a stateful workflow. `brightdata scraper heal` fixes an existing scraper through AI self-healing but stops at an approval gate, and `brightdata scraper approve` exists to release it. That is a two-command dance across two terminal invocations, which is fine interactively and awkward in a cron job or CI step. The design is deliberate, and the approval gate is the safer default, but it means the scraper lifecycle is not fully automatable from the CLI alone.

Long-running work is the second gap. `brightdata browser` navigates, snapshots, clicks and types, but each invocation starts a session. Anything that needs to hold a browser open across many steps has to be scripted around the CLI rather than inside it, and the README does not document a persistent session handle for that.

Credential handling is the third. The README explicitly notes that `brightdata add mcp` ignores `BRIGHTDATA_API_KEY` and the `--api-key` flag. If you standardize on environment variables for CI, that one command breaks the pattern, and the workaround is an interactive login on the machine that runs it.

Finally, this is the wrong tool entirely if you cannot send the target URLs to a third party. Every command routes through Bright Data's infrastructure. There is no offline or self-hosted mode, and the repository contains no crawler to fall back on.

Bright Data CLI against a self-hosted scraper

The natural alternative is a library you run yourself, such as Playwright or Puppeteer with a proxy pool you manage. The difference is where the hard parts live. With the CLI, CAPTCHA handling, JavaScript rendering, anti-bot evasion, geo-targeting and format conversion are all server-side concerns you configure with flags like `--country` and `-f markdown`. With a self-hosted stack, you own all of that: browser binaries, retry logic, proxy rotation, and the ongoing cost of keeping selectors working when sites change.

That trade is not automatically in the CLI's favor. A self-hosted scraper has no per-request billing, no third-party data flow, and no account that can be rate-limited or suspended. For a small number of stable internal pages, running Playwright locally is cheaper and simpler than provisioning zones. The CLI wins when the targets are hostile or numerous, when you need geo-distributed requests without operating proxies, or when you want named dataset extraction for a platform like Amazon or LinkedIn instead of writing and maintaining a parser.

The `pipelines` command is the sharpest illustration. It is not a generic scraper; it is a catalog of pre-built extractors for 40+ platforms. If your target is in that catalog, you skip parser maintenance entirely. If it is not, you are back to `scrape` and doing the parsing yourself.

Maintenance, releases and the MIT licence

The repository is not archived and the last push was on 2026-09-16, which is days before this writing. Recent releases show a steady cadence with small, security-flavored fixes rather than feature churn: v0.3.5 fixed a login auth exit code, v0.3.6 addressed terminal output sanitization, and v0.3.7 added CSV formula injection protection. Those last two are worth reading as a signal about the threat model. A CLI that prints remote page content into your terminal and writes CSV you may open in a spreadsheet has real injection surface, and the maintainers are patching it. The version number is still 0.3.x, so expect breaking changes between minor releases.

The package is MIT licensed, and the repository carries a LICENSE file at the top level. That covers the CLI code. It does not cover the service: your use of the Unlocker, SERP, Browser and Web Scraper APIs is governed by Bright Data's terms and by whatever your targets' terms of service say. The MIT grant tells you nothing about whether a given scrape is permitted, and the README does not address that question. If your use case has legal review attached, the licence is the easy part.

Upgrade cost is low in the ordinary case. It is a global npm package, so `npm install -g @brightdata/cli` moves you to the latest version, and `brightdata init` re-runs the setup wizard if configuration needs to change. The risk is behavioral rather than mechanical: pre-1.0 output formats and flag behavior can shift, so anything parsing CLI output in a script should be pinned and tested before upgrading.

Editorial conclusion

Adopt it if you already pay for Bright Data and want scrape, search, pipelines and browser control in one scriptable command, or if you want to try the 5,000 monthly free credits without writing HTTP code. Skip it if you need a self-hosted scraper with no vendor account, or if you expect `brightdata browser` to run on the free tier: the README states Browser API is excluded from monthly credits and gets only a one-time $2 trial. Before committing, run `brightdata budget` after a week of real use to see which zone each command actually bills against, and check whether your account is on a plan that receives the recurring free credits at all.

Frequently asked questions

What is Bright Data CLI and what does it install?

@brightdata/cli is the official npm package for the Bright Data CLI. It installs the `brightdata` command with `bdata` as a shorthand alias, and requires Node.js 20 or later.

How do I install Bright Data CLI?

On macOS and Linux the README gives a shell installer at cli.brightdata.com/install.sh. On Windows, or manually on any platform, the documented command is `npm install -g @brightdata/cli`, and there is an npx form for running it without installing.

How do I log in to Bright Data CLI?

Run `brightdata login` for an interactive browser flow, `brightdata login --github` if the GitHub CLI is available, or `brightdata login --api-key <your-api-key>` non-interactively. You can also export `BRIGHTDATA_API_KEY` and skip login, though `brightdata add mcp` does not read that variable.

How much is Brightdata?

The README states new accounts get 5,000 free credits per month (about $7.50), renewing on the 1st with no rollover. Proxy products and the Browser API are excluded from those credits and instead get a one-time $2 trial plus a $5 bonus when a payment method is added.

Is Brightdata good?

The README documents command coverage, free tier limits and troubleshooting rather than any evaluation of result quality, so a quality verdict cannot be given here. What is clear is that the CLI is a thin client for Bright Data's paid APIs rather than a standalone scraper.

Official sources

  1. brightdata/cli on GitHub
  2. Issues
  3. License: MIT
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/brightdata-cli.svg)](https://hysenlabs.com/projects/brightdata-cli)