# joeyism/linkedin_scraper: an async Playwright scraper for LinkedIn profiles, companies and jobs

> The library moved from Selenium to Playwright in v3.0.0 and made every method async, which breaks v2 code. It is a good fit if you already have authenticated sessions and need typed Pydantic models, and a poor fit if you wanted a headless service with no browser.

**joeyism/linkedin_scraper** — A library that scrapes Linkedin for user data

- Repository: https://github.com/joeyism/linkedin_scraper
- Stars: 4,558 · Forks: 992
- Language: Python
- License: GPL-3.0
- Published: 2026-09-23 · Updated: 2026-09-23 · Language: en
- Canonical page: https://hysenlabs.com/projects/joeyism-linkedin-scraper

## What joeyism/linkedin_scraper does that a raw HTTP client cannot

LinkedIn does not serve profile, company or job pages as static HTML to an unauthenticated client. The pages that matter sit behind a login and are rendered in the browser. A requests-plus-lxml script gets the login wall, not the profile. This library takes the other route: it drives a real Chromium instance through Playwright, reuses a saved authenticated session, and parses the rendered DOM into typed objects.

The intended audience is a Python developer who needs structured LinkedIn data inside a script or a pipeline and is willing to run a browser to get it. The README lists four scraping targets: person profiles (basic info, work experience, education, skills), company pages (overview, industry, size, headquarters), company posts (text, reaction, comment and repost counts, posted date, images) and job listings (details, requirements, company, application links). Each target has its own scraper class, so you import only what you need.

Two design choices stand out. The data models are Pydantic, which means you get validation and attribute access instead of dictionary key lookups. And the whole surface is async, which is a deliberate break from the older synchronous Selenium API.

## The v3.0.0 rewrite: Playwright, async/await and a new import path

Version 3.0.0 is explicitly not backwards compatible, and the README states this in its own breaking-changes section. Four things changed at once: Selenium was replaced by Playwright, every method became async and requires await, the package structure and import names changed, and the data models moved to Pydantic.

The migration example in the README is the clearest statement of the difference. In v2 you constructed a Person with a URL and a driver and read person.name synchronously. In v3 you open a BrowserManager as an async context manager, load a session file, hand browser.page to a PersonScraper, and await the scrape call.

The upgrade path is a full rewrite of the call sites, not a version bump. If you cannot do that, the README points at the escape hatch: pip install linkedin-scraper==2.11.2 keeps the Selenium-based version. That is a pinned dead end rather than a supported branch, and it will not receive the Playwright fixes.

The async model has a practical consequence. A single BrowserManager wraps one page, and the scrapers take that page. If you want concurrency you have to reason about how many browser contexts you open, because the library does not hide that from you.

## Installing linkedin-scraper and running a first authenticated scrape

Installation is two steps, because the Python package and the browser binary are separate. The README gives both commands. The second one matters: without it, Playwright has no Chromium to launch and the first scrape fails at browser startup rather than at parsing.

```bash
pip install linkedin-scraper
playwright install chromium
```

LinkedIn requires authentication, so before any scraping you need a session file. The README offers a manual login script and a programmatic one. The manual path opens a non-headless browser, waits for you to log in, and writes the session to disk. The timeout is in seconds and the README uses 300.

```python
from linkedin_scraper import BrowserManager, wait_for_manual_login
import asyncio

async def create_session():
    async with BrowserManager(headless=False) as browser:
        await browser.page.goto("https://www.linkedin.com/login")
        await wait_for_manual_login(browser.page, timeout=300)
        await browser.save_session("session.json")

asyncio.run(create_session())
```

With session.json in place, scraping a profile is a few lines. Note headless=False in the README example: the browser is visible while it works.

```python
import asyncio
from linkedin_scraper import BrowserManager, PersonScraper

async def main():
    async with BrowserManager(headless=False) as browser:
        await browser.load_session("session.json")
        scraper = PersonScraper(browser.page)
        person = await scraper.scrape("https://linkedin.com/in/williamhgates/")
        print(f"Name: {person.name}")
        print(f"Experiences: {len(person.experiences)}")

asyncio.run(main())
```

You should see the profile name and an experience count printed. If the session file is stale, the page that loads is the login page and the parsed fields come back empty rather than raising a clear authentication error, so check the browser window on the first run.

The repository also ships runnable samples. Cloning and installing in editable mode, then running the sample scripts, is the fastest way to confirm the whole chain works before you wire it into your own code.

```bash
git clone https://github.com/joeyism/linkedin_scraper.git
cd linkedin_scraper
pip3 install -e .
python3 samples/create_session.py
python3 samples/scrape_person.py
```

For programmatic login the credentials come from the environment. The .env.example file defines LINKEDIN_EMAIL and LINKEDIN_PASSWORD, and notes that LINKEDIN_USERNAME works as an alternative to LINKEDIN_EMAIL.

## Job search and company posts: the scrapers with search parameters

Two scrapers take search input rather than a single URL. JobSearchScraper.search accepts keywords, location and a limit, and returns job objects with title, company, location and linkedin_url. CompanyPostsScraper.scrape takes a company URL and a limit, and returns posts with posted_date, text, reactions_count, comments_count and linkedin_url.

The limit parameter is the part to think about. It caps how many results you ask for, but the library still drives a browser through a logged-in LinkedIn session to collect them, so a limit of 10 and a limit of 500 are very different in wall-clock time and in how much scrolling the page performs. The README does not document rate limiting, retry behaviour or backoff between requests. That silence is the biggest gap in the documentation for anyone planning a bulk run.

Both scrapers follow the same shape as PersonScraper: construct with browser.page, await the method. That consistency is the nice part of the v3 API. The less nice part is that all of them share the one page from BrowserManager, so interleaving a job search and a profile scrape in the same context means they queue on the same browser tab.

## Where joeyism/linkedin_scraper is the wrong tool

The most obvious limitation is architectural: this is a browser automation library, not an HTTP scraper. Every run needs Chromium installed and a display context unless you are confident headless mode works against the current LinkedIn front end. On a small container with no browser, nothing here runs. If your deployment target is a slim serverless function, this library does not fit it.

Authentication is the second constraint, and it is not a small one. You must obtain a session, and the README's own examples use headless=False. The manual login flow assumes a human in front of a browser. The programmatic flow assumes you are comfortable putting account credentials in environment variables and logging in through automation, which is exactly the pattern that tends to end in a challenged account. Nothing in the README describes what happens when LinkedIn serves a captcha or an interstitial; the library has no documented handling for it.

The third limitation is the version split. Because v3 broke compatibility, a large amount of code and tutorial material in the wild is written against the v2 Selenium API. Copying a v2 snippet into a v3 install fails on the import line. The README's migration section is the only bridge, and it covers the Person example, not every scraper.

Finally, the package metadata describes the project as Development Status :: 3 - Alpha. Treat the API surface as capable of changing again between minor versions.

## How it compares to a managed scraping API

The alternative most people weigh against this library is a hosted scraping service, the kind that appears in searches as a LinkedIn scraper API. The difference in approach is where the browser and the session live. With joeyism/linkedin_scraper, you own both: Chromium runs on your machine, session.json is your file, and the credentials in .env are your account. With a managed API, the vendor runs the browsers and manages the accounts, and you send an HTTP request and get JSON back.

That trade is straightforward. The library costs you an install step, a browser binary, a session file and the operational work of keeping a logged-in session alive. In exchange you get no per-request billing, no vendor holding your data pipeline, and full control over parsing because the source is in front of you. The managed service costs money per call and puts a third party between you and the data, but it removes the browser and the login from your infrastructure entirely.

A second alternative is simply not scraping at all and using LinkedIn's own export or an official API where one applies to your use case. That is the only route with no account risk, and it is worth ruling out before you install a browser automation stack.

## Licence, maintenance and the upgrade cost of the v3 API

The repository's LICENSE file and the project metadata in pyproject.toml and setup.py both declare Apache 2.0. The GitHub repository listing shows GPL-3.0. Those two statements do not agree, and the discrepancy is worth resolving with the maintainer or your own legal review before you ship the library inside a product. I am not giving legal advice here; I am pointing at a conflict a reader can see for themselves.

The maintenance picture is concrete. The last push was on 2026-04-10, and the most recent release, v3.1.2, was tagged the same day. Before that, 3.1.1 landed on 2026-01-27 and 3.1.0 on 2026-01-18. So the project saw three releases across roughly three months and then no pushes for the following five months. That is a burst pattern, not a steady cadence, and it is the thing to weigh if you depend on fixes landing quickly.

Upgrade cost is dominated by the async rewrite. If you are starting fresh on v3, the cost is low: the four scrapers share one construction pattern, and the Pydantic models give you stable attribute names. If you are migrating from v2, budget for rewriting every call site, not for a dependency bump. The pinned 2.11.2 install exists, but pinning means you inherit none of the Playwright work.

## Conclusion

Adopt it if you are a Python developer who can run a real browser, already has a logged-in LinkedIn session, and wants typed profile, company, post and job objects rather than raw HTML. Do not adopt it if you need a headless server-side service with no Chromium, if you are still on the v2 Selenium API and cannot rewrite to async, or if you want a managed scraping API that handles login for you. Before writing production code, verify three things: that the package you install reports version 3.x and not 2.11.2, that playwright install chromium completes on your host, and that your session file still authenticates after the first scrape run.

## FAQ

### How do I install joeyism/linkedin_scraper?

Install the package with pip install linkedin-scraper, then run playwright install chromium so the browser binary is present. Without the second command the first scrape fails at browser startup.

### What is joeyism/linkedin_scraper?

It is an async LinkedIn scraper built with Playwright for extracting profile, company, post and job data. It returns Pydantic models rather than raw HTML, and every method must be awaited.

### Is there a free LinkedIn scraper like this one?

This library is open source and installs from PyPI, so there is no per-request fee. The cost it does carry is operational: you run Chromium locally and maintain your own authenticated session file.

### Can I scrape LinkedIn profiles with joeyism/linkedin_scraper?

Yes, PersonScraper.scrape takes a profile URL and returns name, headline, location, experiences and educations. It requires a valid session loaded through BrowserManager.load_session first.

## Sources

- [Issues](https://github.com/joeyism/linkedin_scraper/issues)
- [joeyism/linkedin_scraper on GitHub](https://github.com/joeyism/linkedin_scraper)
- [License: GPL-3.0](https://github.com/joeyism/linkedin_scraper/blob/master/LICENSE)
- [README](https://github.com/joeyism/linkedin_scraper/blob/master/README.md)
- [Releases](https://github.com/joeyism/linkedin_scraper/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/joeyism-linkedin-scraper
