Library / SDK
codelucas/newspaper avatar
codelucas/newspaper

newspaper3k: article text and metadata extraction in Python 3

newspaper3k is a news, full-text, and article metadata extraction in Python 3. Advanced docs:

15,167 stars2,115 forksPythonMIT

At a glance

What is it?
newspaper3k downloads a news URL, parses the article body, byline, date and top image, and optionally runs keyword and summary extraction. It is a small library with a broad dependency list, and the README now spends more space on proxies and third-party APIs than on the parser itself.
Who is it for?
Adopt newspaper3k when you have a known list of news URLs and want article text, authors, publish date and top image without writing per-site extractors. Skip it if you need a general-purpose crawler that discovers articles by topic, or if you cannot accept a wide dependency tree that pulls in lxml, nltk, jieba3k, pythainlp and tinysegmenter.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 15 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What newspaper3k extracts, and who is asking for it

The library targets one job: given the URL of a news article, return the parts a reader would see. The README's opening example shows article.authors, article.publish_date, article.text, article.top_image and article.movies after a download and parse call. A separate nlp() call adds article.keywords and article.summary. That is a narrower contract than a general scraper. You are not describing selectors per site; you hand over a URL and the library decides which part of the DOM is the article body.

The audience is developers building news monitoring, archiving or aggregation pipelines. The README points at a second use case explicitly: newspaper.build('http://cnn.com') returns a paper object whose .articles list holds Article objects, and .category_urls() returns section URLs such as http://lifestyle.cnn.com. So the library covers two layers, discovery of links within one site and extraction from one page. What it does not cover is discovery across sites, which the README addresses by recommending an external search API rather than a built-in mechanism.

How Article.download and Article.parse split the work

The object model is deliberately staged. Article(url) constructs the object with no network activity. download() fetches the page and assigns the raw markup to article.html. parse() then runs the extraction over that stored HTML. The split matters because it lets you keep the raw page and re-parse later, and because it makes the failure point obvious: a 403 or a captcha surfaces at download(), while a bad extraction surfaces at parse().

The README's scale section states that the Config object carries proxies and browser_user_agent, and shows config.proxies as a dict with 'http' and 'https' keys pointing at the same gateway URL. That is the documented answer to rate limiting and blocks. The README attributes the problem to your IP rather than to your code, and the example sets a Chrome user agent string alongside the proxy. Note what is absent: the README does not document a retry policy, a delay between requests, or a robots.txt check. If you need those, they are yours to add around the library.

Language handling is a parameter, not a detection step you control. Article(url, language='zh') is the documented form for Chinese, and newspaper.build('http://www.sina.com.cn/', language='zh') applies the same setting to a whole source. The README says that when no language is specified the library attempts to auto detect one. The dependency list explains the cost of that support: jieba3k, pythainlp and tinysegmenter are installed for Chinese and Thai tokenisation whether or not you parse those languages.

Installing newspaper3k and parsing your first article

The package name on PyPI is newspaper3k, not newspaper. The setup.py file contains a version guard that exits with a warning if you run the Python 3 repository under Python 2, and the message tells you to run pip3 install newspaper3k for Python 3 or pip install newspaper for the Python 2 branch, which the README calls deprecated and buggy. Install with pip3:

bash
pip3 install newspaper3k

Some of the nlp() path depends on NLTK corpora. The repository root contains download_corpora.py, which the README does not describe in the excerpt available, so treat the nlp() step as the part most likely to need extra data before it works. The download and parse path does not need it.

The README's glance section gives this sequence for a single URL. Run it, then inspect the printed fields against the live page:

python
from newspaper import Article

url = 'http://fox13now.com/2013/12/30/new-year-new-laws-obamacare-pot-guns-and-drones/'
article = Article(url)
article.download()
article.parse()

print(article.authors)
print(article.publish_date)
print(article.text[:200])
print(article.top_image)

You should see a list of author names, a datetime, a slice of the article body and a URL for the lead image. If article.text is empty or contains navigation links, parse() picked the wrong container for that page.

If you already have HTML and do not want the download step, the README exposes a one-function path:

python
from newspaper import fulltext

text = fulltext(html)

For a whole site, the README uses newspaper.build and then iterates. The paper object exposes .articles and .category_urls():

python
import newspaper

cnn_paper = newspaper.build('http://cnn.com')

for category in cnn_paper.category_urls():
    print(category)

first = cnn_paper.articles[0]
first.download()
first.parse()
first.nlp()
print(first.title, '--', first.summary[:120])

Expect the first call to take noticeably longer than a single-article parse, because build() fetches and inspects the site's feeds and links. The README does not state how many pages build() requests or whether it honours a crawl delay.

The extraction is heuristic, so some pages come back wrong

Nothing in the README promises a fixed accuracy rate, and that is the honest way to read the library. Extraction works by scoring candidate containers in the parsed DOM. On a standard article page with a single body element it tends to land correctly. On pages that interleave live blogs, photo captions, related-story teasers or comment threads inside the same container, the returned article.text can absorb that material. The README gives no configuration for pinning the body element, so when the heuristic misses, the workaround is to bypass parse() and use fulltext() on HTML you have already narrowed, or to post-process the string.

The metadata fields are weaker than the body text. article.authors and article.publish_date depend on markup conventions that vary between publishers, and the README shows one successful example rather than a coverage statement. Treat both as best-effort values and validate them before they reach a database column.

The wrong-tool case is discovery across publications. newspaper.build() needs a site to start from, so it cannot answer "every article about electric vehicles this week". The README's own answer to that question is to query an external news search service and feed the resulting links into Article. That is a real architectural dependency: your topic coverage is only as good as the third-party search index, and you inherit its key management and quota limits. Two of the services named in that part of the README are commercial, and the links carry tracking parameters, which is worth noticing when you read the section as documentation rather than as a tutorial.

newspaper3k against Goose and readability-style extractors

The repository ships GOOSE-LICENSE.txt alongside its own LICENSE, which tells you where the extraction lineage sits: newspaper3k descends from the Goose extractor, and the alternative family is the readability ports such as python-readability, which wrap the Arc90 readability algorithm. The practical difference is breadth of output. A readability port typically returns cleaned HTML or a text string for a page you have already fetched. newspaper3k returns a structured object with authors, publish date, top image, movies, keywords and a summary, and it also provides the site-level build() step. If all you need is the body text of a page you fetched yourself, the readability approach has a smaller dependency footprint; newspaper3k's requirements.txt pulls in lxml, nltk, Pillow, feedparser, feedfinder2, tldextract and the tokenisers.

The other real alternative is writing per-site extractors on top of lxml or BeautifulSoup. That is more work per publisher and it breaks whenever a layout changes, but it is deterministic and you can test it. newspaper3k trades that determinism for not having to write the extractor at all. For a handful of high-value sources, the hand-written extractor is often the better trade; for hundreds of sources you do not control, it is not.

Maintenance, licence and the cost of upgrading

The repository is not archived, and the last push was on 2026-09-15. The most recent release listed is 0.0.9 from 2014-12-17, labelled "End of Python 2 support", while setup.py declares version 0.3.0. That gap means the release history does not describe what you get from pip; the installable version is the one in setup.py, and the changelog file in the repository root is the place to look for what moved between them. If your deployment pins versions, pin against what pip actually resolves and record it, because a PyPI release page that stops at 0.0.9 gives you no upgrade narrative.

The licence is MIT, and the repository also carries GOOSE-LICENSE.txt for the inherited extractor code. MIT is permissive, but the presence of a second licence file means the codebase is not uniformly under one grant, and the README does not explain which files fall under which. That is a question for whoever reviews dependencies in your organisation; it is not something this article can settle. The upgrade cost is dominated by the dependency list rather than by the library's own API. lxml, Pillow and nltk are the ones that tend to force coordinated upgrades, and jieba3k and tinysegmenter are pinned or effectively frozen (requirements.txt carries tinysegmenter==0.3 with a TODO about relaxing it), so a Python or platform change that breaks them is not something you can fix by bumping a version.

Editorial conclusion

Adopt newspaper3k when you have a known list of news URLs and want article text, authors, publish date and top image without writing per-site extractors. Skip it if you need a general-purpose crawler that discovers articles by topic, or if you cannot accept a wide dependency tree that pulls in lxml, nltk, jieba3k, pythainlp and tinysegmenter. Before committing, run Article(url).download() and .parse() on five real pages from your target sites and check article.text and article.authors against the rendered page, because the README documents no accuracy guarantee and no rollback path.

Frequently asked questions

How do I install newspaper3k in Python?

Run pip3 install newspaper3k. The package name on PyPI is newspaper3k; the separate newspaper package is the deprecated Python 2 branch that setup.py warns against installing on Python 3.

How do I install newspaper3k?

The setup.py guard message points to pip3 install newspaper3k for Python 3. The README's glance section then imports from newspaper, so the import name and the install name differ.

What fields does newspaper3k return for an article?

After download() and parse() the README shows authors, publish_date, text, top_image and movies. Calling nlp() adds keywords and summary.

Can newspaper3k find articles about a topic across many sites?

No. newspaper.build() starts from a site you name and returns that site's articles and category URLs. The README's answer for cross-publication discovery is to query an external news search API first and pass the resulting links to Article.

Does newspaper3k support languages other than English?

Yes. The README shows Article(url, language='zh') for a single Chinese article and newspaper.build('http://www.sina.com.cn/', language='zh') for a whole source, and states that without a language setting the library attempts to auto detect one.

How do I stop newspaper3k from getting blocked while scraping?

The README sets config.proxies to a dict with 'http' and 'https' keys and also sets config.browser_user_agent to a browser string. It documents no retry policy or crawl delay, so pacing is left to your code.

Official sources

  1. codelucas/newspaper on GitHub
  2. License: MIT
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/codelucas-newspaper.svg)](https://hysenlabs.com/projects/codelucas-newspaper)