# news-please promises Python 3.8 and ships a Dockerfile that builds on 3.6.5 from master

> news-please is a news crawler and article extractor that combines Scrapy, newspaper4k and readability, usable three ways: as a library, as a CLI that crawls continuously from a site list, and as a Common Crawl workflow. The packaging metadata, the dependency pins and the container recipe each describe a different version of the same project.

**fhamborg/news-please** — news-please - an integrated web crawler and information extractor for news that just works

- Repository: https://github.com/fhamborg/news-please
- Stars: 2,494 · Forks: 459
- Language: Python
- License: Apache-2.0
- Published: 2026-09-18 · Updated: 2026-09-18 · Language: en
- Canonical page: https://hysenlabs.com/projects/fhamborg-news-please

## The container builds on Python 3.6.5 and installs from an unpinned clone

The Dockerfile starts `FROM python:3.6.5-alpine3.7`, which is two minor versions below the 3.8 floor the README states and four below the newest classifier in the packaging metadata. What it builds next is a full toolchain in the final image, with apk pulling in curl, git, make, gcc, python-dev, musl-dev, libxml2-dev, libxslt-dev, openssl-dev, zlib-dev and jpeg-dev. Then it clones the project and installs the requirements file:

```
RUN git clone https://github.com/fhamborg/news-please.git /news-please
RUN cd /news-please && pip3 install -r requirements.txt
```

No commit, tag or version is pinned on that clone, so an image built twice from the same file can contain different code. The entry point is `docker.sh`, copied to the root and made executable, rather than the package's own console script.

## One dependency is pinned exactly, with a reason in a comment

Most entries in the requirements file use open lower bounds: `Scrapy>=1.1.0`, `elasticsearch>=2.4`, `lxml>=3.3.5`, `beautifulsoup4>=4.3.2`, `warcio>=1.3.3`, `newspaper4k>=0.9.3.1`. One line is different. `readability-lxml` is held at `0.8.1`, exactly, because the comment above it says version 0.8.4.1 is broken and links the upstream issue for python-readability. That is a dependency that will not move on its own, which is the correct treatment for a known bad release but also means this install silently keeps a library from years ago. The file also lists `bs4` as a bare entry right after `beautifulsoup4`, so the same package arrives twice by two names, and it carries platform markers for pywin32 on Windows plus storage clients for MySQL, Postgres, Redis, S3 through boto3, and ElasticSearch.

## Running the CLI with no arguments starts crawling the example sites

The shortest documented run is one word:

```bash
$ news-please
```

With no arguments and no configuration, the crawler does not idle and it does not read standard input. It starts fetching the pages listed in the example configuration, writes extracted results as JSON files into a `data` folder, and in the default setup also saves the original HTML of every page it fetched. You stop it with CTRL+C, at which point the README says shutdown takes between five and sixty seconds, and pressing the key twice kills the process immediately, described as not recommended. For an installed package on a machine with no site list of your own, that is a live network fetch against somebody else's example configuration, which is worth knowing before the first run.

## The Common Crawl workflow needs a clone, not an install

Three capabilities are described: crawling a list of article URLs from library mode, running the crawler from a site list, and pulling articles out of the Common Crawl news archive. The third is the one that does not work from pip. Its instructions are to clone the repository, edit the config section inside `newsplease/examples/commoncrawl.py`, and run the module directly with `python3 -m newsplease.examples.commoncrawl`. Filters can be defined for publisher and for date period. So a user who installed the package with pip and wants archive extraction has to obtain the source tree as a separate step. The library mode equivalents, by contrast, are ordinary calls: `from_url`, `from_urls` with optional request arguments such as a timeout, `from_file` for a list of URLs one per line, `from_html` where an optional original URL improves date extraction, and `from_warc` for a WARC record.

## Results land in JSON, or in a store you enable through a pipeline block

Where output goes is a configuration choice in two files. The site list is `sitelist.hjson`, a JSON with comments format, and the instructions for crawling your own pages point you at the project wiki entry for that file. Storage is chosen through the Scrapy item pipeline configuration in `config.cfg`, which by default lives in `~/news-please/config` and can be pointed elsewhere with the `-c` parameter; if the directory does not exist, a default one is created there. Enabling ElasticSearch means adding its pipeline alongside the extractor:

```cfg
[Scrapy]
ITEM_PIPELINES = {
  'newsplease.pipeline.pipelines.ArticleMasterExtractor':100,
  'newsplease.pipeline.pipelines.ElasticsearchStorage':350
}
```

The numbers are ordering weights, and the extractor runs first. ElasticSearch is also what enables the revisions feature, which crawls articles repeatedly and tracks what changed between crawls.

## The version is 1.6.13, classified as production stable, with no releases

The packaging metadata declares `version="1.6.13"` and carries the trove classifier `Development Status :: 5 - Production/Stable`, alongside Python 3.8 through 3.12 classifiers and an Apache Software License classifier. The project publishes no GitHub releases, so nothing on the hosting side marks a version for a reader to install, and the branch was last pushed on 2026-04-14. The README does carry a badge linking a Zenodo record with a DOI for citation, which gives the project a citable archive even without tags, though that badge points at an http DOI resolver rather than https. Contributors are asked to read the contribution section first, the project keeps a code of conduct and a separate Dockerfile directory, and the licence file is named `LICENSE.txt` rather than `LICENSE`.

## Extraction output is seven fields, and one of them is a guess

The extracted attributes are a short, fixed list: headline, lead paragraph, main text, main image, the names of authors, publication date and language. Six of those are structural. Publication date is the exception, because it is recovered from the page rather than read from a field the publisher wrote for you, which is why `from_html` accepts the original URL as an optional second argument with the stated purpose of increasing the accuracy of extracting the publishing date. Language comes from detection rather than from a tag, and the requirements file installs two detectors for that job, langdetect and faust-cchardet. A sample output file ships in the repository as JSON so the exact shape of the result object can be read before writing any code against it, and the same object serialises through `get_serializable_dict()` when you want to write JSON yourself.

## Conclusion

Adopt it if you want a working article extractor this afternoon and do not mind that configuration lives in two formats and that some capabilities require a source checkout rather than a pip install. The extraction itself is the strongest part: one call returns headline, lead paragraph, main text, main image, authors, publication date and detected language, and WARC input is supported for archive work. Do not adopt the container image as shipped, since it builds on Python 3.6.5, installs from an unpinned clone of master and leaves a compiler toolchain in the final layer. Before you rely on the crawler, put your own URLs in `sitelist.hjson` first, because running the installed CLI with no arguments starts fetching the example sites. And check the readability pin, which is held at an exact version because a later one is recorded as broken.

## FAQ

### How do I install news-please?

Run `pip install news-please`. The README states it runs on Python 3.8 and the packaging metadata classifies Python 3.8 through 3.12, though the bundled Dockerfile builds on Python 3.6.5.

### What does news-please extract from a news article?

Seven attributes: headline, lead paragraph, main text, main image, the names of the authors, publication date and language. A sample of the extracted JSON is kept in the repository under `newsplease/examples/sample.json`.

### Where does the news-please CLI store its results?

By default in JSON files inside a `data` folder, with the original HTML files stored as well. PostgreSQL, ElasticSearch, Redis or your own storage are alternatives, and ElasticSearch additionally enables the revisions feature that tracks changes across repeated crawls.

### Can I use news-please as a library instead of running the crawler?

Yes. `NewsPlease.from_url`, `from_urls` with optional request arguments, `from_file` for a file of URLs, `from_html` for raw markup with an optional URL, and `from_warc` for a WARC record all work in library mode, and all of them block until every URL has been attempted.

### How do I extract news-please articles from the Common Crawl archive?

Clone the repository, adapt the config section in `newsplease/examples/commoncrawl.py`, then run `python3 -m newsplease.examples.commoncrawl`. You can filter by news publisher and by date period. That path needs a source checkout rather than the installed package.

## Sources

- [fhamborg/news-please on GitHub](https://github.com/fhamborg/news-please)
- [Issues](https://github.com/fhamborg/news-please/issues)
- [License: Apache-2.0](https://github.com/fhamborg/news-please/blob/master/LICENSE)
- [README](https://github.com/fhamborg/news-please/blob/master/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/fhamborg-news-please
