python-goose review: article extraction from HTML in Python
Html Content / Article Extractor, web scrapping lib in Python
At a glance
- What is it?
- python-goose takes a news URL or raw HTML and returns the article body, title, meta description, top image and embedded video. It is an Apache-2.0 port of the Java Goose extractor, and it is opinionated about what a news page looks like.
- Who is it for?
- Adopt python-goose if you are scraping news or article-shaped pages and you want title, cleaned_text, top_image and movies from one call, with language-specific stopword classes for Chinese, Arabic and Korean. Do not adopt it for JavaScript-rendered pages, for sites that require cookie handling, or for arbitrary user-submitted HTML that does not look like an article.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Activity is slowing. The repository last received commits 6 months ago.
- What is it written in?
- Mainly HTML, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem python-goose solves, and who it is for
Scraping a news page is mostly boilerplate removal. You fetch HTML, and then you have to decide which of the dozens of divs is the article and which are navigation, related links, comment widgets and ad slots. python-goose exists to make that decision for you. The README states the aim directly: take any news article or article-type web page and extract the main body of the article plus all metadata and the most probable image candidate.
That framing matters, because it also defines the boundary. This is not a general-purpose HTML parser and not a headless browser. It is a heuristic extractor tuned for pages shaped like news articles. The README lists what it tries to return: main text, main image, any YouTube or Vimeo movies embedded in the article, meta description and meta tags. The audience is developers building feeds, archives, summarisation pipelines or content analysis tools over published journalism.
The project is a complete rewrite in Python of Goose, which the README says was originally a Java article extractor converted to a Scala project. Xavier Grangier wrote the Python port. The licence section says Goose is licensed by Gravity.com under Apache 2.0. The setup.py classifiers mark development status as Beta and list Python 3.8 through 3.12, so the port has kept pace with modern interpreter versions even though the extraction heuristics themselves descend from an older tool.
How extraction works: parsers, scoring and language stopwords
The public surface is small. You construct a Goose object, optionally with configuration, and call extract with a url or with raw HTML. The returned article object exposes title, meta_description, cleaned_text, top_image and movies. That is the whole contract, and it is the reason the library is easy to drop into a script.
The interesting part is underneath. The README says Goose can run on either the lxml html parser or the lxml soup parser, with the html parser as the default. That choice is exposed as a configuration key, parser_class, and it is not cosmetic: the two parsers build different trees from the same markup, so the heuristics downstream see different input. On messy real-world HTML it is worth trying both rather than assuming the default is right for your pages.
Language handling is the other visible mechanism. Goose is described as language aware, and the README shows it picking up the correct meta language tag on a Spanish page and returning Spanish text. When a page has no correct meta language tag, you can force the language with use_meta_language set to False and target_language set to a code. For languages without whitespace word boundaries the library expects a dedicated stopword analyser class passed as stopwords_class. The README documents StopWordsChinese, StopWordsArabic and StopWordsKorean, each imported from goose.text. Chinese is called out as harder than occidental languages because segmentation is more difficult, which is why the class has to be supplied rather than inferred.
Video extraction is a separate path. On a French page the README shows article.movies returning a list of Video objects, each with src, embed_code, embed_type, width and height. The embed_code field gives you the iframe markup back, which is convenient if you are republishing rather than just indexing. Note that the README's examples target YouTube and Vimeo style embeds; there is no statement that arbitrary player markup is recognised.
Installing python-goose and running a first extraction
There is no PyPI install line in the README. Setup starts from the repository itself. The README's commands create a virtualenv, clone the project, install the requirements file and run setup.py. The requirements file lists Pillow, lxml, cssselect, jieba, beautifulsoup4 and nltk, so the install pulls a fairly heavy dependency set for what looks like a small library.
mkvirtualenv --no-site-packages goose
git clone https://github.com/grangier/python-goose.git
cd python-goose
pip install -r requirements.txt
python setup.py installOnce installed, the README's own example is the shortest path to a working call: pass a URL and read attributes off the result.
from goose import Goose
url = 'http://edition.cnn.com/2012/02/22/world/europe/uk-occupy-london/index.html?hpt=ieu_c2'
g = Goose()
article = g.extract(url=url)
article.title
article.meta_description
article.cleaned_text[:150]
article.top_image.srcYou should get back the headline, the meta description string, the first part of the cleaned body text and the URL of the lead image. Configuration is passed either as a Configuration object or as a plain dict, which is the more common style in the README. Changing the user agent and switching parsers happen in the same dict.
g = Goose({'browser_user_agent': 'Mozilla', 'parser_class': 'soup'})For non-Latin scripts, the stopword class is passed the same way. The README's Chinese example imports StopWordsChinese from goose.text and passes it as stopwords_class, then prints the first 150 characters of cleaned_text. If your pages lack correct meta language tags, the same dict takes use_meta_language set to False together with target_language, which forces the analyser regardless of what the page declares.
Where python-goose breaks: unicode URLs and cookie-gated sites
The README has a Known issues section, and it is short and honest. The first entry is unicode URLs. The second is cookie handling: some websites need cookies, and the README says the only workaround at the moment is to use raw_html extraction. That means fetching the page yourself with whatever session and cookie jar you need, then handing the HTML to Goose instead of a URL. It is a real limitation, and it pushes session management onto you.
There is a second class of failure the README implies rather than states. The extractor is built for article-type pages. If you point it at a product listing, a forum thread, a documentation page or a single-page application that renders its content in JavaScript, the heuristics have nothing article-shaped to score. Goose fetches HTML, and the README gives no indication of a browser engine, so client-rendered content is out of scope.
Language is a third boundary. Meta language detection can be wrong or absent, which is why the forced target_language option exists. And for Chinese, Arabic and Korean the default analyser is not enough; you must supply the matching stopword class or the segmentation will not do what you want. Treat the language configuration as a required step for those scripts, not an optimisation.
A fourth, quieter risk is that the returned text is only as good as the heuristics' guess about where the article ends. The README shows cleaned_text truncated to 150 characters in every example, which tells you nothing about whether trailing captions, bylines or pull quotes survive intact. That is the first thing to check on your own corpus.
python-goose compared with trafilatura and readability-style extractors
The closest comparison in this space is trafilatura, a Python library built specifically for text and metadata extraction from web pages. The difference in approach is scope. python-goose returns a bundle: cleaned text plus top image plus embedded movies plus meta tags, with the image and video handling wired into the same call. Trafilatura concentrates on the text and metadata side and is designed around large-scale crawling and evaluation against reference corpora. If your pipeline is text-only and you care about extraction quality measurement, trafilatura's focus is the better fit. If you need the lead image URL and the video embed alongside the body, python-goose does that in one pass.
Readability ports are the other common option. Those descend from the browser reading-mode tradition, where the goal is a clean reading view, and they typically return text and a title without a first-class image or video object. python-goose's article.movies with src, embed_code, embed_type, width and height is a genuinely different output shape from what a readability port gives you.
One caveat on the comparison itself: python-goose is a port of an older Java and Scala extractor, and its heuristics were tuned against pages from that era. Trafilatura is a newer codebase. That does not automatically make either one more accurate on your corpus, but it does mean the age of the scoring rules is a variable you should test rather than assume. The practical way to decide is to run both over a sample of your own URLs and diff the cleaned_text output; the README gives no accuracy figures, so there is no published number to lean on.
Maintenance, licence and what upgrading costs
The repository is not archived, and the last push was on 2026-03-10. The README does not describe a release process and no recent releases were retrieved. There is no published changelog, and no version number is quoted in the README. Upgrading therefore means tracking the develop branch, which is the default branch, rather than pinning to tagged releases. That is a real cost: you get the newest fixes, and you also absorb whatever lands on develop between your pins.
The dependency list is the other upgrade surface. requirements.txt pins nothing, listing Pillow, lxml, cssselect, jieba, beautifulsoup4 and nltk without version constraints. lxml and Pillow both ship compiled extensions, so a fresh install on a new Python version depends on wheels being available for that platform. jieba and nltk are pulled in for the text analysis path, which means the Chinese segmentation support is part of the base install even if you never scrape Chinese pages.
On licensing: the README states Goose is licensed by Gravity.com under the Apache 2.0 license and points to the LICENSE file. setup.py carries the Apache Software License classifier and refers to a NOTICE file for contributor copyright information. Apache 2.0 is a permissive licence with an explicit patent grant and notice requirements. This is not legal advice; if you are redistributing the library or its notices inside a product, read LICENSE.txt and the NOTICE reference in setup.py yourself.
Who should adopt python-goose and who should not
Adopt it if your input is published journalism or article-shaped pages and you want one call to return the body text, the headline, the meta description, the lead image and any embedded video. The language stopword classes for Chinese, Arabic and Korean are the strongest reason to pick this over a generic readability port, because they are shipped in the library rather than left to you.
Do not adopt it if your pages are client-rendered, if the site gates content behind cookies, or if the target is not an article at all. The cookie case has a documented escape hatch, raw_html extraction, but that moves fetching and session handling into your own code. The JavaScript case has no escape hatch in the README at all.
What to verify first is narrow and testable: run the extractor over a sample of your own URLs, check that cleaned_text does not clip the tail of long articles, confirm whether your pages declare a correct meta language tag or need use_meta_language set to False, and try parser_class set to 'soup' against the default html parser to see which produces cleaner output on your markup.
Editorial conclusion
Adopt python-goose if you are scraping news or article-shaped pages and you want title, cleaned_text, top_image and movies from one call, with language-specific stopword classes for Chinese, Arabic and Korean. Do not adopt it for JavaScript-rendered pages, for sites that require cookie handling, or for arbitrary user-submitted HTML that does not look like an article. Before committing, verify the behaviour on your own target URLs: check whether your pages carry correct meta language tags, whether the site needs cookies, and whether the default lxml html parser or the soup parser gives better output on your markup. The repository's last push was on 2026-03-10, so the code is not abandoned, but the README itself lists unicode URL problems and cookie handling as known issues.
Frequently asked questions
How do I install python-goose?
The README does not give a PyPI command. It clones the repository, runs pip install -r requirements.txt and then python setup.py install, optionally inside a virtualenv created with mkvirtualenv. The requirements file lists Pillow, lxml, cssselect, jieba, beautifulsoup4 and nltk.
Does python-goose work on Chinese, Arabic or Korean pages?
Yes, but you must pass a dedicated stopword analyser class through the configuration. The README documents StopWordsChinese, StopWordsArabic and StopWordsKorean, all imported from goose.text, and notes that Chinese segmentation is harder than occidental languages, which is why the class has to be supplied.
Can python-goose extract images and videos from an article?
The README says Goose tries to extract the main image of an article and any YouTube or Vimeo movies embedded in it. The article object exposes top_image with a src attribute, and movies returns Video objects carrying src, embed_code, embed_type, width and height.
What are the known issues with python-goose?
The README's Known issues section lists problems with unicode URLs and cookie handling. It states that some websites need cookies and that the only workaround at the moment is to use raw_html extraction, which means fetching the page yourself and passing the HTML in.
Which HTML parser does python-goose use by default?
The README says Goose can run with the lxml html parser or the lxml soup parser, and that the html parser is used by default. You switch parsers by passing parser_class set to 'soup' in the configuration dict.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/grangier-python-goose)