Library / SDK
clips/pattern avatar
clips/pattern

clips/pattern: a Python web mining module that is archived but still installable

Web mining module for Python, with tools for scraping, natural language processing, machine learning, network analysis and visualization.

8,859 stars1,551 forksPythonBSD-3-Clause

At a glance

What is it?
Pattern bundles scraping, part-of-speech tagging, sentiment, WordNet, a vector space model and graph analysis into one Python package. The README now opens with a warning that the repository is no longer maintained, which changes how you should use it.
Who is it for?
Pattern is worth adopting only for prototypes, teaching, or research where its bundled taggers, WordNet and vector model save you from wiring several libraries together, and where you can pin your own environment. Do not adopt it for anything that needs security patches, current Python versions, or a maintained Twitter and Google client, because the README states the repository is no longer maintained and the newest release is 3.7-beta from 2022-08-30.
Can I use it commercially?
Yes. BSD-3-Clause is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 56 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 26, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What clips/pattern solves, and who it is for

Pattern is a web mining module for Python. The README lists four groups of tools: data mining (web services for Google, Twitter and Wikipedia, a web crawler, an HTML DOM parser), natural language processing (part-of-speech taggers, n-gram search, sentiment analysis, WordNet), machine learning (a vector space model, clustering, classification with KNN, SVM and Perceptron), and network analysis (graph centrality and visualization). The value is that these ship in one package rather than as four separate dependencies. A researcher who wants to crawl pages, tag the text, turn it into vectors and cluster the result can do that without reconciling four libraries and four data formats. The audience is the same as it was when the package was published: students, linguists and researchers writing short scripts, and engineers building a one-off pipeline to check whether an idea is worth pursuing. The README's own example is aimed at that reader. It trains a classifier on adjectives mined from Twitter, keeping only words tagged JJ (adjective) and labeling each tweet WIN or FAIL. That single script touches the Twitter client, the tagger and the KNN classifier, which is the pitch in miniature.

How the modules fit together: taggers, vectors, graphs and the corpus

The architecture visible in the repository is a set of submodules under pattern/, each with its own job, plus bundled data. The README's dependency list names the algorithms and datasets that ship inside the package: Brill taggers for English, Dutch, German, Spanish, French and Italian, English pluralization, Spanish and French verb inflection, LIBSVM, LIBLINEAR, NetworkX centrality code, a spelling corrector, and a Graph JavaScript framework. That is why installation is heavier than a pure-Python library: the taggers and the vector classifiers carry model data, and the graph module carries JavaScript for rendering. The data flow in the README example is linear. A Twitter search returns tweet objects. tweet.text is lowercased, then passed to pattern.en.tag, which returns (word, pos) pairs. The list is filtered to adjectives, count() turns it into a dictionary of adjective to frequency, and that dictionary is the vector the KNN classifier trains on. Classification then takes a plain string and returns the learned label. The same vector representation is what the clustering and SVM code use, so the vector space model is the hinge between the NLP side and the machine learning side. The graph side is separate: centrality is computed in Python and the visualization is rendered by the bundled JavaScript framework, which is why the repository has both a pattern/ package and a docs/ folder with a schema image.

Installing pattern and running the Twitter classifier

The README gives two installation paths. The first is to unzip the download and run setup.py from the command line; the second is pip. Note the directory name in the first path, which encodes the Python version the release targets.

bash
cd pattern-3.6
python setup.py install

If pip is available, the README says you can download and install from PyPI instead.

bash
pip install pattern

There is a third fallback documented for cases where neither works: place the pattern folder next to your script, copy it into the site-packages directory for your interpreter (the README lists c:\python36\Lib\site-packages\ on Windows, /Library/Python/3.6/site-packages/ on Mac OS X and /usr/lib/python3.6/site-packages/ on Unix), or append the module location to sys.path before importing.

python
MODULE = '/users/tom/desktop/pattern'
import sys; if MODULE not in sys.path: sys.path.append(MODULE)
from pattern.en import parsetree

For a first real use, the README's example is the shortest path through the package. It collects tweets containing #win or #fail, keeps the adjectives, and trains a KNN classifier on adjective counts.

python
from pattern.web import Twitter
from pattern.en import tag
from pattern.vector import KNN, count

twitter, knn = Twitter(), KNN()

for i in range(1, 3):
    for tweet in twitter.search('#win OR #fail', start=i, count=100):
        s = tweet.text.lower()
        p = '#win' in s and 'WIN' or 'FAIL'
        v = tag(s)
        v = [word for word, pos in v if pos == 'JJ'] # JJ = adjective
        v = count(v) # {'sweet': 1}
        if v:
            knn.train(v, type=p)

print(knn.classify('sweet potato burger'))
print(knn.classify('stupid autocorrect'))

What you should see is two printed labels. The first phrase contains a positive adjective and the second a negative one, so if the training data was collected and labeled as the loop intends, the two outputs should differ. If both print the same label, the usual cause is that the search returned nothing usable and the classifier has too few vectors; the loop only trains on tweets where at least one adjective survived the filter.

The maintenance warning is the first thing in the README

The README begins with a warning block: the repository is no longer maintained, the project is archived, and it will not receive further updates, bug fixes or security patches, with issues and pull requests possibly unreviewed. The repository metadata on the page says Archived: no, and the last push was on 2026-08-05, but the README text is the project's own statement about its status and it is unambiguous. Anyone deciding whether to depend on Pattern should treat the README warning as authoritative over the badge. The practical consequences are concrete. The newest release listed is 3.7-beta from 2022-08-30, while the README's Version section says 3.6, so the documented version and the published release do not match. The supported interpreters named in the README are Python 2.7 and Python 3.6. Python 3.6 reached end of life years ago, and the setup.py directory name pattern-3.6 reinforces that this is the target. A security patch for a bundled dependency, or a fix for a site change that breaks the Twitter or Google client, will not arrive. The web services module is the most exposed part: it talks to third-party endpoints whose request formats and authentication rules change independently of this package. The NLP, vector and graph modules are less exposed because they run locally, but they still carry the interpreter constraint.

Where Pattern is the wrong tool

Pattern is the wrong choice when your pipeline depends on a live third-party API. The README lists web services for Google, Twitter and Wikipedia. Those clients were written against the APIs of their time, and the README's own example relies on an unauthenticated Twitter search that returns tweets by hashtag. If your project needs current API behavior, you are asking an archived package to track a moving target, and the README says no updates will come. It is also the wrong tool when you need a modern NLP stack. Pattern's taggers are Brill taggers bundled with the package, and the README presents them as part of the value. If your work requires transformer models, tokenizers with subword vocabularies, or GPU inference, nothing in the README describes that, and the package predates it. The third case is deployment on current Python. If your environment is pinned to a recent interpreter and you cannot add an older one, the documented support for Python 2.7 and 3.6 is a real constraint, not a footnote. Finally, if you need a maintained dependency for a compliance review, an archived project with a no-security-patches notice is a poor fit regardless of how well the code works.

NLTK, spaCy and scikit-learn as alternatives, and what actually differs

The honest comparison is not one library against another but one design decision against another. Pattern's decision is bundling: taggers, WordNet, a vector space model, classifiers, a crawler and graph centrality in one install, with a shared vector representation. NLTK makes the opposite choice. It is a toolkit of separate components with its own corpora and its own download step, and it does not ship a web crawler or a graph renderer as part of the same package. If you want to swap tokenizers or bring your own model, that separation helps; if you want the README's four-line path from tweet text to a trained classifier, it costs you assembly work. spaCy takes a third position: it is built around pretrained statistical pipelines and a fixed processing chain, which gives you faster and more accurate tagging than a bundled Brill tagger, but it does not include WordNet, a crawler, or graph centrality, so the Pattern example would become several dependencies. scikit-learn covers the machine learning half of Pattern's feature list with a much wider set of estimators, but it has no NLP or web mining layer at all. The practical difference is that Pattern is a single script's worth of integration and the alternatives are a stack you compose. For a one-off research script, that integration is the whole point. For a production service, the composition is what lets you replace the parts that go stale.

Licence, upgrade cost and what the repository tells you about both

Pattern is licensed under BSD-3-Clause, and the README links to LICENSE.txt for the full text. The README also states that the source code is licensed under BSD. That is a permissive licence, which matters here because the package bundles third-party algorithms and datasets: Brill taggers for six languages, LIBSVM, LIBLINEAR, NetworkX centrality code, a spelling corrector and a Graph JavaScript framework. The README lists these under Bundled dependencies with attribution, but it does not state the licence of each bundled component individually. If you redistribute Pattern or ship it inside a product, read LICENSE.txt and the individual attributions rather than assuming the top-level BSD-3-Clause covers everything; that is a question for your own legal review, not something the README settles. The upgrade cost is the mirror image of the maintenance warning. There is no upgrade path to plan for, because the README says the project is archived and will receive no further updates. The newest published release is 3.7-beta from 2022-08-30 while the README documents version 3.6, so pinning is not optional: record the exact version you install and keep the artifact. The repository layout gives you one more lever. Because the README documents a sys.path fallback, you can vendor the pattern folder next to your script instead of installing into site-packages, which keeps the dependency visible in your tree and avoids fighting an installer that expects Python 3.6.

Editorial conclusion

Pattern is worth adopting only for prototypes, teaching, or research where its bundled taggers, WordNet and vector model save you from wiring several libraries together, and where you can pin your own environment. Do not adopt it for anything that needs security patches, current Python versions, or a maintained Twitter and Google client, because the README states the repository is no longer maintained and the newest release is 3.7-beta from 2022-08-30. Before writing code, verify one thing: that pip install pattern succeeds on your interpreter and that from pattern.en import parsetree imports, since the README lists Python 2.7 and Python 3.6 as the supported versions and later interpreters are not covered.

Frequently asked questions

How do I install the clips/pattern Python module?

The README gives two paths: unzip the download and run python setup.py install from the pattern-3.6 directory, or run pip install pattern to fetch it from PyPI. If neither works, the README describes placing the pattern folder next to your script, copying it into the site-packages directory for your interpreter, or appending its location to sys.path before importing.

Is clips/pattern still maintained?

The README opens with a warning that the repository is no longer maintained, that the project is archived, and that it will not receive further updates, bug fixes or security patches. The newest release listed is 3.7-beta from 2022-08-30, and the README's Version section says 3.6.

Which Python versions does clips/pattern support?

The README states that Pattern supports Python 2.7 and Python 3.6. The installation example uses a directory named pattern-3.6, and the site-packages paths it lists are for Python 3.6 on Windows, Mac OS X and Unix.

Official sources

  1. clips/pattern on GitHub
  2. License: BSD-3-Clause
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/clips-pattern.svg)](https://hysenlabs.com/projects/clips-pattern)