Library / SDK
pemistahl/lingua-py avatar
pemistahl/lingua-py

Lingua's accuracy data is single-language only, and its two badges point at the Rust repository

The most accurate natural language detection library for Python, suitable for short text and mixed-language text

1,806 stars62 forksPythonApache-2.0

At a glance

What is it?
Lingua is a natural language detector for Python, Apache licensed, published as lingua-language-detector on the package index and built since version 2 as compiled bindings to a Rust implementation. The repository commits its benchmarks as images, tables and test corpora, and the visible accuracy data never mixes languages.
Who is it for?
Lingua suits a preprocessing step that needs a language label on short text without a network call and without loading a machine learning framework, and the 1.x line stays available for environments that cannot load a native extension. It suits a mixed-language corpus less well than its repository description implies, because the bundled accuracy data is single-language throughout and the page never says how code-switched text is measured.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 75 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 3, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The package name, the repository name and the two badges all point somewhere different

Three identifiers are in play. The repository is lingua-py, the distribution name on the package index is lingua-language-detector, and the native implementation it wraps lives in a separate repository called lingua-rs.

The build metadata in the project file settles the install story:

toml
[tool.poetry]
name = "lingua-language-detector"
version = "2.2.0"
description = "An accurate natural language detection library, suitable for long and short text alike"

The version there matches the newest release, and the homepage and repository fields both point back at the Python repository, even though the repository's own homepage field is left empty.

The two badges at the top of the page are the clearest sign of how the work is split. The build badge links to a workflow inside lingua-rs, and the coverage badge links to that repository's coverage page. So the continuous integration and coverage evidence shown for a Python package are the Rust project's, which follows from the architecture rather than from an error, but it does mean the badge pair tells you nothing about the Python bindings themselves.

Three descriptions, and the mixed-language claim appears in only one of them

The project describes itself three different ways.

The repository description calls it the most accurate natural language detection library for Python, suitable for short text and mixed-language text. The distribution description calls it an accurate natural language detection library, suitable for long and short text alike, with no superlative and no mention of mixed text. The page itself opens by saying its task is simple: it tells you which language some text is written in.

Read together, the mixed-language promise exists in exactly one place, and the superlative claim is dropped everywhere it is not marketing copy. Nothing in the body of the page discusses code-switched text, a text with sentences in two languages, or how such input would be handled. The two stated technical advantages of the method are about method rather than mixture: it combines rule-based and statistical naive Bayes, and it uses neither neural networks nor dictionaries of words.

What it does promise is offline operation. No external service is contacted, and once the package is downloaded it works without a connection, which for a preprocessing step is often the deciding feature rather than accuracy.

The bundled accuracy data contains no text in two languages at all

The accuracy section is built on test data shipped with the library, and it is split three ways for every supported language: a list of single words, a list of word pairs, and a list of complete grammatical sentences of various lengths.

All three parts are single-language by construction. There is no mixed set, no code-switched set, and no sentence assembled from two sources. So the claim in the repository description about mixed-language text is not measured anywhere in the data the project publishes.

The data itself is well specified, and the specification is worth reading because it is unusually careful. The language models and the test data were built from separate documents of the Wortschatz corpora offered by Leipzig University. Training text was crawled from various news websites, with one million sentences per language corpus. Test corpora were drawn from arbitrarily chosen websites at ten thousand sentences each, and from each one a random unsorted subset of a thousand single words, a thousand word pairs and a thousand sentences was taken.

One thing that follows from the training source and is not commented on anywhere: news text is a specific register, and a detector trained on it inherits that register's vocabulary and its sentence length distribution.

No word under five characters is in the benchmark, which is the hard band for short text

The single-word test set has a minimum length of five characters, and the word-pair set has a minimum length of ten. The sentence set has no length rule beyond being a complete grammatical sentence.

That lower bound removes the band where short-text detection is hardest. The two to and four character words are the ones where a detector has little signal to work with and where closely related languages collide, and none of them appear in the published benchmark. For a library whose central argument is that it works on short text, the short-text measurement starts at five characters.

The comparison itself has a second structural caveat, stated in the page. Detection was run for this library and for four alternatives over the data for the languages this library supports, and languages that an alternative does not support are simply ignored for it. So every alternative is measured on the subset it can handle, and the subsets differ, which is a fair thing to do and worth knowing when reading a bar that shows an alternative scoring well on a language with a short test set.

The result presentation is described rather than tabulated: a bar plot per language for detailed accuracy, and box plots per classifier where the box is the middle half of the values and the horizontal line marks the median.

Two Python lines ship side by side, and the 1.x one is still being cut

The history of the library is a short account of two failed storage strategies followed by a rewrite. Language models were first held in dictionaries at runtime, which was fast and cost more than three gigabytes of memory. They were then held in NumPy arrays, which brought memory down to roughly eight hundred megabytes and cost a large amount of CPU time. Neither was acceptable, so from version 2.0.0 the pure Python implementation was replaced by compiled bindings to the Rust implementation.

The pure Python code is not dead. It lives on a separate branch in the same repository and the page says it will be kept up to date in subsequent 1.x releases, with the stated reason that some environments do not support native extensions, a mobile runtime named as one example. Both lines remain on the package index.

The release list shows the dual line being maintained deliberately rather than as a leftover. The 1.4.2 and 2.1.1 tags were published on the same day about half an hour apart, which is what cutting a maintenance release and a feature release together looks like. A reader who needs the pure Python path has to know that 1.4.2 is the ceiling of it, and that the newest feature release is two minor versions ahead on the other line.

The benchmarks are committed as directories, and the page shows empty placeholders

The accuracy evidence lives in the repository rather than in the prose. Four root directories carry it: a directory of accuracy reports, a directory of tables, a directory of language test data, and a directory of images.

That layout has a visible consequence in the page itself. Each of the three accuracy subsections, for single words, for word pairs and for sentences, opens with a heading, some line breaks and a collapsed element whose summary reads bar plot, with nothing after it. The plots are referenced rather than reproduced as text, and no numbers appear in the page at all.

So the strongest claim the project makes, that it is more accurate than four named alternatives across 75 languages, is not checkable from the page. It is checkable from the repository, and a reader willing to open the tables directory gets per-language figures rather than a single verdict.

The rest of the project is conventional. The build is Poetry based with a lockfile committed, there is a release notes file at the root, and the licence file carries a name that differs from the usual one while the package metadata names the Apache licence and the classifier list repeats it.

The language list is sorted by endonym, so Bokmal sits under B and Nynorsk under N

The list of supported languages is alphabetised by the language's own name rather than by its English name, and the effect is visible in the headings. Norwegian Bokmal appears under the letter B, while Norwegian Nynorsk appears under N. The list uses Persian where English usage would often say Farsi, and Slovene where it would say Slovenian.

The count is exact. The page claims seventy five languages and the list holds seventy five, with one entry for Chinese.

That last point is where the package metadata runs ahead of the list. The classifier list in the project file names both Chinese (Simplified) and Chinese (Traditional) as separate natural language values, and also lists Esperanto, while the language list itself has a single Chinese entry. A classifier is a hint to an index crawler, so the mismatch costs nothing at runtime, but it does mean the metadata advertises a granularity the detector does not expose.

The stated philosophy is quality over quantity, described as getting detection right for a small set of languages before adding new ones. Applied to a list that includes Latin, Esperanto, Welsh, Maori and a cluster of southern and eastern African languages, it is a claim about the training data rather than about the size of the list.

The library is personified throughout, and the pages describe the person, not the API

The page writes about the library in the feminine and does so throughout. It says she aims at eliminating these problems, that she nearly does not need any configuration, and that she yields accurate results on long and short text. The rest of the writing follows, with a short history section, a section on which languages are supported, a numbered accuracy section, and a comparison against named alternatives.

What is missing from all of it is the interface. Nothing in the visible page shows how detection is called, what the return value looks like, how a language is selected or excluded, or how a low-confidence result is represented. The page describes what the library decides, on what data, and with what method, and leaves the call surface to the documentation elsewhere.

The two stated drawbacks of the alternatives are worth repeating because they define the target. Detection that needs lengthy fragments fails on the length of a short message, and accuracy falls as more languages take part in the decision. Both are problems of a design that scores every supported language and then picks a winner, and a library that claims seventy five languages has to answer the second one in a different way rather than by supporting fewer languages.

Editorial conclusion

Lingua suits a preprocessing step that needs a language label on short text without a network call and without loading a machine learning framework, and the 1.x line stays available for environments that cannot load a native extension. It suits a mixed-language corpus less well than its repository description implies, because the bundled accuracy data is single-language throughout and the page never says how code-switched text is measured. Before adopting it, check four things. Decide which line you want, since 2.x needs a native extension and 1.x is pure Python on a separate branch. Count the languages you actually need against the list, which is sorted by endonym rather than by English name. Read the committed accuracy reports rather than the summary sentence, since the summary makes a superlative claim and the reports are per language. And treat any mixed-language requirement as unverified until you measure it on your own data.

Frequently asked questions

How do I install Lingua for Python?

The distribution name on the package index is lingua-language-detector, taken from the build metadata, while the repository is named lingua-py. From version 2.0.0 the package is compiled Python bindings to a Rust implementation, so an environment without native extension support needs the 1.x line instead.

How many languages does Lingua detect?

Seventy five, and the published list holds exactly seventy five entries. The list is sorted by endonym, so Norwegian Bokmal appears under the letter B and Norwegian Nynorsk under N, and Chinese appears as a single entry even though the package classifier list names both Simplified and Traditional variants.

Does Lingua handle text that mixes two languages?

The repository description says the library is suitable for short text and mixed-language text, but the bundled accuracy data is single-language throughout: single words, word pairs and complete sentences, with no mixed set. The page does not describe how code-switched text is measured.

What method does Lingua use for detection?

It combines rule-based and statistical naive Bayes methods, and it uses neither neural networks nor dictionaries of words. It also needs no connection to an external service, so it runs entirely offline once the package is downloaded.

How was Lingua's accuracy measured?

Test data is bundled per language and split into single words of at least five characters, word pairs of at least ten characters, and complete grammatical sentences. The models and the test data were built from separate documents of the Wortschatz corpora from Leipzig University, with training text crawled from news sites at one million sentences per language and test subsets of a thousand items per category.

Which alternatives does Lingua compare itself against?

Four are named: two Google compact language detectors, a language identification library, a simplemma library and langdetect. The page attributes two drawbacks to all but the last, namely that detection needs lengthy text and that accuracy falls as more languages take part, and it notes that languages an alternative does not support are ignored for it during the comparison.

Official sources

  1. Issues
  2. License: Apache-2.0
  3. pemistahl/lingua-py on GitHub
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/pemistahl-lingua-py.svg)](https://hysenlabs.com/projects/pemistahl-lingua-py)