Open-source project
chardet/chardet avatar
chardet/chardet

chardet 7: what changed in the encoding detector rewrite

Python character encoding detector

2,676 stars307 forksPython0BSD

At a glance

What is it?
chardet 7 replaces the library Mark Pilgrim started in 2006 with a from-scratch rewrite that keeps the package name and the public API but swaps the licence to 0BSD, returns language and MIME type with every guess, and adds era filtering plus an optional compiled build.
Who is it for?
chardet 7 fits code that needs an encoding verdict on bytes it already holds, or a stream it can feed a line at a time. Two things are worth checking before you upgrade an existing pin.
Can I use it commercially?
Yes. 0BSD is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 37 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 4, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The same package name now installs different code

Since version 7, installing chardet no longer installs the library Mark Pilgrim wrote. What ships now is a ground-up rewrite under the same package name and the same public API, positioned as a drop-in replacement for 5.x and 6.x, with the licence underneath changed from LGPL to 0BSD. The history section makes the handoff explicit. Pilgrim created the project in 2006 as a Python port of Mozilla's universal charset detection library, published 1.0 that year and 1.0.1 in 2008, then worked on an unreleased Python 3 port numbered 2.0.1 hosted on Google Code. After he deleted his online accounts in 2011 the project was carried on by David Cramer, Erik Rose, Toshio Kuratomi, Ian Cordasco and Dan Blanchard. The closing sentence of that history, the one beginning In 2026, Dan Blanchard rewrote c, stops mid-word, so the account of the rewrite itself is left to a separate blog post linked from the top of the page rather than told in the file.

The comparison table measures 7.6.1.dev while the newest tag is 7.6.0

The benchmark block is the project's own, and its first column is headed chardet 7.6.1.dev (compiled), a version with no tag behind it: the latest releases are 7.6.0, 7.5.1 and 7.5.0, all three published on 14 August 2026 within one second of each other. That column reports 99.7% accuracy over 3,138 test files against 84.4% for chardet 6.0.0 and 86.6% for charset-normalizer 3.5.1, language detection at 91.8% against 38.7% and 54.8%, and 99 supported encodings against 84 for the old chardet. The speed row reads 2,641 files/s compiled with 641 files/s in pure Python, next to 9 files/s for chardet 6.0.0 and 2,250 files/s for charset-normalizer. Dividing those two published speed numbers gives a smaller factor than the 315x claimed in the sentence above the table, so the headline and the grid are measurements from different runs rather than one figure quoted twice. Peak memory is listed as 27.7 MiB against 71.0 MiB, and the last three rows, streaming detection, encoding era filtering and MIME type detection, are marked no for both competitors. Behind those numbers sits a 13-stage pipeline that names its own steps: BOM detection, magic number identification, structural probing, byte validity filtering and bigram statistical models, with the 91.8% language figure covering 49 languages. The optional compiled build is the second lever on the speed row, mypyc plus a Cython scoring kernel described as 4.8x additional speedup on CPython, which is why the first column is labelled compiled while the pure Python figure shares the same cell. Every figure here is the project describing itself, which makes the table a claim about a dev build rather than a neutral comparison.

Large input is validated over every byte instead of sampled

Big buffers are the case the full-file call was reshaped for. `detect(data, max_bytes=len(data))` on a 272 MiB file is reported to finish in about 0.13 to 0.23 seconds, and the UTF-8 verdict is validated over every byte examined rather than over a prefix, so passing the whole buffer does not trade accuracy for speed. Combined with the 27.7 MiB peak in the table above, that is what makes the call usable on a request path. The everyday version is one line, and the return value is a dict with four keys where 5.x returned encoding and confidence:

python
import chardet

chardet.detect(b"Python is a great programming language for beginners and experts alike.")
# {'encoding': 'ascii', 'confidence': 1.0, 'language': 'en', 'mime_type': 'text/plain'}

Code that reads the old keys by name keeps working. Code that unpacks the result as a pair, or that iterates the dict expecting two items, does not, and that break is the one upgrade cost the drop-in claim does not cover.

Era filtering deletes candidates instead of rescoreing them

`detect_all()` returns ranked candidates, and a Russian sentence encoded in windows-1251 shows the default shape: Windows-1251 at 0.46, MacCyrillic at 0.42, KZ1048 at 0.2 and ptcp154 at 0.2, four answers pulled from four different eras. Passing `encoding_era=EncodingEra.MODERN_WEB` reduces that to the single Windows-1251 line, and the winner keeps its 0.46. Nothing is remeasured; the filter removes candidates from the set, so a result that barely beat three rivals can be promoted by narrowing the era rather than by matching better. For the legacy families in that candidate list, which is where the value sits, since MacCyrillic and ptcp154 have no business in a web response and the tie disappears once they are gone. The markdown of this example is also broken at the end: the fence is closed by a lone backtick on its own line instead of three, so the block runs straight into the next heading when it renders.

include_encodings and exclude_encodings are a separate lever

Narrowing by era and narrowing by encoding name are different operations, and the named form is what handles EBCDIC. Both parameters sit on the ordinary call:

python
# Only consider UTF-8 and Windows-1252
chardet.detect(data, include_encodings=["utf-8", "windows-1252"])

# Consider everything except EBCDIC
chardet.detect(data, exclude_encodings=["cp037", "cp500"])

With 99 encodings in the candidate set, covering the EBCDIC, Mac, DOS and Baltic families, an era filter alone cannot express a web-only policy whenever an era still holds code pages nobody serves. Both lines in the example refer to a variable named `data` that the snippet never defines, so the pair is lifted out of a longer session rather than runnable on its own. The command line exposes the include list through `-i`, which takes a comma separated string, and there is no era option there at all, so a pipeline driven from the shell can filter by name but not by era.

UniversalDetector stops reading once the verdict is settled

For a file or a socket that should not be held in memory, the streaming path takes one chunk at a time and stops on its own:

python
from chardet import UniversalDetector

detector = UniversalDetector()
with open("unknown.txt", "rb") as f:
    for line in f:
        detector.feed(line)
        if detector.done:
            break
result = detector.close()
print(result)

The break is the point of the class: `detector.done` turns true once enough evidence has accumulated and `close()` hands back the result. Feeding a line at a time instead of a large block makes that early exit arrive sooner, which is the behaviour the streaming row records as yes for both chardet versions and no for charset-normalizer. `detect()` and `detect_all()` are also called out as safe to use concurrently and as scaling on free-threaded Python, which matters when a web worker detects encodings per request. What the streaming path cannot give you is a ranking while you are still reading: candidates only come from `detect_all()` on data you already hold. How many bytes a verdict cost stays inside that loop too, since nothing in the result reports it back.

chardetect prints one verdict line and no version

The console script is named `chardetect` rather than `chardet`, a leftover from the package's older layout, and the packaging declares exactly one entry point at `chardet.cli:main`. What it prints is fixed:

bash
chardetect somefile.txt
# somefile.txt: utf-8 with confidence 0.99

chardetect --minimal somefile.txt
# utf-8

# Include detected language
chardetect -l somefile.txt
# somefile.txt: utf-8 en (English) with confidence 0.99

# Only consider specific encodings
chardetect -i utf-8,windows-1252 somefile.txt
# somefile.txt: utf-8 with confidence 0.99

# Pipe from stdin
cat somefile.txt | chardetect
# stdin: utf-8 with confidence 0.99

Input from a pipe is labelled stdin rather than a path, so a loop over files gives one line per file with its source named. Five behaviours are shown and none of them prints a version, the candidate list, an era filter or a machine readable format, so anything past that single verdict line is written against the Python API instead. The flag vocabulary is also narrower than the library's: `-i` maps to include_encodings, and nothing maps to exclude_encodings or to EncodingEra.

The version number is read out of git tags at build time

Nothing in the tree holds a version string to edit. The project metadata declares `dynamic = ["version"]`, points the version source at `source = "vcs"`, and installs a build hook that writes `src/chardet/_version.py` while building, with hatchling and hatch-vcs as the only build requirements. A checkout with no tags therefore has no version to report, which is the cost of not keeping a number in two places by hand. The wheel packages `src/chardet`, and the sdist excludes `tests/data/` and `data/`, so the corpus behind the 3,138 file accuracy figure is not in the sdist a user downloads. Runtime requirements are `requires-python = ">=3.10"` with no dependency list at all, matching the zero-dependency claim and the PyPy support, and the classifiers stop at Python 3.14 while still carrying `Development Status :: 5 - Production/Stable`. Two top-level files have no counterpart in a library of this shape: `hatch_build_cython.py`, which drives the optional mypyc and Cython build behind the 4.8x compiled speedup, and `prek.toml`, a repository hook configuration that sits beside the `.claude/` and `.devcontainer/` directories. Nothing in that metadata pins the mypyc path, so a plain install from PyPI lands on the pure Python figure rather than the compiled one.

Editorial conclusion

chardet 7 fits code that needs an encoding verdict on bytes it already holds, or a stream it can feed a line at a time. Two things are worth checking before you upgrade an existing pin. The library behind the name is new code, not the 5.x implementation, so a version range that crossed 5 to 7 picked up different candidate ranking and two extra keys in the result dict. And the headline numbers, including the 315x speed claim, come from the project measuring a 7.6.1.dev build against chardet 6.0.0. If your pipeline depends on encoding lists rather than a single best guess, check the include_encodings and exclude_encodings parameters against your own data before switching, since a detector with 99 candidates can return a different family than the one you got before.

Frequently asked questions

What is Chardet and how is it used in Python?

It is a character encoding detector for Python, rewritten from the ground up in version 7 under a 0BSD licence. You install it with `pip install chardet`, import the package, and call `chardet.detect()` on a bytes object, which returns the encoding, a confidence value, the detected language and, for binary input, a MIME type.

how to use chardet in python

Call `chardet.detect()` on a bytes object for one verdict, `chardet.detect_all()` when you want ranked candidates, or feed a file line by line to `chardet.UniversalDetector` and read the answer from `close()`. The package needs Python 3.10 or newer and declares no runtime dependencies.

how to install chardet

`pip install chardet` pulls the current release from PyPI. It targets Python 3.10 and above, ships no runtime dependencies, and the project states that it works on PyPy as well as CPython.

chardet vs charset normalizer

On the project's own figures, chardet 7.6.1.dev compiled reports 99.7% accuracy over 3,138 files against 86.6% for charset-normalizer 3.5.1, 2,641 files/s against 2,250 files/s, 27.7 MiB peak memory against 71.0 MiB, and streaming detection, which the comparison marks as absent in charset-normalizer.

what does chardet do

It guesses the character encoding of a block of bytes and returns that guess with a confidence value. Every result also carries a detected language, and binary files get a MIME type read from magic number signatures across 40 or more formats, plus markup types such as text/html and text/xml.

Official sources

  1. chardet/chardet on GitHub
  2. Issues
  3. License: 0BSD
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/chardet-chardet.svg)](https://hysenlabs.com/projects/chardet-chardet)