Library / SDK
codespell-project/codespell avatar
codespell-project/codespell

codespell: a dictionary of common misspellings, not a spell checker

check code for common misspellings

2,430 stars530 forksPythonGPL-2.0

At a glance

What is it?
A single Python package with no runtime dependencies that catches adn and wrod, ships its own dictionaries, and gives you four different ways to say not this one.
Who is it for?
codespell earns its place by knowing what it is not trying to do. It holds a curated list of common misspellings rather than a complete dictionary of the language, so it catches `adn` and ignores `adnasdfasdf`, and that design choice is what keeps it from flagging the jargon in your own codebase.
Can I use it commercially?
Yes, with conditions. GPL-2.0 is a copyleft licence: if you distribute software that includes it, you must release that software's source code under the same licence. Running it internally without distributing it does not trigger that obligation.
Is it still maintained?
Yes. The repository last received commits 8 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 28, 2026, and from our analysis. They are not legal advice.

Editorial analysis

A typo list, not a dictionary lookup

The README states the design constraint in its first paragraph, and it explains most of how the tool behaves. codespell does not check for word membership in a complete dictionary, it looks for a set of common misspellings. Therefore it should catch errors like `adn`, but it will not catch `adnasdfasdf`. The same property is stated as a benefit: it should not generate false positives when you use a niche term it does not know about.

That is a real engineering decision rather than a limitation dressed up. A tool that validated every word against a full English dictionary would have to know your domain vocabulary to stay quiet, and any domain vocabulary you have is exactly what a general dictionary lacks. codespell inverts the problem. It knows the perhaps tens of thousands of ways English gets misspelled, and it stays silent about everything else, so a variable named after an obscure protocol or a chemical is never a question.

The scope is described as text files, designed primarily for checking misspelled words in source code, with one specific accommodation: backslash escapes are skipped. That matters more than it sounds, since a C or Python string literal full of escape sequences is a rich source of tokens that look like words and are not.

The dictionaries themselves have a lineage worth knowing. The project ships a collection of dictionaries described as an improved version of the list available on Wikipedia, after applying them in projects like the Linux Kernel, EFL and oFono. You can supply your own dictionary, and patches for new or different entries are explicitly welcome. The files live in `codespell_lib/data/dictionary*.txt`, and you can test a candidate word against the current set by piping it in:

bash
echo "word" | codespell -
echo "1stword,2ndword" | codespell -

That question, whether a word you want to add is already there, comes up constantly when working on the dictionary, which is why the README documents a one-line way to ask it. Optional builtin dictionaries can be selected with `--builtin`, and `--builtin=all` enables all of them, though that only works without custom `-D` dictionaries.

Installing a package with no runtime dependencies

Installation is a single pip command:

bash
pip install codespell

Python 3.9 or above is required. The `pyproject.toml` shows a project with `dependencies = []`, which is unusual and worth pausing on: nothing is installed alongside codespell when you install it. The optional extras exist instead. There is a `hard-encoding-detection` extra pulling in `chardet`, and a `toml` extra pulling in `tomli` for Python versions below 3.11, since `tomllib` is only in the standard library from 3.11 onward.

The package is licensed GPL-2.0-only and authored by Lucas De Marchi. That license is a consideration if you want to run codespell as part of a closed-source build, since it is copyleft rather than permissive. The classifiers list Windows, POSIX, Unix and MacOS, and Python 3.9 through 3.13, which matches the requirement floor.

Running it needs no arguments in the common case. `codespell` with no arguments checks all files in the current directory, and `codespell some_file some_dir/ *.ext` checks specific files or directories named on the command line or matched by glob. The README is clear that `codespell -h` is where the exhaustive list of options lives, so what follows is the subset that matters in practice.

Four ways to say do not report this one

False positives are the central problem for a tool like this, and codespell gives you four distinct mechanisms, which is generous. The first is a file of allowed words, one word per line:

bash
codespell -I FILE, --ignore-words=FILE

The second is a comma-separated list on the command line, which is the quick version for one-off use:

bash
codespell -L word1,word2,word3,word4

The third is an exclude file, which ignores whole lines matching those in the given file exactly. The fourth is inline, in the source itself, which is the one that scales best in a large repository:

python
def wrod() # codespell:ignore wrod
    pass

The inline form takes a list of words, and there is a bare form that ignores every misspelling on the line. There is also a separate directive for the next line, which the README recommends when a formatter pushes comments around:

python
# codespell:ignore-next-line wrod
def wrod():
    pass

The bare version of that ignores all misspellings on the following line. That the project provides two directives rather than one is a small sign that it has been used on real codebases where the comment placement fights you.

There is a subtlety in the ignore lists that will bite someone within a week. Spelling errors are matched case-insensitively, but words to ignore are case-sensitive. The dictionary entry `wrod` also matches the typo `Wrod`, and to silence that you must pass `wrod` in lowercase to match the case of the dictionary entry. That asymmetry is defensible for a tool where the ignore list lives in version control, but it is not obvious and the README does spend a paragraph on it.

Skipping files, directories and quiet output

Skipping is handled by `--skip`, which takes a comma-separated list of files and accepts globs as well. To skip `.eps` and `.txt` files it is `codespell --skip="*.eps,*.txt"`, and to skip directories it is `codespell --skip="./src/3rd-Party,./src/Test"`. There is also `-x` or `--exclude-file` for ignoring whole lines that match a file's contents exactly.

The README then gives a command that is worth copying into any repository you are introducing this to, because it shows what a real invocation looks like:

bash
codespell -d -q 3 --skip="*.po,*.ts,./src/3rdParty,./src/Test"

The description is that this lists all typos found except translation files and some directories, displayed without terminal colors and at a quiet level of 3. The `-d` flag produces diff output, which matters more than it sounds: a diff tells you what codespell intends to change, and reviewing intent is the whole safety story of an automatic fixer.

Writing changes is the `-w` or `--write-changes` flag, and without it the run is a dry run. The README recommends running it with `-i` or `--interactive`, and gives the combination `codespell -i 3 -w` as interactive mode at level 3 with changes written to file. Interactive mode is the difference between a tool you trust and a tool you run once and then disable.

Configuration that lives in the repository

Command line options can also be specified in a config file, which is what turns this from a personal habit into a project setting. On each run codespell checks the current directory for an INI file named `setup.cfg` or `.codespellrc`, or a file given by `--config`, containing an entry named `[codespell]`:

ini
[codespell]
skip = *.po,*.ts,./src/3rdParty,./src/Test
count =
quiet-level = 3

Each command line argument can be specified in the file without the preceding dashes. The format is defined by Python's `configparser`, so comments start with `;` or `#`. The README also notes that codespell checks for a `pyproject.toml` file in the current directory, with the option to point it elsewhere, so a project already standardised on pyproject does not need a second config file.

There is one more suppression mode that is specific to prose rather than code. The `--ignore-sic` option tells codespell to skip a misspelling that is followed by the editorial `[sic]` marker, case-insensitively. Only the single occurrence preceding the marker is ignored, so other misspellings on the same line are still reported. A closing quote may sit between the word and the marker, which is the common case when documenting a corrected typo:

text
correct the "wrod" [sic] typo in a changelog entry

The README draws the distinction from the inline comment clearly: unlike `codespell:ignore`, the marker is part of the prose itself and does not require naming the word in a tooling comment. For a changelog or a specification document that quotes a typo on purpose, that is the difference between editing the file and leaving it alone.

A Makefile that treats the dictionary as a build artifact

The most interesting part of the repository is not in the README. The tree has a `Makefile` that treats the dictionaries as generated output, which tells you the project sees them as data with a pipeline rather than as a pile of hand-maintained lines.

The generated artifact is `codespell_lib/data/dictionary_en_to_en-OX_AUTOGENERATED.txt`, produced from `dictionary_en-GB_to_en-US.txt` by `tools/gen_OX.sh`, and the name says what it is: a British to American spelling mapping generated into the general dictionary so codespell can flag `colour` where the project wants `color`. The default target runs the generation, then `check-dictionaries`, then builds the man page.

The man page is also generated. The rule runs `help2man` over the `codespell --help` output using `codespell.1.include` and strips the usage section afterwards, so the flag reference cannot drift from the actual flags. The `check` target chains the generation, dictionary checks, a distribution check, pytest and ruff, which is the full local gate.

Dictionary hygiene has its own targets. `sort-dictionaries` runs a pre-commit hook called `file-contents-sorter` over all files. `check-dictionaries` greps every dictionary for leading, trailing or blank lines and fails with a message telling you to run `make trim-dictionaries`, and it runs `codespell_lib/tests/test_dictionary.py` through pytest. `trim-dictionaries` fixes exactly those whitespace problems with sed. `DICTIONARIES` globs both `codespell_lib/data/dictionary*.txt` and the wordlists under `codespell_lib/tests/data/`, so test fixtures are held to the same standard as the shipped data.

Around that, the repository carries the usual markers of a Python project with real infrastructure: `.pre-commit-config.yaml` and `.pre-commit-hooks.yaml` so codespell can lint other projects' repositories, a `tox.ini`, `codecov.yml`, a `.coveragerc`, a `.devcontainer/`, and a `.git-blame-ignore-revs` file, which is the tell that the dictionary has had mass reformattings that would otherwise pollute `git blame`. There is a `snap/` directory for Snapcraft packaging, and an `example/` directory with a C file and a dictionary. The release notes show the shape of the work too: v2.4.2 was mostly a chardet 7 compatibility fix and dictionary corrections, and v2.4.3 is mostly individual dictionary additions such as `radback` and `repetirion`, mixed with pre-commit autoupdates from a bot.

Editorial conclusion

codespell earns its place by knowing what it is not trying to do. It holds a curated list of common misspellings rather than a complete dictionary of the language, so it catches `adn` and ignores `adnasdfasdf`, and that design choice is what keeps it from flagging the jargon in your own codebase. The practical surface is small: `pip install codespell`, run it with no arguments in a directory, and add `-w` to write the fixes. The real work is in the exclusion machinery, since a checker you cannot silence is a checker you eventually turn off. That machinery comes in four forms, an ignore list file, a comma-separated list, an exclude file for whole lines, and inline `codespell:ignore` comments, plus a `--ignore-sic` mode for prose where the typo is intentional. Version 2.4.3 shipped on 2026-07-15 and the last push was on 2026-09-28. Start with `codespell -d -q 3 --skip="*.po,*.ts,./src/3rdParty,./src/Test"` to see the noise floor in a repository, then move configuration into `pyproject.toml` or `.codespellrc` so the next run is identical on every machine.

Frequently asked questions

How do I install codespell?

Run `pip install codespell`, which needs Python 3.9 or above. The package has no required runtime dependencies. Encoding detection via chardet and TOML parsing on Python below 3.11 are available as optional extras if you need them.

How do I stop codespell from flagging a word I know is correct?

There are four ways: a file of allowed words passed with `-I`, a comma-separated list with `-L`, an exclude file for whole lines with `-x`, or an inline comment in the source such as `codespell:ignore word`. Note that ignore words are case-sensitive even though the misspellings themselves are matched case-insensitively.

Will codespell rewrite my files when I run it?

Not by default. Without `-w` the run is a dry run, and `-d` shows a diff of what would change. The README recommends combining writing with interactive mode, for example `codespell -i 3 -w`, so you see each change before it is kept.

Where do codespell's misspelling lists come from?

They are shipped in `codespell_lib/data/dictionary*.txt` and are described as an improved version of the public list of common misspellings, refined by applying them in projects such as the Linux Kernel, EFL and oFono. You can add your own with `-D`, and the README documents how to test whether a word is already present.

Official sources

  1. codespell-project/codespell on GitHub
  2. Issues
  3. License: GPL-2.0
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/codespell-project-codespell.svg)](https://hysenlabs.com/projects/codespell-project-codespell)