Library / SDK
sepandhaghighi/pycm avatar
sepandhaghighi/pycm

PyCM reads vectors or a finished matrix, and its four install paths do not ship the same code

Multi-class confusion matrix library in Python

1,508 stars126 forksPythonMIT

At a glance

What is it?
sepandhaghighi/pycm turns predictions into per class counts, normalized tables and error indices, but the declared dependencies, the optional plotting backends and the version each install path resolves to do not line up.
Who is it for?
PyCM earns its place when you already hold predictions as label vectors or as a finished matrix and want per class counts, an overall statistics block and error indices in one object rather than three separate reports. Before you commit, check two things.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 4 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 8, 2026, and from our analysis. They are not legal advice.

Editorial analysis

Two constructors, and two different ideas about what a class is

PyCM takes input in two shapes, and the shape decides what a class is called. Hand it two label vectors and it works out the class list itself, so integer labels come back sorted as `[0, 1, 2]`. Hand it a nested dict and the keys you typed become the class list untouched, which is how you keep names like `Class1`, or your own domain labels, instead of losing them to a sort.

pycon
>>> from pycm import *
>>> y_actu = [2, 0, 2, 2, 0, 1, 1, 2, 2, 0, 1, 2]
>>> y_pred = [0, 0, 2, 1, 0, 2, 1, 0, 2, 0, 2, 2]
>>> cm = ConfusionMatrix(actual_vector=y_actu, predict_vector=y_pred)
>>> cm.classes
[0, 1, 2]
>>> cm.table
{0: {0: 3, 1: 0, 2: 0}, 1: {0: 0, 1: 1, 2: 2}, 2: {0: 2, 1: 1, 2: 3}}
>>> cm.print_matrix()
Predict 0       1       2
Actual
0       3       0       0

1       0       1       2

2       2       1       3

>>> cm.print_normalized_matrix()
Predict       0             1             2
Actual
0             1.0           0.0

The counts live in `cm.table` either way, one dict per actual class holding one dict per predicted class. Both shapes print the same two tables: raw counts, then row normalized values where each actual class divides by its own total. That normalization is worth checking once against your own data, because a class with few observations produces large normalized fractions from very few events. The vector example on the README page ends one row into that normalized printout, so the remaining rows are not shown there.

transpose is a direct CM only option

The dict form is not just a shortcut for people who already hold counts. It is the only form with the extra input parameters attached, and `transpose`, added in version 1.2, is documented as working only in Direct CM mode. An object built from vectors therefore cannot be flipped after the fact and has to be constructed again from the dict form to get the other orientation. If you wrap PyCM in a pipeline that sometimes receives vectors and sometimes receives matrices, that asymmetry is the detail to design around.

pycon
>>> from pycm import *
>>> cm2 = ConfusionMatrix(matrix={"Class1": {"Class1": 1, "Class2": 2}, "Class2": {"Class1": 0, "Class2": 5}})
>>> cm2
pycm.ConfusionMatrix(classes: ['Class1', 'Class2'])
>>> cm2.classes
['Class1', 'Class2']
>>> cm2.print_matrix()
Predict      Class1       Class2
Actual
Class1       1            2

Class2       0            5

>>> cm2.print_normalized_matrix()
Predict       Class1        Class2
Actual
Class1        0.33333       0.66667

Class2        0.0           1.0

>>> cm2.stat(summary=True)
Overall Statistics

The same constructor argument set carries `sample_weight`, added in version 1.2, `relabel` from version 1.5, `file` from version 0.9.5 for loading a `.obj` written by the `save_obj` method, and `threshold` from version 0.9 for real value predictions. Overall numbers are reached through `cm.stat(summary=True)`, which prints under the heading Overall Statistics.

relabel rewrites class names, and one_vs_all then reads them

Class names are editable after construction. `relabel` arrived in version 1.5 and takes a mapping from old names to new ones, so a matrix built from integers 0, 1 and 2 reports classes `L1`, `L2`, `L3` afterwards without the counts moving underneath. That matters more than it first appears, because the one versus all view in `to_array` is addressed by class name, and `L1` is a label that exists only because the relabel step ran first.

pycon
>>> cm.relabel(mapping={0: "L1", 1: "L2", 2: "L3"})
>>> cm
pycm.ConfusionMatrix(classes: ['L1', 'L2', 'L3'])
>>> cm.to_array()
array([[3, 0, 0],
       [0, 1, 2],
       [2, 1, 3]])
>>> cm.to_array(normalized=True)
array([[1.     , 0.     , 0.     ],
       [0.     , 0.33333, 0.66667],
       [0.33333, 0.16667, 0.5    ]])
>>> cm.to_array(normalized=True, one_vs_all=True, class_name="L1")
array([[1.     , 0.     ],
       [0.22222, 0.77778]])

The second row of that two by two view pools every other class into a single row: nine of the twelve sample observations are not class 0, and two of those nine were predicted 0, which produces the 0.22222 and 0.77778 pair. Each row is normalized on its own.

position hands back a fourth list, TN, for every class

Added in version 2.8, `position` goes back to the prediction vector and reports the index of every observation that landed in each of the four outcomes, per class.

pycon
>>> cm.position()
{0: {'FN': [], 'FP': [0, 7], 'TP': [1, 4, 9], 'TN': [2, 3, 5, 6, 8, 10, 11]}, 1: {'FN': [5, 10], 'FP': [3], 'TP': [6], 'TN': [0, 1, 2, 4, 7, 8, 9, 11]}, 2: {'FN': [0, 3, 7], 'FP': [5, 10], 'TP': [2, 8, 11], 'TN': [1, 4, 6, 9]}}

Class 0 lists true positives at indices 1, 4 and 9, false positives at 0 and 7, no false negatives at all, and seven true negatives at the remaining positions. The true negative list is what separates this from the usual per class report, because in a multiclass setting every observation is a true negative for each class it is not, so each class carries its own list and the lists do not have the same length. For the same run class 2 lists three true positives against three false negatives and two false positives. If you are auditing a classifier rather than scoring it, these index lists point straight at the rows of your own input vector behind a given error.

combine adds two matrices cell by cell

Version 3.0 added `combine`, framed on the project page as useful for mini batch learning: evaluate each batch on its own and merge the resulting matrices afterwards instead of holding every observation at once. The merge is plain addition.

pycon
>>> cm_combined = cm2.combine(cm3)
>>> cm_combined.print_matrix()
Predict      Class1       Class2
Actual
Class1       2            4

Class2       0            10

In that printed example the first matrix holds one and two across its top row, then zero and five, while the combined table holds two and four, then zero and ten. Every cell grew by exactly what the second matrix contributed. The example passes two matrices that already share the labels `Class1` and `Class2`, and the shown interface carries no label mapping of its own, so the class sets have to agree before you call it. The result is a separate object rather than a change to the first matrix, which is what makes it reasonable to keep the per batch copies around.

Four install paths, and the code they hand you is not the same

Four routes are offered and only one of them pins a version. The PyPI line asks for `pip install pycm==4.6`, which matches the version string in `setup.py` and the newest tag, v4.6, dated 2026-03-09. The source route offers two archives, the v4.6 zip and a dev branch zip, then runs `pip install .`, so choosing the second one installs branch code rather than the tagged release. Commits have landed on the default branch since that tag, the most recent dated 2026-09-18, and the repository is not archived, so the two archives are not interchangeable.

bash
pip install pycm==4.6
pip install .
conda install -c sepandhaghighi pycm
pip install pycm

The conda line resolves against the maintainer's own channel and names no version at all. The MATLAB route installs with a bare `pip install pycm`, also unpinned. Two of the four therefore hand you whatever the index or the channel holds at that moment, which matters most when your evaluation numbers have to be reproduced later.

requirements.txt names two packages, and setup.py splits one file to find them

The runtime dependency list is two lines long, and that is the whole file.

text
art>=1.8
numpy>=1.9.0

Plotting is separate. Version 3.0 added a `plot` method that draws through Matplotlib or Seaborn, and the install notes put the floors at Matplotlib 3.0.0 or Seaborn 0.9.1. Neither package appears in that file, so a clean install can build every table and every statistic and still have nothing to draw with until you add a backend yourself. `setup.py` declares no dependencies inline: it opens the same file, splits the contents on whitespace, drops empty strings and passes the result to `setup()`, so those two floors are decided by the text file and nothing else. The split is on any whitespace rather than on line boundaries, so a comment written into it would reach the installer as a requirement name; the file carries no comments today, so this is a property of the mechanism rather than a fault in it.

The same module reads `README.md` and appends `CHANGELOG.md` into the long description, with a hardcoded paragraph about multi class confusion matrices as the fallback if either read fails. Installed packages are limited to `packages=['pycm']`, so the `docker/`, `MATLAB/`, `Test/` and `paper/` directories visible in the tree stay outside the installed package.

The Python floors step down from 4.3 to 2.4

Support for old interpreters is recorded as a list of last versions rather than enforced by a runtime check. Version 4.3 is the last release to run on Python 3.6, version 3.9 is the last for Python 3.5, and version 2.4 is the last for Python 2.7 and Python 3.4 together. Anything newer than those lines needs a current interpreter.

matlab
>> pyversion PYTHON_EXECUTABLE_FULL_PATH

The MATLAB route states its own floors separately: MATLAB 8.5 or newer, and Python 3.7 or newer installed with both the `Add to PATH` and `Install pip` options ticked, after which the interpreter is named explicitly before any call goes through. That route exists because the library is also driven from MATLAB, which is why the tree carries a `MATLAB/` directory of examples and why the setup module declares a single `pycm` package rather than a tree of subpackages. The two routes together draw a plain line: a current interpreter for the Python side, MATLAB 8.5 with Python 3.7 for the MATLAB side, and a documented older floor for each past release behind those.

Editorial conclusion

PyCM earns its place when you already hold predictions as label vectors or as a finished matrix and want per class counts, an overall statistics block and error indices in one object rather than three separate reports. Before you commit, check two things. The four install paths do not all give you the same code, because only one of them pins a version and the dev branch archive sits ahead of the newest tag. And plotting needs a backend you install yourself, since requirements.txt names only art and numpy. Pin what you test against, and relabel before you ask for one_vs_all numbers.

Frequently asked questions

What does PyCM need installed before it can compute a matrix?

Two packages, art 1.8 or newer and numpy 1.9.0 or newer, listed in requirements.txt. setup.py reads that file, splits it on whitespace and hands the result to setup(), so nothing else is pulled in as a runtime dependency.

Can PyCM draw a matrix, and do I need an extra package for that?

Yes. The plot method arrived in version 3.0 and draws through Matplotlib or Seaborn, with floors of Matplotlib 3.0.0 and Seaborn 0.9.1. Neither backend is listed in requirements.txt, so you install one yourself.

Which Python versions can still run PyCM?

Version 4.3 is the last to support Python 3.6, version 3.9 is the last for Python 3.5, and version 2.4 is the last for Python 2.7 and Python 3.4. The MATLAB route asks for Python 3.7 or newer alongside MATLAB 8.5.

How do I get the indices of the errors PyCM found?

Call position, added in version 2.8. It returns, for each class, the indices in the prediction vector that produced the TP, TN, FP and FN lists, so class 0 in the twelve observation example reports true positives at 1, 4 and 9 and false positives at 0 and 7.

Does PyCM let me merge the results of separate mini batches?

The combine method, added in version 3.0, merges two confusion matrices by adding them cell by cell. In the printed example every cell of the result is larger than the first matrix by exactly the second matrix's contribution.

Official sources

  1. License: MIT
  2. Project website
  3. README
  4. Releases
  5. sepandhaghighi/pycm on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/sepandhaghighi-pycm.svg)](https://hysenlabs.com/projects/sepandhaghighi-pycm)