Library / SDK
commaai/panda avatar
commaai/panda

commaai/panda: CAN bus firmware where the test suite is itself mutation tested

firmware powering the comma.ai panda

1,701 stars984 forksCMIT

At a glance

What is it?
This is the firmware for a small hardware device that speaks two generations of the controller area network bus and enforces a per-vehicle safety model compiled in from another repository. The most unusual thing in its readme is a single line near the end: the tests that check the safety logic are themselves verified by deliberately breaking them, which is the only way to establish that a test suite means anything.
Who is it for?
panda is worth reading as an example of verification discipline rather than as something to install, because the practices here, a published standards coverage table, strict compiler warnings promoted to errors, per-variant regression tests and mutation testing of the tests themselves, are applicable to any project where a bug is expensive.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 17 days ago.
What is it written in?
Mainly C, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 20, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The tests are tested, which is the rarest claim in the readme

Building and testing is two scripts:

bash
git clone https://github.com/commaai/panda.git
cd panda
./setup.sh
./test.sh

The verification section builds up in layers, and the last layer is the one worth stopping on.

There is static analysis by a C linter, with an addon specifically for checking a widely used automotive coding standard, and a coverage table for that check is published in the repository. Compiler warnings are promoted to errors with a set of flags including a strict-prototypes warning that most projects do not enable, because it catches a class of implicit-declaration bug. The safety logic has unit tests per supported vehicle variant, whose stated purpose is to ensure behaviour does not change. And there is a hardware-in-the-loop test that exercises the device itself across all active hardware variants, covering more than CAN messaging: it checks the safety model again, compiles and flashes both the bootloader stub and the application, and measures loopback and latency through the serial peripheral.

Then the last line. The tests are themselves tested, by a mutation test, applied to two things: the measurement of standards coverage, and the vehicle-specific safety logic.

Mutation testing is a well-known idea and almost nobody does it. You take a test suite, change the code in a way that should break a test, and check that a test actually failed. A test suite with no failing mutation is a test suite that does not test anything, and this is the only way to find that out. Most coverage tools tell you which lines ran. Almost none of them can tell you whether a line running means an assertion was checked.

Applying it here, to the two places where the claim is most load bearing, is the correct priority. The standards coverage measurement is the thing the project publishes as evidence, so if the coverage tool has a gap in how it detects a violation, the published table overstates compliance. The safety logic is the thing that prevents the device from sending unsafe CAN frames, so if a unit test for it does not actually fail when the logic changes, the regression protection is decorative.

The hardware-in-the-loop test deserves a separate note because it is the only one that touches real silicon. Everything else is a build-time check. A device that compiles, passes unit tests and then cannot actually forward a message on a bus is a device that passes CI and fails in a car, and the loopback and latency measurements through the serial peripheral are the kind of check that catches firmware configuration errors which no amount of unit testing would.

A published standards coverage table means the gaps are visible

The readme states that the C linter has a specific addon for checking violations of a widely used automotive coding standard, and links to a coverage table. It is worth being precise about what that table is, because it is the difference between claiming compliance and measuring it.

The standard in question is a set of language-level rules written for code whose failure mode is a malfunction rather than a crash. It covers things a general-purpose linter does not: implicit conversions, recursion, pointer arithmetic, and a long list of requirements about how types may be combined. It is prescriptive to the point of being unusual, because its authors assumed the code would be written by people who did not know C and who would be working on something they could not test.

A linter addon for it will catch a subset. Some rules need semantic information the linter does not have. Some are advisory rather than mandatory and a project may have decided to skip them. Some conflict with a codebase that predates the standard. All of those are legitimate reasons for a rule to be unchecked, and all of them are reasons a coverage number without a table is meaningless.

So the table is the artefact. It exists in the repository, it is a document rather than a log line, and it can be diffed between commits. That means when a new file is added, a reviewer can see whether the newly introduced violations are inside the checked subset, and when the tool is upgraded, a change in the number is visible as a change in a file rather than as a build log that nobody reads.

The same reasoning applies to the strict compiler flags, and it is worth noticing what one of them buys. A strict-prototypes warning catches a function declared without a parameter list, which in C means the compiler assumes unspecified arguments. That single omission turns a type mismatch into a silent memory error, and it is invisible at review time because the code looks correct. Promoting it to an error is a one-flag change that eliminates a class of bug permanently, and it is the cheapest safety measure in the entire repository.

The readme also says the code rigor requirement applies specifically to the application code in the board directory, rather than to the whole project. That scoping is a judgement rather than an omission: the C that talks to the bus is held to a different standard from the tooling around it, which is the correct way round and is worth naming because many projects either apply the strictest standard everywhere and get ignored, or nowhere and get audited.

The safety model is a runtime gate, not a document

The safety section is short and points elsewhere, which is the correct design and also the part a reader has to work hardest to understand.

The device is compiled with vehicle-specific safety logic that comes from a separate repository, one shared with the larger driver-assistance system it works with. The details about the car models, about how the safety logic is tested, and about the rigor of that code all live in that repository rather than here. So the safety model is not a set of comments in this firmware. It is a compiled artefact produced elsewhere and linked in.

That has a concrete consequence visible in the usage example, and it is the most important line in the whole readme for anybody who wants to use the device. You cannot send a message until you have selected a safety mode. The Python interface shows it directly: you import the vehicle parameter struct, call a function that sets the safety mode to a named model, and only then call the function that transmits. There is no overload that sends without a mode.

That ordering is the whole safety story compressed into two lines of example code. The safety model decides which messages are permitted given the car's configuration, and it is enforced by the device rather than by the caller. A caller who sets the wrong mode gets the wrong permission, which is a real risk, but a caller who forgets to set a mode at all gets nothing, which is not. The failure mode is a wrong configuration rather than an absent one, and that is a much better shape for an automotive device.

It also explains why the unit tests are per vehicle variant. The safety model is a different piece of logic compiled into each car's build, so a change in the shared repository can change behaviour for one car and not another, and the only way to catch that is to test each variant. The readme states the purpose of those tests as ensuring behaviour remains unchanged, which is the phrasing of a regression test rather than a functional test, and that is the right kind.

The examples directory reinforces the point. Alongside the obvious logging and bit transition scripts there is a script that queries the vehicle identification number and statistics, and one named for a specific car manufacturer. A VIN read over a bus is exactly the kind of operation that a safety model exists to gate, so its presence as a documented example is a reasonable signal about what the device is used for.

The compiler and the linter arrive as pinned packages

Look at the development dependencies and there is a decision that deserves more attention than it gets.

The compiler for the embedded target is not assumed to be on the machine. It is a Python package, versioned, in the development dependency group, along with the static analysis tool. The build is a Python-centric workflow: a setup script, a test script, and a configuration manifest that declares what the tooling is.

The reason to do this is reproducibility of verification, not of the artefact. If your static analysis tool is whatever version the machine happens to have, then a codebase that passes today can fail tomorrow when somebody's operating system updates the tool, and the failure looks like a code problem. If your compiler is whatever is installed, then a warning that was a warning in one build can be an error in the next, and the build becomes non-deterministic in a way nobody can debug.

Pinning both as packages means the analysis result is a function of the lock file rather than of the environment. Two developers on two machines get the same verdict, and a year from now the project can still reproduce the check that passed. For a project whose entire claim is code rigor, that is a foundational decision and it appears in the manifest rather than in a wiki page, which is where it belongs.

The build system is a Python-adjacent one rather than a shell script or a makefile, and the repository has the configuration files a cross-compilation build needs. That is consistent with the same approach: the build is a program with declared inputs rather than a sequence of shell commands someone typed.

The one oddity is the ruff configuration, which sets a line length of one hundred and sixty characters. That is very long, and it is a deliberate choice for a codebase where a long line is often a table of register names or a bit mask. It is also the kind of choice that a reviewer would question once and then stop questioning, which is arguably the correct outcome for a linter setting that has no effect on correctness.

An unpinned dependency from a branch, in a safety-critical project

The manifest declares a Python floor and a ceiling, and the ceiling has a comment attached explaining it. macOS does not work with the newer Python release because of a data serialisation library that arrives transitively from the shared repository. So the ceiling is not a policy decision, it is a workaround for somebody else's package on one operating system, and it is documented in place.

That is good practice. Documenting why a constraint exists is the difference between a constraint somebody will respect and one somebody will delete in a hurry because it looks arbitrary.

The same file contains a decision that is harder to defend and more interesting. One of the three runtime dependencies is resolved from a git URL pointing at a branch, not at a tag or a version. The branch is the default one. So every install of this library, on any machine, resolves to whatever the current tip of another repository is at that moment.

For an ordinary library that would be an obvious problem. For this one, the tension is real and worth stating properly. That other repository is where the per-vehicle safety models live, and the models change as new car models are supported and as bugs in existing ones are found. Pinning it would mean a device firmware build could carry a safety model that is missing a fix. Not pinning it means the build is not reproducible, and a codebase that publishes a standards coverage table and mutation tests has rather explicitly staked its claim on reproducibility.

There is a reasonable position on the other side, and this project has clearly taken it deliberately. In safety work, being current on the model matters more than being reproducible on the model, because the failure you are protecting against is a known defect that has already been fixed upstream. A pinned build is a build with a known-stale safety model.

The linter configuration, meanwhile, sets a language target one version below the declared Python floor. So the linter will accept syntax the runtime will refuse to run. That is harmless in practice because the ceiling and floor are enforced by the installer, and the linter target is about which syntax the linter's own parser understands rather than about what the project supports. It is still a small inconsistency in a manifest that is otherwise unusually well documented, and it is the kind of thing that makes a reader check the rest of the file rather than trust it.

World-writable access to a vehicle bus, and five hardware IDs for one device

The Linux setup instructions include a device rules block, and it is worth reading carefully because it says two things about the product at once.

The first is the mode. Every rule sets the device to world-readable and world-writable. That is the standard way to let a non-root user talk to a USB device, and it is the correct choice for a single-user machine, where requiring root to send CAN frames would be unusable. It is also, on a shared machine, a grant of the ability to transmit arbitrary frames on a vehicle bus to every local account, which is a much larger capability than it looks. The safety model limits what the device will send, so the exposure is bounded by that model, but the model is only as good as the mode it was selected in.

The second is the number of rules. There are five, covering two vendor identifiers and two product identifiers. One of the vendor identifiers belongs to the microcontroller manufacturer, which makes sense for a board sold under the chip vendor's USB identity. The other is the project's own, which appears with both product identifiers. So the same logical device has at least two hardware revisions in the field, and each revision has to be listed for the rules to match.

That is a small, concrete illustration of what supporting hardware means. The firmware has to work across those revisions, the hardware-in-the-loop test has to run on each active variant, and the installation instructions have to name every combination. None of it appears in a feature list and all of it is real work.

There is one more detail in that block worth mentioning, because it is the kind of thing that gets missed. After writing the rules, the instructions reload and retrigger them. Skipping that step is the most common reason a freshly installed device is not found, and the fact that it is written down saves the twenty minutes of debugging that would otherwise follow.

The Python package naming follows the older convention, where the distribution name and the import name differ. The distribution carries a different name entirely, with a version that exists only in the manifest, because there are no release tags. Which raises the question of how you know what firmware is on a device, and the examples directory answers it: there is a script whose whole purpose is to query firmware versions. For firmware, that is the right design. You do not release firmware, you build it and flash it, and the version that matters is the one on the device, so you ask the device.

Editorial conclusion

panda is worth reading as an example of verification discipline rather than as something to install, because the practices here, a published standards coverage table, strict compiler warnings promoted to errors, per-variant regression tests and mutation testing of the tests themselves, are applicable to any project where a bug is expensive. It is not something most people should run, since the device is a node on a vehicle bus and installing the user-space library requires granting world-writable access to the USB device. If you are evaluating the safety model itself, read it in the separate repository it is compiled from, and treat the coverage table as the first place to look for what is not checked rather than as reassurance.

Frequently asked questions

What is the commaai panda device?

It is a small hardware device that speaks both generations of the controller area network bus used in vehicles, running firmware on a 32-bit microcontroller from a major semiconductor manufacturer. The repository holds that C firmware plus a Python user-space library, tests, build scripts and example scripts for using the device in a car.

What does mutation testing of the panda test suite actually mean?

It means deliberately breaking the code in ways that should fail a test, and checking that a test actually fails. The repository applies this to the two places where its claims are load bearing: the measurement of coding-standard coverage, and the vehicle-specific safety logic. A suite where no mutation causes a failure is a suite that does not test anything, and this is the only way to find that out.

How does panda's safety model work?

The device is compiled with per-vehicle safety logic supplied by a separate repository, and the logic is enforced by the device rather than documented for the caller. The Python interface makes it a gate: you must select a named safety model before any message can be transmitted, and there is no call that sends without one.

What verification does panda run on real hardware?

A hardware-in-the-loop test covering all active device variants. Beyond the static analysis and unit tests, it performs additional safety model checks, compiles and flashes both the bootloader stub and the application, exercises receiving, sending and forwarding messages on all buses, and measures loopback and latency through the serial peripheral.

Why does panda's Python package have an upper version bound?

The manifest pins a narrow range and the comment on the ceiling explains it: the newer Python release does not work on macOS because of a data serialisation library that arrives transitively from the shared repository. The bound is a workaround for an upstream package on one platform, documented in place rather than left as an unexplained restriction.

Official sources

  1. commaai/panda on GitHub
  2. Issues
  3. License: MIT
  4. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/commaai-panda.svg)](https://hysenlabs.com/projects/commaai-panda)