ScanCode's release tags name a sub-project, and its requirements pin every package exactly
:mag: ScanCode detects licenses, copyrights, dependencies by "scanning code" ... to discover and inventory open source and third-party packages used in your code. Sponsored by NLnet, the Google Summer of Code, Azure credits, nexB and other generous sponsors!
At a glance
- What is it?
- aboutcode-org/scancode-toolkit detects licences, copyrights and dependencies in source and binary files, and the repository is unusually explicit about how it is put together: six separate packaging manifests driven by a build backend of its own, a requirements list of about a hundred entries pinned to exact versions including two packages that are themselves patches of other people's libraries, a container image pinned to one processor architecture whose build trains a Markov chain model for gibberish detection, and four badge definitions that point at the same project under a different organisation. The three newest releases on record are all for one of its internal libraries rather than for the toolkit.
- Who is it for?
- Use ScanCode Toolkit when you need a defensible inventory of what licences are in a codebase, because its central claim is that it compares against licence texts rather than guessing from patterns, and that claim is the reason to choose it over a regex scanner. Three things to check.
- Can I use it commercially?
- Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
- Is it still maintained?
- Yes. The repository last received commits 7 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 5, 2026, and from our analysis. They are not legal advice.
Editorial analysis
Six manifests and a build backend of its own
The repository root holds six files whose names differ only in the middle, and each one builds a different artefact: the toolkit, a reduced toolkit, two data packages for licence text and licence index, one for a shared internal library, and one more. The manifest at the root says in a comment that it exists only for development, names the project with a development suffix, and carries a release candidate version, while the other five are what actually produce the wheels and archives. The build backend is not the usual packaging library either; it is a project of the same organisation with a pinned minimum version, and the comment explains that it is the tool that drives the other manifests. That arrangement is coherent for a project that ships several related distributions from one tree, and it means a contributor who edits the root manifest expecting a release has edited the wrong file.
The release tags name a sub-project, not the toolkit
The three newest releases on record are all named after one of the internal libraries rather than after the toolkit, and they carry that library's own version line. The toolkit's own version lives somewhere else entirely: the development manifest in the root declares a release candidate in the low thirties, while the newest release of anything you can see on the page belongs to the shared library and sits several versions lower. The consequence is that the tag list is not a release history for the thing most people install, and a reader who reads the version off a tag will install a different component than they expected. The tree reinforces the arrangement by carrying separate changelog files for the toolkit and for the internal library, which is the honest way to publish several versions at once, and also means the toolkit's own history has to be read somewhere other than the releases page.
The requirements list pins everything, including two patched packages
The install list is about a hundred lines and every one is pinned to an exact version, with no ranges at all. Among them are several of the project's own sibling libraries, each pinned too, which is how the toolkit keeps a set of related packages in step across a release. Two entries have names that announce what they are: one is a patched distribution of a parameter expansion library and another is a patched distribution of a Maven client. A file at the root configures the tool that vendors and patches dependencies, so this is a maintained practice rather than a workaround someone forgot to remove, and the trade is deliberate: an install that cannot drift is an install you can reproduce for an audit, which is exactly what this project is for. The cost is a long list to review, and two entries a reviewer has to go and look up separately.
The image is one architecture and its build trains a model
The container file pins its base image to a single processor architecture, which means the image it produces will not run everywhere the documentation says the software runs, and the repository carries separate dependency lists for a native install and for Linux to compensate. The more surprising part is what the build does. After installing dependencies it runs three preparation commands: one builds the base licence index, one builds the cache of package patterns, and one trains a Markov chain model used for gibberish detection. So building this image is not a packaging exercise, it is a build with training in it, and it will take correspondingly longer and produce a different result each time unless the training input is pinned. Both facts are stated plainly in the file, which is the right way to write it; they are simply the first two things to know before you put the image into a pipeline.
Four badge definitions point at a different organisation
The page has a four cell table of build badges, and the image substitutions that fill them point at continuous integration definitions under a project owned by a company rather than under the organisation that hosts the repository now. All four do, including the release badge, so the badge a reader sees is reporting on a build that may or may not be the same commit as the one they are reading. Two consequences follow. A green release badge is not evidence about this repository's default branch, which is a development branch that the badge definitions do query, so part of it is right and part of it is not. And anyone debugging a failing build from the badge lands in the wrong place, which is a small thing that costs a maintainer an occasional confused issue report.
The claims on the page are superlatives without numbers
Two paragraphs make strong statements and neither gives a measurement. One calls the project the leading tool in scanning depth and accuracy, used by hundreds of software teams, and another calls its licence detection engine the most accurate available, on the grounds that it compares your code against a database of licence texts rather than relying on patterns, edit distance or machine learning. The second claim is the interesting one, because it is a statement about method rather than about a benchmark, and method claims are checkable: you can read how the comparison works and decide whether it suits your situation. The first claim is not checkable at all. Neither is disqualifying, and the method argument is a good reason to try the tool, but a reader evaluating it for a compliance workflow should note which sentences are evidence and which are adjectives.
Two typos and a compound licence in the same page
The licence section is the part to read carefully, because it is not one licence. The project overall is under a permissive licence, the reference datasets are under a creative commons attribution licence, and third-party components and the test suite's code and data carry a set of secondary licences including copyleft ones. There is a notice file, and the project uses per-directory files that document the origin and licence of vendored code, which is a stronger practice than a single notice. The same page that does that has a misspelling in the build status paragraph and another in the sentence offering commercial support. That contrast is the useful part: the licence handling is careful where it has to be, and the prose is not edited to the same standard. Two more files at the root, a roadmap for the toolkit and one for the wider organisation's work, suggest the same split between careful machinery and lightly edited pages.
Editorial conclusion
Use ScanCode Toolkit when you need a defensible inventory of what licences are in a codebase, because its central claim is that it compares against licence texts rather than guessing from patterns, and that claim is the reason to choose it over a regex scanner. Three things to check. The licence is compound, with a different licence for the reference datasets than for the code, so read the data licence before redistributing a scan's reference material. The install is a configure script plus a fully pinned dependency list, which means reproducible and slow, and two of those pinned packages are patched forks you should look at. And the release tags are per internal library, so pinning the toolkit by its newest tag will not work the way you expect; read the changelog rather than assuming the tag you can see is the toolkit.
Frequently asked questions
What is scancode-toolkit?
A set of code scanning tools that detect the origin of code, meaning its copyrights, its licences and its vulnerabilities, in both source code and binary files. You can use it as a command line tool or as a library.
What licence is ScanCode Toolkit under?
A compound one: a permissive licence for the project overall, a creative commons attribution licence for the reference datasets, and a set of secondary permissive and copyleft licences for third-party components and for the test suite's code and data, documented in a notice file and per-directory files.
Does ScanCode Toolkit run on Windows, macOS and Linux?
Yes, and the project tests on all three. Its container image is built for one processor architecture only, and the repository carries separate dependency lists for a native install and for Linux.
How does ScanCode Toolkit install its dependencies?
Through a configure script at the repository root that creates a virtual environment, with a flag for a development install. The shipped requirements list pins every package to an exact version, including several of the toolkit's own internal libraries and two that are patched distributions.
What output formats does ScanCode Toolkit write?
JSON, YAML and HTML, the two software bill of materials formats which are CycloneDX and SPDX, and your own format if you write it as a template.
Why does building the ScanCode container image train a model?
Because the build runs three preparation commands after installing dependencies: one builds the licence index, one builds the package patterns cache, and one trains a Markov chain model used for gibberish detection.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/aboutcode-org-scancode-toolkit)