AlDanial/cloc: counting code, comments and blanks across 200+ languages
cloc counts blank lines, comment lines, and physical lines of source code in many programming languages.
At a glance
- What is it?
- cloc is a Perl command line tool that reports blank, comment and physical source lines per language. It reads files, directories, archives and git commits, and it is best treated as a fast inventory tool rather than a precise parser.
- Who is it for?
- Adopt cloc if you need a quick per-language line inventory of a directory, an archive or a git commit, and you can accept that the numbers come from language-specific regex rules rather than a full parser. Do not adopt it if your question is about logical statements, complexity or test coverage, or if you need per-author attribution, which the README does not document.
- Can I use it commercially?
- Yes, with conditions. GPL-2.0 is a copyleft licence: if you distribute software that includes it, you must release that software's source code under the same licence. Running it internally without distributing it does not trigger that obligation.
- Is it still maintained?
- Yes. The repository last received commits 10 days ago.
- What is it written in?
- Mainly Perl, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
DEEP OPEN-SOURCE ANALYSIS
What cloc actually answers, and who asks that question
The README states the scope in one line: cloc counts blank lines, comment lines, and physical lines of source code in many programming languages. That is narrower than it sounds. It is a physical line counter, not a statement counter, so a single logical expression spread over five lines counts as five. The output table separates those three buckets per language and adds a SUM row, which is the number most people quote afterwards.
The people who get value from this are the ones who need a defensible inventory rather than a metric. A due diligence reviewer handed a tarball wants to know how much of it is Perl versus how much is vendored JavaScript. A maintainer taking over an unfamiliar repository wants the same thing. cloc reads archives directly, so the tarball does not need to be unpacked first. The README also shows it accepting a git commit hash, which means the inventory can be pinned to a revision instead of to whatever happens to be in the working tree.
What it is not is a code quality tool. Nothing in the output says anything about whether the code is good, tested or maintained. The comment-to-code ratio is the closest thing to a judgement call, and that ratio is only as meaningful as the language's commenting conventions allow.
How the counting works: regex rules per language, not a parser
cloc ships as a single Perl script. The repository root holds the `cloc` file itself, alongside `Unix/`, `tests/`, `sqlite_formatter` and two Dockerfiles. There is no build step for the source version beyond having a Perl interpreter available, and the README notes that the Windows executable has no requirements at all.
Internally the tool carries per-language definitions for how to strip comments and how to recognise blank lines. That is why the language list is long and why adding a language is a matter of describing its comment syntax. It is also why the numbers are approximations when a language has awkward constructs: heredocs, embedded SQL strings, template languages layered inside HTML, or a comment marker that appears inside a string literal. cloc does not build an abstract syntax tree, so it cannot know that a `#` is inside a quoted string unless the rules for that language account for it.
The practical consequence is that cloc is fast and dependency-light in exchange for being approximate on edge cases. For a rough size comparison between two trees, the approximation is fine. For a claim that a specific file has exactly N comment lines, it is not the right instrument.
Archive handling and git handling are separate code paths. The README demonstrates both: `cloc master.zip` pulls the zip from GitHub and counts the Perl inside, and `cloc 6be804e07a5db` inside a cloned repository counts the tree at that commit. The `--vcs git` flag appears in the README's shell loop example, where it is used to count each project directory under a parent.
Installing cloc: release zip, package manager, or Docker
The README names three routes. Install from GitHub Releases, install via a package manager, or run the Docker image. The source version needs a Perl interpreter. The Windows executable does not. The Docker version needs Docker.
The executable name depends on which route you took. A development checkout gives you `cloc`. A released source file is named with its version, for example `cloc-2.10.pl`. The Windows build is `cloc-2.10.exe`. The README says `cloc` is used on the page as the generic term for any of them, so a command shown as `cloc` may need renaming on your machine.
The Dockerfile in the repository builds in stages from `perl:slim`. A builder stage installs `gcc` and pulls four CPAN modules: `Algorithm::Diff`, `Digest::MD5`, `Parallel::ForkManager` and `Regexp::Common`. The final stage installs `git` and `unzip`, copies the prepared Perl library tree and source, sets the working directory to `/tmp`, and sets the entry point to `/usr/src/cloc` with a default command of `--help`. That default is why a bare `docker run` on the image prints usage rather than doing anything.
A first real use is counting a single file, which the README shows with `hello.c`. The output reports how many text files were seen, how many were unique, how many were ignored, then a table with columns for files, blank, comment and code.
A first run and reading the three numbers that matter
The README's quick start is three steps: install, open a terminal, invoke cloc on files, directories, archives or git commits. The smallest example is a single C file.
cloc hello.cThe header lines report file counts before the table. In the README's example that is one text file, one unique file, zero files ignored. Then the table gives one row for C with 0 blank, 7 comment and 5 code lines. Those three columns are the whole product.
The ignored count is the field most people skip and later regret skipping. It is the tool telling you what it declined to examine. If you point cloc at a repository root and the ignored count is large, the SUM row describes a subset of the tree, not the tree.
Counting a directory follows the same shape, and the README's example on `gcc-5.2.0/gcc/c` shows why the language column matters: it separates C from C/C++ Header rather than folding them together, so the header-to-source ratio is visible in the output.
For a git repository, the README clones a project and then passes a commit hash as the argument. The header line in that example reports 48 text files, 41 unique files and 8 files ignored, and the table breaks the result across Python, reStructuredText, YAML, Bourne Shell, Text and make. That is the useful shape for reviewing a pinned revision.
git clone https://github.com/inducer/pudb.git
cd pudb
cloc 6be804e07a5dbTo walk several project directories in one pass, the README uses a shell loop with `--vcs git`, changing into each subdirectory and echoing its name before counting.
Where cloc gives you the wrong answer
The first limitation is the one already described: comment detection is rule-based, so any construct that hides a comment marker inside a string, or a string inside a comment, can be misclassified. The README does not claim parser-grade accuracy, and the per-language rule approach is the reason.
The second is the ignore list. cloc decides what to skip, and the README's own examples show non-zero ignored counts on ordinary inputs. If your tree contains generated files, minified bundles or checked-in dependencies, the code column will reflect whatever cloc chose to read. The number is not wrong so much as undefined until you look at the ignored count.
The third is scope. cloc counts physical lines. If you need logical statements, cyclomatic complexity, or the ratio of tested to untested code, cloc cannot produce those and does not pretend to. Teams that want a single quality number should not reach for this tool, because every number it produces is a size measurement.
The fourth is attribution. The README documents counting a commit, not counting who wrote what. If your question is per-author contribution, cloc's documented interface does not answer it. The `--vcs git` flag is shown in the context of counting a working tree, not of producing a per-contributor breakdown.
cloc against scc, and why the difference is not cosmetic
The search data around this project includes a question about scc, which stands for SLOC, Cloc and Code. The comparison is worth making concrete rather than gestural. cloc is a Perl script with per-language regex rules and no build step for the source version. It reads files, directories, archives and git commit hashes, and the Docker image is built from `perl:slim` with four CPAN dependencies.
scc is a different implementation of the same broad idea. The README here does not discuss it, so the honest statement is that cloc's documentation does not position itself against any alternative. What the repository does tell you is what you are choosing: a single Perl file you can read and modify, with language definitions you can extend, and an install path that includes a Windows executable requiring nothing else.
That matters for two kinds of user. If you are on a machine where you cannot install a compiled binary, cloc's source form is a Perl script and the Windows form is a standalone executable. If you need to add a language that cloc does not know, the rule-based design is an invitation rather than an obstacle. If instead you want a single static binary with no interpreter anywhere in the chain, that is a different trade and cloc is not it.
Maintenance, licence and the cost of an upgrade
The repository is not archived and the last push was on 2026-09-20. The most recent release listed is v2.10 from 2026-07-04, preceded by v2.08 on 2026-01-25 and v2.06 on 2025-06-25. That is a release cadence measured in months, not weeks, which is consistent with a mature tool whose language rules change slowly.
Upgrade cost is low for the source route because there is nothing to compile. You replace the script. The version is visible in the output header, which the README's examples show as a URL followed by a version string, so you can confirm which build produced a table by reading the first line. That is a small but genuinely useful property when a number is being cited months later.
The Docker route carries more upgrade surface. The image installs `git` and `unzip` in the final stage and copies a Perl library tree built from four CPAN modules. Rebuilding picks up whatever those modules resolve to at build time, so pinning the image tag is the difference between a reproducible count and a moving one. The Dockerfile also includes a test stage that clones `cloc_submodule_test` and runs `make test` from `Unix/`, which tells you the project tests the git integration path rather than only the file-reading path.
The licence is GPL-2.0. That is a copyleft licence, and the practical question for most readers is whether they are distributing cloc or a modified version of it, not whether they are running it internally. Running a GPL tool over your own source does not change your source's licence, but if you embed cloc in a product you ship, the terms apply to what you ship. Read the LICENSE file in the repository root rather than taking this paragraph as a legal position.
Editorial conclusion
Adopt cloc if you need a quick per-language line inventory of a directory, an archive or a git commit, and you can accept that the numbers come from language-specific regex rules rather than a full parser. Do not adopt it if your question is about logical statements, complexity or test coverage, or if you need per-author attribution, which the README does not document. Before relying on a number, verify that the language you care about is listed in the supported set, and check the files-ignored count in the header, since that is where vendored or generated trees show up. Start with `cloc --vcs git` on the repository you already have checked out, then compare the SUM row against the file count you expected.
Frequently asked questions
What is cloc used for?
It counts blank lines, comment lines and physical lines of source code across many programming languages. It accepts files, directories, archives and git commits as input.
How do I use cloc to count lines of code?
Install it from GitHub Releases, a package manager, or the Docker image, then pass a path or an archive as the argument, for example `cloc hello.c` or `cloc master.zip`. The output is a table with files, blank, comment and code columns per language plus a SUM row.
What is the difference between cloc and scc?
The cloc README does not discuss scc, so no direct comparison is documented here. What the repository does show is that cloc is a single Perl script with per-language regex rules, and that its install routes include a Windows executable with no requirements and a Docker image built from `perl:slim`.
What does LOC mean in code?
In cloc's output the code column counts physical lines of source, separate from blank lines and comment lines. It is a physical line measure, so one logical statement spanning several lines counts as several.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/aldanial-cloc)
Community notes