github-linguist/linguist: what it decides about your repository's language breakdown
Language Savant. If your repository's language is being reported incorrectly, send us a pull request!
At a glance
- What is it?
- Linguist is the Ruby library GitHub uses to label blob languages, skip vendored and generated files, and build language graphs. It is easy to install and easy to misread, because most of its behaviour lives in its heuristics rather than in its output.
- Who is it for?
- Adopt github-linguist if you need the same language attribution GitHub shows, or if you are debugging a repository whose language bar looks wrong. Do not adopt it as a general-purpose syntax highlighter or as a classifier for arbitrary text; it works on Git blobs, and its answers come from filename rules, extensions and heuristics rather than from parsing.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository received new commits within the last day.
- What is it written in?
- Mainly Ruby, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
DEEP OPEN-SOURCE ANALYSIS
The problem Linguist solves on GitHub, and who actually needs it
GitHub has to answer a narrow question for every repository it hosts: which language is this file, and should it count toward the language bar at the top of the page? The README states the library is used on GitHub.com to detect blob languages, ignore binary or vendored files, suppress generated files in diffs, and generate language breakdown graphs. That is the whole scope. Linguist is not a parser and not a highlighter; it is a classifier plus a set of exclusion rules, and the two are tangled together.
The people who need it are narrower than the download numbers suggest. Maintainers whose repository shows the wrong primary language are the obvious audience, and the repository description invites exactly that: if your repository's language is being reported incorrectly, send a pull request. Beyond them, anyone building a tool that wants to reproduce GitHub's numbers, or CI that fails when a vendored directory starts dominating the language graph, has a reason to depend on the gem rather than reimplement the rules. If you only want to colour a code block, this is the wrong dependency.
How detection works: strategies, vendored paths and generated files
Linguist walks the Git objects in a revision rather than the working tree. The Ruby example in the README constructs a Rugged repository, takes repo.head.target_id, and passes both into Linguist::Repository. The result exposes project.language for the single dominant language and project.languages as a hash of language name to byte size. That byte count, not a line count, is what drives the percentages.
Each blob gets a strategy, and the --strategies flag prints it. In the README's own output on this repository, Gemfile is tagged [Filename] while bin/git-linguist is tagged [Extension]. So a file with no useful extension can still be classified by name alone. The interesting case is .gitattributes. When an override applies, the strategy line records both the original detection method and whether the attribute changed or merely confirmed it: demo.ts shows [Heuristics (overridden by .gitattributes)] and demo.js shows [Extension (confirmed by .gitattributes)]. That distinction matters when you are trying to work out whether your own override is doing anything.
The suppression side is the part people underestimate. Binary and vendored files are excluded before percentages are computed, and generated files are suppressed in diffs. The README does not enumerate those paths or patterns, so if a directory of yours is silently ignored you will not find the reason in the README; you have to read the repository's own data files under lib/ and the docs directory. Treat the exclusion lists as the real configuration surface of this project.
Installing github-linguist and running it on a real repository
The README gives one install command, and it is a Ruby gem. You need a recent Ruby; the README warns that the macOS or Xcode-supplied Ruby causes problems installing some dependencies and recommends Homebrew, rbenv, rvm, ruby-build or asdf instead.
gem install github-linguistTwo native dependencies carry their own system requirements. charlock_holmes handles character encoding and needs cmake, pkg-config, ICU and zlib. rugged provides the libgit2 bindings and needs libcurl and OpenSSL. On macOS with Homebrew the README suggests installing the build tools first:
brew install cmake pkg-config icu4cOn Ubuntu the README lists the full set in one line, including the Ruby development headers:
sudo apt-get install build-essential cmake pkg-config libicu-dev zlib1g-dev libcurl4-openssl-dev libssl-dev ruby-devOnce the executable is on your PATH, change into a repository and run it with no arguments. The README says this prints the language breakdown by percentage and file size:
cd /path-to-repository
github-linguistYou should see one line per language, with a percentage, a byte count and the language name. To find out which files produced those numbers, add the breakdown flag, and add strategies when a classification looks wrong:
github-linguist --breakdown --strategiesThe README notes that --strategies implies --breakdown unless --json is also given, and that --strategies has no effect when --json is present. For machine consumption, --json emits a single object keyed by language, and combining it with --breakdown adds a files array per language:
github-linguist --breakdown --jsonTo analyse a different revision, such as a tag or a branch, pass --rev with any gitrevisions-compatible value. The README's example runs the tool against Jekyll's gh-pages branch and reports 100.00% HTML, which is a good illustration of how completely a branch can change the answer:
github-linguist jekyll --rev origin/gh-pagesWhere Linguist gets it wrong, and when it is the wrong tool
The first limitation is the one the project itself advertises: detection is not always right, which is why the repository description asks for pull requests when a language is reported incorrectly. If your repository's language is wrong, you are expected to fix the data upstream rather than configure your way out of it locally.
The second is the byte-based denominator. Because percentages come from blob sizes, a single large generated file that survives the exclusion rules can outweigh hundreds of small source files. The --breakdown output is the only way to see this, and the README does not describe a threshold or a warning for it.
The third is the interaction between overrides and strategies. The README shows that an override can confirm rather than change a result, and that --strategies is silently ignored under --json. If your pipeline parses JSON, you lose the provenance information entirely, which is exactly the information you need when the numbers move.
Finally, scope. Linguist analyses Git repositories and single files. It is not a general text classifier, it does not parse your code, and it does not tell you anything about code quality, dependency weight or how much of a language is actually written by humans rather than emitted by a generator. Using it to gate a release on "our repo is 80% Go" is a misuse; using it to check that a vendored directory has not crept into the language bar is not.
Linguist against a syntax highlighter, and against writing your own rules
The closest thing to a real alternative is a syntax highlighting library, for example Rouge or Pygments. The difference in approach is fundamental. A highlighter takes text and returns tokens; it has to handle anything you hand it and it does not care where the file came from. Linguist takes Git blobs, decides a single language label per file, and then decides whether that file should count at all. Highlighting and classification overlap only in the extension-to-language mapping.
That overlap is why the topics list on the repository includes both language-statistics and syntax-highlighting, but the two jobs are not interchangeable. If you want coloured output, a highlighter is the right dependency and Linguist will not help you. If you want to know why GitHub says your repository is 60% C, a highlighter will not tell you, because the answer depends on vendored-file exclusion and byte weighting that only Linguist implements.
The other alternative is maintaining your own extension map and ignore list. That is viable for a single repository with a stable stack, and it avoids the native dependency chain of charlock_holmes and rugged. It stops being viable the moment you want your numbers to match GitHub's, because matching means tracking the same data files and the same heuristics, which is what the gem exists to package.
Maintenance, releases and what the MIT licence does not cover
The repository is not archived, and the last push was on 2026-09-21. Releases are frequent and versioned: v9.5.0 on 2026-03-18, v9.6.0 on 2026-06-08 and v9.7.0 on 2026-08-26. If you pin the gem, expect to move the pin several times a year, because language definitions and heuristics change between minor versions and those changes alter your percentages without any change to your code.
The repository's own Dockerfile shows how the maintainers handle that pinning. It takes LINGUIST_VERSION as an ARG and installs ${LINGUIST_VERSION:+-v ${LINGUIST_VERSION#v}}, with a comment explaining that a bare gem install would leave the RUN instruction text unchanged and let the registry build cache replay the same layer on every release. The #v strips the leading v from the tag for RubyGems. If you build your own image, that comment is the reason to pass an explicit version rather than rely on the default.
The gem is MIT licensed, and the repository ships a LICENSE file at the top level. The native dependencies are separate projects with their own licences, and the README does not discuss them; ICU, libcurl, OpenSSL, zlib and libgit2 each carry their own terms, and if you redistribute a container built from this Dockerfile you are redistributing those too. That is a packaging question for your own legal review, not something this repository settles.
Editorial conclusion
Adopt github-linguist if you need the same language attribution GitHub shows, or if you are debugging a repository whose language bar looks wrong. Do not adopt it as a general-purpose syntax highlighter or as a classifier for arbitrary text; it works on Git blobs, and its answers come from filename rules, extensions and heuristics rather than from parsing. Before relying on the numbers, run github-linguist --breakdown --strategies on the repository and confirm which files were counted and why, then check whether a .gitattributes override or a vendored path is responsible for anything that looks off.
Frequently asked questions
How do I install github-linguist?
Install it as a Ruby gem with gem install github-linguist, using a recent Ruby rather than the macOS or Xcode-supplied one. Two native dependencies, charlock_holmes and rugged, need system packages such as cmake, pkg-config, ICU, zlib, libcurl and OpenSSL before the gem will build.
What does github-linguist output by default?
Run with no arguments inside a repository, it prints the language breakdown by percentage and file size, one line per language. Adding --breakdown lists the files behind each language, and --json returns the same data as a single JSON object.
Why does github-linguist report a different language for my file?
Detection comes from filename rules, extensions and heuristics, and the README notes that a .gitattributes override can change or confirm the result. Run github-linguist --breakdown --strategies to see the strategy used for each file, including whether .gitattributes overrode it.
Does the --strategies flag work with --json in github-linguist?
No. The README states that the --strategies flag has no effect when the --json flag is present, and that unless --json is specified, --strategies sets --breakdown implicitly.
Can I use github-linguist to analyse a specific branch or tag?
Yes. The --rev REV flag changes the revision being analysed to any gitrevisions-compatible value. The README demonstrates it against Jekyll's gh-pages branch, where the result is 100.00% HTML.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/github-linguist-linguist)
Community notes