PMD: a multilanguage static analyzer with 400+ rules and a copy-paste detector
An extensible multilanguage static code analyzer.
At a glance
- What is it?
- PMD parses Java, Apex and 16 other languages into ASTs and runs XPath or Java rules against them. It ships with CPD for duplicated code, and the README points to a binary zip rather than a package manager.
- Who is it for?
- Adopt PMD if you need a rule engine you can extend with XPath or Java, or if you need CPD's duplicate detection across Java, Apex, C/C++, Python and the other languages listed in the README. Skip it if you want a zero-configuration linter, or if your codebase is in a language the rule set does not cover: Scala is parsed but the README states there are currently no Scala rules available.
- Can I use it commercially?
- Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
- Is it still maintained?
- Yes. The repository received new commits within the last day.
- What is it written in?
- Mainly Java, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What PMD actually checks, and who ends up using it
PMD is a static analyzer that reads source files without executing them. The README describes its output as common programming flaws: unused variables, empty catch blocks, unnecessary object creation, and so forth. Those are the kinds of findings a reviewer would otherwise catch by reading a diff, which is why the tool fits into a pre-commit hook or a build step rather than into a runtime monitor.
The README states PMD is mainly concerned with Java and Apex, and supports 16 other languages. Java and Apex are where the rule coverage is deepest; the other languages are parsed and analyzed but with thinner rule sets. The audience is therefore Java shops, Salesforce teams working in Apex and Visualforce, and anyone who wants one analyzer covering a polyglot repository instead of one linter per language.
A second audience exists and is easy to miss: people who need custom rules. Because rules can be written in Java or as XPath queries, PMD can be pointed at conventions that no general-purpose linter knows about, such as a house rule about how a particular framework annotation must be used. That extensibility is the reason the project describes itself as extensible rather than as a fixed rule pack.
How the parsing pipeline and rule engine fit together
The mechanism is a two-stage pipeline. First, source files are parsed into abstract syntax trees using JavaCC and Antlr, as the README states. Second, rules run against those trees and report violations. Nothing in this design requires the code to compile, which is why PMD can analyze a file in isolation and why it can report on a partial checkout.
The repository layout mirrors that split. pmd-core holds the engine, pmd-cli holds the command line entry point, and there is a module per language: pmd-java, pmd-apex, pmd-kotlin, pmd-swift, pmd-plsql, pmd-go, pmd-ruby and many more. pmd-languages-deps and pmd-lang-test exist to wire the language modules together and to provide shared test scaffolding. pmd-test and pmd-test-schema support rule testing, and pmd-doc generates documentation. The practical consequence is that adding a language is a module, not a fork of the core.
The two rule styles have different costs. An XPath rule is declarative and lives in a rule set XML file, so it can be reviewed and changed without a build. A Java rule is compiled code and gets access to the full API, which is more powerful but ties the rule to the PMD version it was written against. Teams that only ever write XPath rules avoid that coupling; teams that write Java rules inherit an upgrade obligation. The README does not describe a compatibility policy for Java rules across major versions, so that obligation has to be checked against the release notes at upgrade time.
CPD, the copy-paste detector, is a separate tool bundled in the same project. It works on tokens rather than on ASTs, which is why its language list is longer than the analyzer's and includes C/C++, C#, Dart, Fortran, Gherkin, Julia, Lua, Matlab, Objective-C, Perl, PHP, Python, Ruby, T-SQL and TypeScript.
Installing PMD from the release zip and running a first check
The README does not describe a package manager install. It says to download the latest binary zip from the releases page and extract it somewhere, then run the launcher script. On a Unix-like system that is bin/pmd; on Windows it is bin\pmd.bat. The subcommand shown in the README is check.
bin/pmd checkThat is the whole invocation the README gives, and the Windows equivalent is the same subcommand through the batch file:
bin\pmd.bat checkThe README does not enumerate flags. It links to a Getting Started page in the documentation, and that page is the reliable place to learn what your version accepts, because flag names differ between PMD versions and an example copied from an older release can fail silently.
What the reader should see from a check run is one violation per line, each naming the file, the line, the rule and a short message. If the output is empty, that is a result too: either the rule set found nothing or the target did not contain parsable files. Pointing the command at the wrong path produces an empty report rather than an error in most CLI designs, so verify the path before concluding the code is clean.
The README also notes that there are plugins for Maven and Gradle as well as for various IDEs, and links to a Tools / Integrations page. For a Maven or Gradle build, the plugin is usually a better fit than shelling out to the zip, because the plugin manages the PMD version alongside the build.
Where PMD stops being the right tool
The clearest limitation is stated in the README itself: Scala is supported, but there are currently no Scala rules available. A Scala project can therefore run CPD and can use the parser, but gets nothing from the analyzer's built-in rule set. The same asymmetry applies more broadly: the language list is long, but the README's own framing puts Java and Apex at the center, so a project in one of the peripheral languages should check rule coverage before assuming parity.
A second limitation is the rule set itself. PMD reports stylistic and structural issues, not type errors, not security dataflow, and not dependency problems. The examples in the README are unused variables and empty catch blocks. A team expecting a security scanner will be disappointed by the category of findings, and a team that already runs a compiler with strict warnings will find some overlap.
A third issue is triage cost. Four hundred plus built-in rules is a large surface, and enabling a broad rule set on an existing codebase produces a backlog rather than a clean build. The README does not describe a baseline or suppression workflow, so the practical approach is to start with one rule set, fix or suppress what it reports, and add rule sets incrementally. Suppression is a per-file or per-line concern in most analyzers of this kind, and PMD's documentation, not the README, is where the syntax lives.
Finally, the license field in the repository metadata is NOASSERTION while the README links to a BSD Style LICENSE file. That mismatch matters to anyone running an automated license check: the tooling that reads the metadata will not classify it, even though the project states its license in prose.
PMD against a compiler-warning-only workflow
The obvious alternative is to rely on the Java compiler's own lint warnings plus whatever the IDE already reports. The difference in approach is scope. A compiler warning is derived from the type system and the language specification; it can only flag things the compiler knows are suspicious, such as an unchecked cast or a deprecated call. PMD works from the syntax tree and from token streams, so it can flag patterns that are perfectly legal and type-correct: an empty catch block, an object created inside a loop, a duplicated block of twenty lines copied between two classes. Those are exactly the findings a compiler will never produce.
The trade-off runs the other way too. Compiler warnings have no false-positive problem in the sense PMD does, because they are grounded in semantics. A PMD rule that matches a syntactic pattern can fire on code that is intentional. That is why the XPath rule style is both the project's strength and its maintenance burden: a rule you wrote is a rule you have to tune.
For duplicate detection specifically, CPD has little competition inside a standard Java toolchain. The compiler has no notion of copy-paste, and IDE refactoring tools suggest duplication only within a file or a narrow scope. CPD's token-based approach and its long language list make it usable in mixed repositories where a single-language tool would need a second installation.
Maintenance, release cadence and what an upgrade costs
The repository is not archived, and the last push was on 2026-09-22. Releases are frequent: the material lists pmd_releases/7.27.0 dated 2026-08-28, pmd_releases/7.26.0 dated 2026-06-29, and a 7.28.0-SNAPSHOT line. A snapshot release exists alongside tagged ones, which tells you the project publishes development builds rather than holding everything until a tag.
The upgrade cost depends on how you extend PMD. If your rules are XPath queries in rule set XML, an upgrade is mostly a matter of checking whether a rule was renamed, moved between categories, or removed. If your rules are Java classes compiled against the PMD API, an upgrade can require recompilation and API adjustments, and the README gives no compatibility guarantee for that surface. Custom rule authors should read the release notes for the version they are moving to before changing the dependency.
The project also publishes a Docker image, referenced by the Docker badge in the README pointing at pmdcode/pmd on Docker Hub. For CI that wants a pinned analyzer without a JDK setup step, that is an alternative to the zip, though the README does not document the image's entry point or its default command.
On licensing, the README links to a BSD Style LICENSE file and the repository metadata reports NOASSERTION. The BSD style is permissive, which generally means embedding the analyzer in a build pipeline does not trigger copyleft obligations, but the metadata mismatch means automated scanners may flag the dependency for manual review. This is a description of what the files say, not legal advice; a team with a compliance process should read the LICENSE and NOTICE files directly.
Editorial conclusion
Adopt PMD if you need a rule engine you can extend with XPath or Java, or if you need CPD's duplicate detection across Java, Apex, C/C++, Python and the other languages listed in the README. Skip it if you want a zero-configuration linter, or if your codebase is in a language the rule set does not cover: Scala is parsed but the README states there are currently no Scala rules available. Before rolling it out, verify the two things the README leaves open: which rule sets your build actually enables, and whether the reported violations are worth the triage time on your code. Start with bin/pmd check on a single module and read the output before wiring it into CI.
Frequently asked questions
What does PMD stand for?
The README does not expand the acronym. It presents PMD only as the name of the source code analyzer and links to the project documentation, which is where an expansion would appear if the project gives one.
What is PMD used for?
PMD is an extensible multilanguage static code analyzer. It finds common programming flaws such as unused variables, empty catch blocks and unnecessary object creation, and it includes CPD, a copy-paste detector for duplicated code.
How do I install PMD and run it on a project?
Download the latest binary zip from the releases page, extract it, then run bin/pmd check on Unix-like systems or bin\pmd.bat check on Windows. The README points to the Getting Started page in the documentation for the full setup.
Which languages does PMD support?
The README lists Java, JavaScript, Apex and Visualforce, Kotlin, Swift, Modelica, PL/SQL, Apache Velocity, JSP, WSDL, Maven POM, HTML, XML and XSL, with Scala parsed but currently having no Scala rules available. CPD covers a longer list that adds C/C++, C#, Go, Python, Ruby, TypeScript and others.
Can I write my own PMD rules?
Yes. The README states that rules can be written in Java or using an XPath query, which is what makes the analyzer extensible. Java rules compile against the PMD API, while XPath rules live in rule set files.
Is there a Maven or Gradle plugin for PMD?
The README states there are plugins for Maven and Gradle as well as for various IDEs, and links to a Tools / Integrations documentation page. It does not list plugin coordinates or versions.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/pmd-pmd)