# jscpd: clones found by tokens, ranked by kind, reported everywhere

> jscpd is an MIT-licensed copy/paste detector for source code, written in Rust and covering 224 language formats through per-language tokenization and a rolling Rabin-Karp hash. Beyond exact duplicates it detects renamed and near-miss clones, experimental semantic duplicates via code embeddings, dead code, complexity and git-history duplication trends, and it ships as one self-contained binary with fifteen reporters and an MCP server for AI agents.

**kucherenko/jscpd** — Copy/paste detector for source code. 220+ languages, Rust engine, SARIF/HTML/badge reporters, GitHub Action, MCP server for AI agents.

- Repository: https://github.com/kucherenko/jscpd
- Website: https://jscpd.dev
- Stars: 6,282 · Forks: 265
- Language: Rust
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/kucherenko-jscpd

## A clone is repeated tokens, never repeated text

jscpd's foundational decision is what it compares. It tokenizes each of its 224 supported formats the way that language defines it, applying its own comment and string rules rather than treating source as generic text, so a comment change or a reformatted string literal does not create or destroy a clone. For JavaScript, TypeScript, JSX and TSX it uses the oxc parser; for Vue single-file components, Svelte, Astro, Markdown and Razor it extracts the embedded languages and analyzes them properly; and token classification into keywords, identifiers and literals means the comparison happens at language granularity. The matching itself runs a rolling Rabin-Karp hash over token sequences, the classic algorithm for finding repeated substrings at scale. Cross-format detection extends the idea to mixed codebases, with --cross-formats groups matching clones across JavaScript and TypeScript. The full mechanism lives in the documentation's How detection works section, and the complete format list with file extensions is maintained in FORMATS.md.

## Three clone taxonomies, opt-in one at a time

Above exact duplication sits a taxonomy the tool makes explicit. Type-2 clones, blocks differing only in names, literal values or annotations, appear through --ignore-identifiers, --ignore-literals and --ignore-annotations, reported under the kind renamed. Type-3 near-miss clones arrive two ways: --max-gap-lines N merges a copy containing a few inserted or changed lines into one similar clone carrying a similarity score, while --similarity 0.85 compares whole JavaScript and TypeScript functions by syntax-tree structure, catching renames and scattered edits. The experimental Type-4 pass, --semantic, embeds functions using a code embedding model, either one jscpd runs itself after a one-time --semantic-download or any OpenAI-compatible API, and reports functions that do the same thing written differently, the README's example being a rule enforced in a Rust backend and repeated in a Svelte frontend, across seventeen languages from JavaScript and Python to Swift and Scala. Defaults stay conservative, reporting only exact clones, and the --kind filter keeps the statistics and threshold aligned with whichever kinds you actually care about. Because the threshold follows the filter, a team can hold exact clones to a strict budget while trialing renamed detection at a looser one, which is how stricter detection usually earns its way into a pipeline without a noisy debut.

## --compare: port parity between two codebases

The semantic machinery enables a second, more unusual mode: jscpd --compare python/ typescript/ pairs the functions of two folders with the same embedding model and lists which functions of each have a counterpart in the other. The use cases the README names are concrete, what remains to be ported from one language or platform to another, or what the Android version of an application has that the iOS one lacks. This reframes a copy/paste detector as a migration instrument: instead of asking where one codebase repeats itself, it asks where two codebases already agree, and the complement of that set is the porting backlog. For teams maintaining parallel implementations, a polyglot service and its rewrite, or client apps on competing platforms, the mode turns an afternoon of diffing and guessing into a generated list, and its experimental label is the honest caveat that the pairing quality depends on the embedding model's judgment of what counts as the same function. Teams tend to reach for this mode at the exact moment a port begins, which is also when the list of unmatched functions is worth the most, before the two codebases drift further apart.

## Dead code, complexity, git history and one health score

Version 5 grows past duplication into a small codebase health instrument. --dead-code finds unreferenced code, --complexity ranks files by complexity, --history tracks duplication over the git history to show whether the codebase is getting better or worse rather than merely how it stands, and --health rolls the measurements into a single score for the whole repository. The documentation table confirms the operational surface around these, baseline files for comparing against an accepted level, summary output, a dashboard, and blame integration attributing clones to commits. The badge reporter turns the health score into an embeddable image for the README, which closes the loop: the tool that measures the codebase also publishes the measurement. For a team deciding where refactoring investment pays off, the combination of clone kinds, complexity ranking and duplication trend is a triangulation rather than a single dubious number, and the health score is the executive summary of the three. The history view also disciplines arguments, since a duplication figure with a trend line ends the recurring debate about whether the codebase is degrading, replacing opinion with a direction anyone can read.

## One binary, every package manager

Distribution is unusually thorough for a Rust tool. The quick start offers three shapes, a shell script for macOS and Linux, a PowerShell one-liner for Windows, and a no-install run through npx:

```bash
# macOS / Linux
curl -fsSL https://jscpd.dev/install.sh | bash

# Windows (PowerShell)
irm https://jscpd.dev/install.ps1 | iex

# No install — run once with npx (Node.js)
npx jscpd .
```

The install table then covers npm under two names, jscpd and cpd exposing the same binary, PyPI with plain pip, pipx, uv tool and uvx variants, cargo builds from crates.io, Homebrew, Nix, and a multi-arch Docker image at ghcr.io built from the release binaries. The npm package specifically ships a prebuilt binary with no Node.js requirement at runtime, the second npm name exists for muscle memory, cpd being the traditional name in this tool category, and the Docker image's runtime stage is distroless, containing only the verified binary. This breadth is a policy statement: a duplication check that is trivial to install gets adopted, and one that requires a toolchain decision does not.

## CI by default: SARIF, thresholds and pre-commit

The GitHub Action is two lines plus a threshold:

```yaml
- uses: kucherenko/jscpd@v5
  with:
    threshold: 5
```

By default it uploads SARIF results to GitHub Code Scanning, so clones surface as code scanning findings rather than as console noise a developer has to scroll past, and the threshold fails the check when duplication exceeds the configured percentage. Fifteen reporters cover the remaining destinations, console and console-full, json, xml, csv, html, markdown, badge, sarif, codeclimate, openmetrics, an ai reporter, xcode and threshold variants, with clone kinds carried through everywhere, JSON gaining kind, similarity and method fields and SARIF distinguishing jscpd/duplicate-code, renamed-code and similar-code rules. Pre-commit integration has a subtle packaging story documented in the repository's own pyproject.toml: the root project is unpublished and exists so pre-commit can pip install the repository and receive the jscpd binary from the matching PyPI wheel, with a sync-version script keeping the pins aligned.

## An MCP server for agents, and a tool that dogfoods

The AI-facing surface is first-class rather than bolted on. jscpd --mcp runs a stdio MCP server, and the Docker image's default stage is that server, with a glama.json feeding the MCP listing services that score such endpoints, an arrangement the Dockerfile comments explain outright. An AI reporter formats results token-efficiently for agents, and a skills/ directory carries agent skills, the README badge pointing at a skills registry. The repository also dogfoods visibly, a .jscpd.json at the root runs the detector on itself, benchmark/ tracks the engine's own performance, fixtures/ contains one demo directory per feature with expected outputs, and CITATION.cff supports academic citation. Release cadence is brisk, v5.3.1 on 2026-09-21, v5.3.2 on 2026-09-23 and v5.3.3 on 2026-09-28, the same day as the last push, and the v4 README is kept in the tree for teams still on the previous major line, a courtesy with real value since major-version jumps in analysis tools change what thresholds mean and migrating teams need the old reference beside the new one.

## Conclusion

Use jscpd when duplication should be a measured, tracked property of the codebase rather than an anecdote, with exact, renamed and near-miss clones classified separately and SARIF flowing into Code Scanning. Prefer PMD's Copy Paste Detector when your stack is JVM-centric and already carries PMD, since jscpd's edge is breadth of formats, packaging and reporters rather than raw Java-depth. Verify first which clone kinds your thresholds should count, since default runs report exact clones only, budget the semantic pass's model download or API endpoint explicitly, and pin the version consistently across npm, pip and pre-commit, which the project's own sync tooling exists to keep aligned.

## FAQ

### What is jscpd?

jscpd is an MIT-licensed copy/paste detector for source code, supporting 224 language formats through per-language tokenization and a rolling Rabin-Karp hash. Written in Rust, it ships as a self-contained binary with fifteen reporters including SARIF and HTML, plus an MCP server for AI agents.

### How do you use jscpd?

Run npx jscpd . inside a project for a one-off scan, or install through curl, npm, pip, cargo, brew, nix or Docker, then run jscpd /path/to/code. In CI, the kucherenko/jscpd@v5 GitHub Action uploads SARIF results to GitHub Code Scanning by default.

### Which languages does jscpd support?

It tokenizes 224 formats with per-language comment and string rules, using the oxc parser for JavaScript, TypeScript, JSX and TSX, and extracting embedded languages from Vue, Svelte, Astro, Markdown and Razor files. The full list with file extensions is maintained in FORMATS.md.

## Sources

- [kucherenko/jscpd on GitHub](https://github.com/kucherenko/jscpd)
- [License: MIT](https://github.com/kucherenko/jscpd/blob/master/LICENSE)
- [Project website](https://jscpd.dev)
- [README](https://github.com/kucherenko/jscpd/blob/master/README.md)
- [Releases](https://github.com/kucherenko/jscpd/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/kucherenko-jscpd
