Slopo solves problems two and three, and the file says which one it skips
Embedding-based code duplication detector
At a glance
- What is it?
- An embedding-based duplication detector for code that was never copy-pasted, ranked so that distant copies surface first. The default model comes from a benchmark the author wrote, the free tier is non-commercial, and the free key may simply not be issued.
- Who is it for?
- This fits a specific workflow and not a general one. The value is in large or messy codebases where an agent re-implements something that already exists three modules away, and where reviewing the whole tree for near-duplicates is not something to ask a model to do.
- Can I use it commercially?
- Yes, with strict conditions. AGPL-3.0 is a network copyleft licence: if people use a modified version over a network, for example as a hosted service, you must offer them its source code under the same licence.
- Is it still maintained?
- Yes. The repository last received commits 8 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 2, 2026, and from our analysis. They are not legal advice.
Editorial analysis
Two of the three scale problems, and the third one is left alone
The case is made as a numbered escalation, and then the tool picks its place in it.
One: reviewing a small change, an agent can make a focused analysis and find similar implementations. Two: working on a large change in a messy codebase, review becomes less reliable and more costly. Three: analysing the whole codebase to identify clusters of similar code is not a task for agents at all. The sentence that follows says plainly which of those are addressed: Slopo solves two and three.
So the small-change case is deliberately out of scope, and for a defensible reason. Not all code duplication is the hardest to detect kind, and coding agents are already able to spot much of it, as the additional benefits section puts it. There are two other claims attached to that decision, both about token spend rather than detection. Embedding API access costs practically nothing next to the models agents run on, and an agent handed a report of duplicates that already exist does not spend tokens rediscovering them. The second claim is about precision rather than cost: agents do their best work on one focused task, and a single cluster with a narrow instruction is that shape.
The further apart the two copies sit, the higher they rank
The output is not a list of near matches, it is clusters, and there are two sort keys.
The first is similarity, which comes from the embeddings themselves. A vector represents the meaning of a piece of text, and when two snippets have close vectors the tool treats them as similar and therefore as potential duplicates, even when they are written in different ways. That is the whole reason embeddings are used here instead of text comparison, and it is what catches a second implementation of the same problem rather than a copied block.
The second key is distance in the codebase, and it is the one that distinguishes this from a normal similarity search. Alongside detecting non-exact duplication, the focus is on code sitting far apart, often spread across different modules or separated within a large file, and the further apart it is the higher the priority. A cluster of two near-identical helpers twenty lines from each other is uninteresting; the same two helpers in two packages is the case worth a reviewer's attention.
The clusters are framed as input rather than as verdicts. They are meant to be handed to an agent that checks whether a cluster is a real duplicate, and the file describes agents being instructed on what exactly to do with results rather than just to report them, so that non-actionable similarity gets discarded quickly.
The default model comes from a benchmark the author wrote
The configuration shipped as the sample is described as the one that gives the best results, and the results come from a linked article written by the author of the tool. That is worth stating plainly rather than reading past: this is a self-benchmark, and the defaults encode its conclusion.
The article covers what embedding models allow to detect, where they are weak, and which work best. The shortlist that comes out of it has two providers, and the second one carries the counter-intuitive result of the whole file. Jina AI is recommended, with code-focused models available through an API and for local use. Voyage AI is also recommended, but only their general-purpose model: their code-focused models are described as not suitable here.
A model tuned on code losing to a general-purpose model is the kind of result that makes a benchmark worth reading, and it is also the kind of result a vendor's own model card would not produce. If you take one thing from the linked research, take that one, and do not assume a code-specific embedding model is the safe default here.
Wider than the shortlist, any provider supported by LiteLLM can be configured, and so can any OpenAI-compatible server, so the choice is not locked to the two names.
Ten million free tokens, non-commercial only, and a guess about why they vanish
The free path is the interesting one and it has three conditions attached to it.
Jina offers API keys with free tokens for non-commercial use, without registration. Paid options exist for commercial use. You open the main page and the token is generated, and the file tells you what you should see: a line saying you have 10,000,000 tokens left in the API key below. And then it warns you that free tokens are not always granted.
The explanation offered for that is, in the file's own words, a guess and not official information: abuse prevention mechanisms based on IP address, or other measures. That candour is worth crediting, because the failure mode is silent in the sense that a key gets generated and then does not work well. A pipeline that assumes a free key exists has to have a paid fallback ready.
Commercial use is where this meets the licence. The tool is AGPL-3.0-or-later, and the free tokens are explicitly for non-commercial use, so a company running this on its own repository needs a paid embedding key on top of accepting the copyleft terms. The two constraints are independent and both apply.
Fourteen grammar packages covering fifteen languages
The supported language list has fifteen entries: Python, TypeScript, TSX, JavaScript, Java, Kotlin, C, C++, C#, Go, Rust, PHP, Elixir, Ruby and Swift. The dependency list has fourteen tree-sitter grammar packages.
The difference is TSX, which arrives with the TypeScript grammar rather than needing its own. Everything else maps one to one: `tree-sitter-c`, `tree-sitter-c-sharp`, `tree-sitter-cpp`, `tree-sitter-elixir`, `tree-sitter-go`, `tree-sitter-java`, `tree-sitter-javascript`, `tree-sitter-kotlin`, `tree-sitter-php`, `tree-sitter-python`, `tree-sitter-ruby`, `tree-sitter-rust`, `tree-sitter-swift` and `tree-sitter-typescript`.
The pinning style is unusual and consistent. Every single dependency uses the compatible release operator rather than a floor: `litellm~=1.102.1`, `numpy~=2.5.3`, `pathspec~=1.1.1`, `python-dotenv~=1.2.3`, `pyyaml~=6.0.3` and `tree-sitter~=0.26.0` alongside every grammar. So a resolver can move within a minor band and no further, and the embedding client in particular is confined to the 1.102 series.
The rest of the packaging is modern and narrow. It needs Python 3.12 or newer, builds with `uv_build` behind a `>=0.11.15,<0.13.0` constraint, exposes a single entry point at `slopo.cli:main`, type checks `src` and `tests` together under mypy, and excludes one directory from ruff, the per-language parsing fixtures under tests.
One uv command to install, and no automatic updates
Installation is two commands and no Python prerequisite.
uv tool install slopoThat uses uv, described as a Python package manager, to install the package from PyPI into an isolated virtual environment, and the file is explicit that you do not need to install Python separately. Upgrading is the second command, and it comes with a warning attached.
uv tool upgrade slopoThere are no automatic updates and no notifications about them, which means a pinned tool stays pinned until someone remembers. There is also no mention of a source install, a container, or a pip path, so uv is the only installation route the file documents.
Configuration starts with `slopo init`, which writes a config file template that carries further instructions inside it. Only two things are required: the directory holding the code to analyse, and the embedding model configuration. Everything else in the generated file is optional, and the file's own framing of the order of difficulty is worth repeating, because it is the opposite of most tools: once you have figured out the model, the file says, the rest is quick.
Three routes to a model, and the API can be asleep on the first call
The API path is the recommended default and it has a behaviour worth planning for. The first request is sometimes delayed, because the servers may be offline and need to start, and subsequent requests are faster. For a tool that embeds a whole codebase on the first run that is a one-off delay, and for a tool used interactively on a small diff it is a stall in the wrong place.
The local path is Ollama with `jina-embeddings-v2-base-code`. It is described as a small model that also runs on a CPU, which makes it the option for a machine with no GPU, and the file is honest that it may be significantly slower than the API depending on the hardware you use. The model itself is pulled from a third-party Ollama namespace rather than the official one, so the image provenance is a decision you are making rather than one the tool makes for you.
The third route is the open one: anything LiteLLM supports, plus any OpenAI-compatible server. That covers self-hosted models and internal gateways, and it is the route that removes the free-token constraint entirely, since a local or self-hosted endpoint has no non-commercial clause attached to it.
What sits on the other side of all three is the same requirement, that every snippet in the analysed tree is embedded and compared. On a large codebase that is the cost that determines whether the free tier is enough, and the file gives no figure for how many tokens a repository of a given size consumes.
Editorial conclusion
This fits a specific workflow and not a general one. The value is in large or messy codebases where an agent re-implements something that already exists three modules away, and where reviewing the whole tree for near-duplicates is not something to ask a model to do. It is not a linter for the easy cases, and the file says so. Before you build on it, resolve three things. The licensing and the model cost pull in opposite directions: the tool is AGPL, and the ten million free tokens are for non-commercial use, so a company run means a paid key. The default model is chosen from the author's own benchmark, so read that article and pick a configuration yourself rather than accepting the sample. And decide whether the cold first request matters, because the API path can sit idle before it starts and a CPU-only local model trades that delay for per-query slowness instead.
Frequently asked questions
What does Slopo detect that a normal duplication tool misses?
Code that was never copy-pasted but is a second implementation of the same problem, written differently. It embeds each snippet so that two pieces of code meaning the same thing land close together in vector space, then returns clusters ranked by similarity and by how far apart the two copies sit in the codebase, with distant pairs ranked higher.
Which embedding model should Slopo use?
Two providers come out of a benchmark written by the tool's own author: Jina AI, whose code-focused models are available through an API or locally, and Voyage AI, but only their general-purpose model, since their code-focused models are described as not suitable here. Any LiteLLM-supported provider or OpenAI-compatible server also works. The local option is Ollama with jina-embeddings-v2-base-code, which runs on a CPU but may be significantly slower than the API.
How do I install and configure Slopo?
`uv tool install slopo` pulls it from PyPI into an isolated environment, and no separate Python install is needed. `uv tool upgrade slopo` is the upgrade path, and there are no automatic updates or notifications. `slopo init` writes a config template where only two things are required: the directory of code to analyse and the embedding model configuration.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/rafal-qa-slopo)