# Stop counting duplicated lines: I rank copy-paste by what it costs me

> Source: <https://dev.to/jgauffin/stop-counting-duplicated-lines-i-rank-copy-paste-by-what-it-costs-me-3mp4>
> Published: 2026-09-29 21:33:55+00:00

Every CI pipeline I have worked with had a duplication report. I cannot remember the last time I read one.

The number moves from 4.2% to 4.5%, a list of blocks shows up, and half of them are constructors, guard clauses and property mappings that look alike because that is how the language is written. Meanwhile the copy that actually hurt me never made the list: someone (oops, me) copied a function, renamed three variables, and a bug fix later landed in one copy but not the other.

We blame tight deadlines, but can't blame StackOverflow anymore, so maybe it's time to clean the code bases.

So I built **dry-mcp**, an MCP server that finds duplicated code by meaning instead of by token, ranks it by what it costs to keep, and gives the result to my AI agent instead of a dashboard.

SonarQube is good at what it is built for, so let's be precise about what that is. According to [its documentation](https://docs.sonarsource.com/sonarqube-server/10.4/user-guide/metric-definitions), a block counts as duplicated when there are at least 100 successive duplicated tokens spread over at least 10 lines (Java uses 10 successive statements instead). Differences in indentation and string literals are ignored.

Two consequences follow:

That is the right design for a quality gate ("new code may not exceed X% duplication"). It is the wrong design for the question I actually have: **where is the duplication that is worth an afternoon?**

Matching runs in two passes.

First, every block is normalized (formatting and comments stripped) and hashed. Identical hashes are exact copies. That is free and never wrong.

What is left is compared using embeddings from [`jina-embeddings-v2-base-code`](https://huggingface.co/jinaai/jina-embeddings-v2-base-code), a model trained on code. Two blocks that do the same thing land close together even when every name differs. Measured against that model:

| Pair of blocks | Similarity | 
|---|---|
| Same logic, every name changed | ~0.53 | 
| Same logic, different language | ~0.87 | 
| Unrelated code | ≤ 0.22 | 

Yes, the second row means a helper that was ported to another language is still recognized as the same code. I did not set out to build that, but it falls out of matching by meaning.

A 60-line block copied four times is a real problem: every change has to be made four times, and one will be forgotten. A 3-line fragment repeated forty times is almost always an idiom.

So the ranking counts block size for more than the number of copies. Want the most-copied code instead? Ask for `orderBy: "frequency"`.

Repetition that is just how the language is written gets demoted instead of ranked:

Nothing is thrown away. `includeSuppressed: true` returns the demoted groups with the reason for each, so the rules can be checked.

Devs are lazy, including me. I don't want to read a duplication report. I want my agent to find the worst copy and fix it. That changes what the output needs to carry.

**Confidence on every finding.** Near-miss matching is deliberately inclusive, because missing a large repeated block is worse than flagging a coincidence. So instead of silently filtering, each finding is marked `certain`, `high`, `moderate` or `low`, and the reply explains the scale. The agent knows what it can act on and what it has to read first.

**Honesty about completeness.** Indexing runs in the background (more on that below). Any reply built from an incomplete index says so, with numbers. A partial answer is never dressed up as a full one.

**Self-service scope.** Without configuration, every source file under the root is analysed. When that looks too wide, the reply says so, names the largest folders, and tells the agent what to write in `duplication.config.json`. The agent edits the file, and the next question uses the new scope. No restart.

The whole surface is four tools:

| Tool | Answers | 
|---|---|
| `detect_duplication` | Where is the duplication, worst first? | 
| `duplication_status` | Is the index ready? | 
| `explain_duplication` | Show me every copy. | 
| `reindex` | Start over. | 

A trimmed finding looks like this:

```
duplications:
 - id: 622069fe1577
   occurrences:
    - file: src/analysis/clusterer.ts
      startLine: 58
      endLine: 177
      lines: 78
    - file: src/analysis/duplication-service.ts
      startLine: 322
      endLine: 441
      lines: 81
    # ...three more
   frequency: 5
   medianLines: 81
   removableLines: 324
   severity: 1692.69
   similarity: 0.763
   confidence: high
summary:
 clustersFound: 98
 byConfidence:
  certain: 4
  high: 50
  moderate: 44
```

The agent picks the top finding, calls `explain_duplication` with its id to get the source of every copy, and decides how to merge them.

There is no grammar to install. Block boundaries are inferred from braces and indentation, so C#, TypeScript, Java, Go, Rust, Python, Ruby, PHP, SQL, shell, CSS and friends all work out of the box.

The trade-off: line ranges are approximate. The returned source is authoritative, the numbers around it are a pointer. For an agent, that's not a problem since it will scan all and decide what to fix.

The embedding model is not bundled. Even the smallest weights are around 160 MB, which is too much to push through `npm install`, and a download that arrives unannounced in the middle of a question is worse than being told once to run a command:

```
node dist/index.js download-model          # int8, ~160 MB
node dist/index.js download-model --fp32   # ~640 MB, for accuracy
```

int8 is the default because it is the precision CPUs actually accelerate. x86 cores without AVX512-FP16 have no native fp16 compute, so an fp16 model often runs *slower* than fp32, while int8 uses the VNNI instructions directly.

Without the model the server still starts. Queries return an empty result with the command to run, not an error.

Everything runs locally on the CPU. No API key, no code leaving the machine. The price is that embedding a whole project takes minutes, not seconds.

It is slower still by design. I typically have several projects open, each with its own server, on the same machine I am compiling on. Left alone, embedding would take every core it can get. So background indexing runs at a **20% duty cycle**: it works in short slices and rests in between. All servers together cost less than one core, and a project still finishes within an editor session.

While indexing is in progress, replies carry a progress block:

```
progress:
 filesInScope: 3510
 filesEmbedded: 1204
 pendingFiles: 2306
 percentComplete: 34
```

For projects larger than a couple of hundred files, you get *only* that block until half the project is embedded. A ranking drawn from a third of a codebase is not an early version of the real ranking. The worst duplication is most likely in the part not read yet, while the reply would look like an answer and invite acting on it.

Two things make this bearable:

`include` list in `duplication.config.json` before the first run. Fewer files, faster sync, and less vendored code in the results anyway.
Node 22.5 or later.

```
git clone https://github.com/jgauffin/dry-mcp
cd duplication-mcp
npm install && npm run build
node dist/index.js download-model
```

Add it to `.mcp.json` in your project:

```
{
  "mcpServers": {
    "duplication": {
      "command": "node",
      "args": ["/path/to/duplication-mcp/dist/index.js", "."]
    }
  }
}
```

Optionally scope it:

```
{
  "include": ["src/**", "lib/**"],
  "exclude": ["**/*.generated.*", "**/migrations/**"]
}
```

Then ask your agent where the worst duplication is. If it tells you it is still indexing, that is the honest answer. Ask again in a few minutes.

I'd love to hear what it finds in your codebase, and especially where it is wrong.
