cd /news/ai-agents/stop-counting-duplicated-lines-i-ran… · home › topics › ai-agents › article
[ARTICLE · art-142079] src=dev.to ↗ pub= topic=ai-agents verified=true sentiment=↑ positive

Stop counting duplicated lines: I rank copy-paste by what it costs me

A developer built dry-mcp, an MCP server that detects duplicated code by semantic similarity rather than token matching, ranking findings by maintenance cost and delivering results to AI agents instead of a dashboard. The tool combines exact hashing with embeddings from jina-embeddings-v2-base-code, which the developer measured at roughly 0.53 similarity for renamed copies and 0.87 for the same logic ported across languages, versus 0.22 or less for unrelated code. Each finding carries a confidence level (certain, high, moderate, low), and replies disclose when the background index is incomplete.

by read6 min views1 publishedSep 29, 2026

Every CI pipeline I have worked with had a duplication report. I cannot remember the last time I read one.

The number moves from 4.2% to 4.5%, a list of blocks shows up, and half of them are constructors, guard clauses and property mappings that look alike because that is how the language is written. Meanwhile the copy that actually hurt me never made the list: someone (oops, me) copied a function, renamed three variables, and a bug fix later landed in one copy but not the other.

We blame tight deadlines, but can't blame StackOverflow anymore, so maybe it's time to clean the code bases.

So I built dry-mcp, an MCP server that finds duplicated code by meaning instead of by token, ranks it by what it costs to keep, and gives the result to my AI agent instead of a dashboard.

SonarQube is good at what it is built for, so let's be precise about what that is. According to its documentation, a block counts as duplicated when there are at least 100 successive duplicated tokens spread over at least 10 lines (Java uses 10 successive statements instead). Differences in indentation and string literals are ignored.

Two consequences follow:

That is the right design for a quality gate ("new code may not exceed X% duplication"). It is the wrong design for the question I actually have: where is the duplication that is worth an afternoon?

Matching runs in two passes.

First, every block is normalized (formatting and comments stripped) and hashed. Identical hashes are exact copies. That is free and never wrong.

What is left is compared using embeddings from jina-embeddings-v2-base-code, a model trained on code. Two blocks that do the same thing land close together even when every name differs. Measured against that model:

Pair of blocks Similarity
Same logic, every name changed ~0.53
Same logic, different language ~0.87
Unrelated code ≤ 0.22

Yes, the second row means a helper that was ported to another language is still recognized as the same code. I did not set out to build that, but it falls out of matching by meaning.

A 60-line block copied four times is a real problem: every change has to be made four times, and one will be forgotten. A 3-line fragment repeated forty times is almost always an idiom.

So the ranking counts block size for more than the number of copies. Want the most-copied code instead? Ask for orderBy: "frequency".

Repetition that is just how the language is written gets demoted instead of ranked:

Nothing is thrown away. includeSuppressed: true returns the demoted groups with the reason for each, so the rules can be checked.

Devs are lazy, including me. I don't want to read a duplication report. I want my agent to find the worst copy and fix it. That changes what the output needs to carry.

Confidence on every finding. Near-miss matching is deliberately inclusive, because missing a large repeated block is worse than flagging a coincidence. So instead of silently filtering, each finding is marked certain, high, moderate or low, and the reply explains the scale. The agent knows what it can act on and what it has to read first.

Honesty about completeness. Indexing runs in the background (more on that below). Any reply built from an incomplete index says so, with numbers. A partial answer is never dressed up as a full one.

Self-service scope. Without configuration, every source file under the root is analysed. When that looks too wide, the reply says so, names the largest folders, and tells the agent what to write in duplication.config.json. The agent edits the file, and the next question uses the new scope. No restart.

The whole surface is four tools:

Tool Answers
detect_duplication Where is the duplication, worst first?
duplication_status Is the index ready?
explain_duplication Show me every copy.
reindex Start over.

A trimmed finding looks like this:

duplications:
 - id: 622069fe1577
   occurrences:
    - file: src/analysis/clusterer.ts
      startLine: 58
      endLine: 177
      lines: 78
    - file: src/analysis/duplication-service.ts
      startLine: 322
      endLine: 441
      lines: 81
   frequency: 5
   medianLines: 81
   removableLines: 324
   severity: 1692.69
   similarity: 0.763
   confidence: high
summary:
 clustersFound: 98
 byConfidence:
  certain: 4
  high: 50
  moderate: 44

The agent picks the top finding, calls explain_duplication with its id to get the source of every copy, and decides how to merge them.

There is no grammar to install. Block boundaries are inferred from braces and indentation, so C#, TypeScript, Java, Go, Rust, Python, Ruby, PHP, SQL, shell, CSS and friends all work out of the box.

The trade-off: line ranges are approximate. The returned source is authoritative, the numbers around it are a pointer. For an agent, that's not a problem since it will scan all and decide what to fix.

The embedding model is not bundled. Even the smallest weights are around 160 MB, which is too much to push through npm install, and a download that arrives unannounced in the middle of a question is worse than being told once to run a command:

node dist/index.js download-model          # int8, ~160 MB
node dist/index.js download-model --fp32   # ~640 MB, for accuracy

int8 is the default because it is the precision CPUs actually accelerate. x86 cores without AVX512-FP16 have no native fp16 compute, so an fp16 model often runs slower than fp32, while int8 uses the VNNI instructions directly.

Without the model the server still starts. Queries return an empty result with the command to run, not an error.

Everything runs locally on the CPU. No API key, no code leaving the machine. The price is that embedding a whole project takes minutes, not seconds.

It is slower still by design. I typically have several projects open, each with its own server, on the same machine I am compiling on. Left alone, embedding would take every core it can get. So background indexing runs at a 20% duty cycle: it works in short slices and rests in between. All servers together cost less than one core, and a project still finishes within an editor session.

While indexing is in progress, replies carry a progress block:

progress:
 filesInScope: 3510
 filesEmbedded: 1204
 pendingFiles: 2306
 percentComplete: 34

For projects larger than a couple of hundred files, you get only that block until half the project is embedded. A ranking drawn from a third of a codebase is not an early version of the real ranking. The worst duplication is most likely in the part not read yet, while the reply would look like an answer and invite acting on it.

Two things make this bearable:

include list in duplication.config.json before the first run. Fewer files, faster sync, and less vendored code in the results anyway. Node 22.5 or later.

git clone https://github.com/jgauffin/dry-mcp
cd duplication-mcp
npm install && npm run build
node dist/index.js download-model

Add it to .mcp.json in your project:

{
  "mcpServers": {
    "duplication": {
      "command": "node",
      "args": ["/path/to/duplication-mcp/dist/index.js", "."]
    }
  }
}

Optionally scope it:

{
  "include": ["src/**", "lib/**"],
  "exclude": ["**/*.generated.*", "**/migrations/**"]
}

Then ask your agent where the worst duplication is. If it tells you it is still indexing, that is the honest answer. Ask again in a few minutes.

I'd love to hear what it finds in your codebase, and especially where it is wrong.

── more in #ai-agents 4 stories · sorted by recency
── more on @dry-mcp 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/stop-counting-duplic…] indexed:0 read:6min 2026-09-29 · —