{"slug": "stop-counting-duplicated-lines-i-rank-copy-paste-by-what-it-costs-me", "title": "Stop counting duplicated lines: I rank copy-paste by what it costs me", "summary": "A developer built dry-mcp, an MCP server that detects duplicated code by semantic similarity rather than token matching, ranking findings by maintenance cost and delivering results to AI agents instead of a dashboard. The tool combines exact hashing with embeddings from jina-embeddings-v2-base-code, which the developer measured at roughly 0.53 similarity for renamed copies and 0.87 for the same logic ported across languages, versus 0.22 or less for unrelated code. Each finding carries a confidence level (certain, high, moderate, low), and replies disclose when the background index is incomplete.", "body_md": "Every CI pipeline I have worked with had a duplication report. I cannot remember the last time I read one.\n\nThe number moves from 4.2% to 4.5%, a list of blocks shows up, and half of them are constructors, guard clauses and property mappings that look alike because that is how the language is written. Meanwhile the copy that actually hurt me never made the list: someone (oops, me) copied a function, renamed three variables, and a bug fix later landed in one copy but not the other.\n\nWe blame tight deadlines, but can't blame StackOverflow anymore, so maybe it's time to clean the code bases.\n\nSo I built **dry-mcp**, an MCP server that finds duplicated code by meaning instead of by token, ranks it by what it costs to keep, and gives the result to my AI agent instead of a dashboard.\n\nSonarQube is good at what it is built for, so let's be precise about what that is. According to [its documentation](https://docs.sonarsource.com/sonarqube-server/10.4/user-guide/metric-definitions), a block counts as duplicated when there are at least 100 successive duplicated tokens spread over at least 10 lines (Java uses 10 successive statements instead). Differences in indentation and string literals are ignored.\n\nTwo consequences follow:\n\nThat is the right design for a quality gate (\"new code may not exceed X% duplication\"). It is the wrong design for the question I actually have: **where is the duplication that is worth an afternoon?**\n\nMatching runs in two passes.\n\nFirst, every block is normalized (formatting and comments stripped) and hashed. Identical hashes are exact copies. That is free and never wrong.\n\nWhat is left is compared using embeddings from [`jina-embeddings-v2-base-code`](https://huggingface.co/jinaai/jina-embeddings-v2-base-code), a model trained on code. Two blocks that do the same thing land close together even when every name differs. Measured against that model:\n\n| Pair of blocks | Similarity | \n|---|---|\n| Same logic, every name changed | ~0.53 | \n| Same logic, different language | ~0.87 | \n| Unrelated code | ≤ 0.22 | \n\nYes, the second row means a helper that was ported to another language is still recognized as the same code. I did not set out to build that, but it falls out of matching by meaning.\n\nA 60-line block copied four times is a real problem: every change has to be made four times, and one will be forgotten. A 3-line fragment repeated forty times is almost always an idiom.\n\nSo the ranking counts block size for more than the number of copies. Want the most-copied code instead? Ask for `orderBy: \"frequency\"`.\n\nRepetition that is just how the language is written gets demoted instead of ranked:\n\nNothing is thrown away. `includeSuppressed: true` returns the demoted groups with the reason for each, so the rules can be checked.\n\nDevs are lazy, including me. I don't want to read a duplication report. I want my agent to find the worst copy and fix it. That changes what the output needs to carry.\n\n**Confidence on every finding.** Near-miss matching is deliberately inclusive, because missing a large repeated block is worse than flagging a coincidence. So instead of silently filtering, each finding is marked `certain`, `high`, `moderate` or `low`, and the reply explains the scale. The agent knows what it can act on and what it has to read first.\n\n**Honesty about completeness.** Indexing runs in the background (more on that below). Any reply built from an incomplete index says so, with numbers. A partial answer is never dressed up as a full one.\n\n**Self-service scope.** Without configuration, every source file under the root is analysed. When that looks too wide, the reply says so, names the largest folders, and tells the agent what to write in `duplication.config.json`. The agent edits the file, and the next question uses the new scope. No restart.\n\nThe whole surface is four tools:\n\n| Tool | Answers | \n|---|---|\n| `detect_duplication` | Where is the duplication, worst first? | \n| `duplication_status` | Is the index ready? | \n| `explain_duplication` | Show me every copy. | \n| `reindex` | Start over. | \n\nA trimmed finding looks like this:\n\n```\nduplications:\n - id: 622069fe1577\n   occurrences:\n    - file: src/analysis/clusterer.ts\n      startLine: 58\n      endLine: 177\n      lines: 78\n    - file: src/analysis/duplication-service.ts\n      startLine: 322\n      endLine: 441\n      lines: 81\n    # ...three more\n   frequency: 5\n   medianLines: 81\n   removableLines: 324\n   severity: 1692.69\n   similarity: 0.763\n   confidence: high\nsummary:\n clustersFound: 98\n byConfidence:\n  certain: 4\n  high: 50\n  moderate: 44\n```\n\nThe agent picks the top finding, calls `explain_duplication` with its id to get the source of every copy, and decides how to merge them.\n\nThere is no grammar to install. Block boundaries are inferred from braces and indentation, so C#, TypeScript, Java, Go, Rust, Python, Ruby, PHP, SQL, shell, CSS and friends all work out of the box.\n\nThe trade-off: line ranges are approximate. The returned source is authoritative, the numbers around it are a pointer. For an agent, that's not a problem since it will scan all and decide what to fix.\n\nThe embedding model is not bundled. Even the smallest weights are around 160 MB, which is too much to push through `npm install`, and a download that arrives unannounced in the middle of a question is worse than being told once to run a command:\n\n```\nnode dist/index.js download-model          # int8, ~160 MB\nnode dist/index.js download-model --fp32   # ~640 MB, for accuracy\n```\n\nint8 is the default because it is the precision CPUs actually accelerate. x86 cores without AVX512-FP16 have no native fp16 compute, so an fp16 model often runs *slower* than fp32, while int8 uses the VNNI instructions directly.\n\nWithout the model the server still starts. Queries return an empty result with the command to run, not an error.\n\nEverything runs locally on the CPU. No API key, no code leaving the machine. The price is that embedding a whole project takes minutes, not seconds.\n\nIt is slower still by design. I typically have several projects open, each with its own server, on the same machine I am compiling on. Left alone, embedding would take every core it can get. So background indexing runs at a **20% duty cycle**: it works in short slices and rests in between. All servers together cost less than one core, and a project still finishes within an editor session.\n\nWhile indexing is in progress, replies carry a progress block:\n\n```\nprogress:\n filesInScope: 3510\n filesEmbedded: 1204\n pendingFiles: 2306\n percentComplete: 34\n```\n\nFor projects larger than a couple of hundred files, you get *only* that block until half the project is embedded. A ranking drawn from a third of a codebase is not an early version of the real ranking. The worst duplication is most likely in the part not read yet, while the reply would look like an answer and invite acting on it.\n\nTwo things make this bearable:\n\n`include` list in `duplication.config.json` before the first run. Fewer files, faster sync, and less vendored code in the results anyway.\nNode 22.5 or later.\n\n```\ngit clone https://github.com/jgauffin/dry-mcp\ncd duplication-mcp\nnpm install && npm run build\nnode dist/index.js download-model\n```\n\nAdd it to `.mcp.json` in your project:\n\n```\n{\n  \"mcpServers\": {\n    \"duplication\": {\n      \"command\": \"node\",\n      \"args\": [\"/path/to/duplication-mcp/dist/index.js\", \".\"]\n    }\n  }\n}\n```\n\nOptionally scope it:\n\n```\n{\n  \"include\": [\"src/**\", \"lib/**\"],\n  \"exclude\": [\"**/*.generated.*\", \"**/migrations/**\"]\n}\n```\n\nThen ask your agent where the worst duplication is. If it tells you it is still indexing, that is the honest answer. Ask again in a few minutes.\n\nI'd love to hear what it finds in your codebase, and especially where it is wrong.", "url": "https://wpnews.pro/news/stop-counting-duplicated-lines-i-rank-copy-paste-by-what-it-costs-me", "canonical_source": "https://dev.to/jgauffin/stop-counting-duplicated-lines-i-rank-copy-paste-by-what-it-costs-me-3mp4", "published_at": "2026-09-29 21:33:55+00:00", "updated_at": "2026-09-29 21:46:45.260045+00:00", "lang": "en", "topics": ["ai-agents", "agent-protocols", "developer-tools", "ai-tools", "mlops"], "entities": ["dry-mcp", "SonarQube", "jina-embeddings-v2-base-code", "Jina AI", "MCP"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/stop-counting-duplicated-lines-i-rank-copy-paste-by-what-it-costs-me", "markdown": "https://wpnews.pro/news/stop-counting-duplicated-lines-i-rank-copy-paste-by-what-it-costs-me.md", "text": "https://wpnews.pro/news/stop-counting-duplicated-lines-i-rank-copy-paste-by-what-it-costs-me.txt", "jsonld": "https://wpnews.pro/news/stop-counting-duplicated-lines-i-rank-copy-paste-by-what-it-costs-me.jsonld"}}