Fuzzy Duplicate Finder for Spreadsheets CleanMySheet released a fuzzy duplicate finder that groups near-duplicate spreadsheet rows using normalized Levenshtein similarity at a default threshold of about 85%, blocking on a short key prefix to keep large files usable. The tool parses CSV and Excel files on the user's device, surfaces matches in a review drawer, and removes nothing until a user marks a pair as a duplicate, with an undo option available from the recipe list. CleanMySheet states the similarity measure is a string distance rather than an LLM guess. CleanMySheet guide Find near-duplicate rows fuzzy matching CleanMySheet groups similar rows using normalized Levenshtein similarity default ~85% after blocking on a short key prefix so large files stay usable. Matches appear in a review drawer; nothing is removed until you mark a pair as a duplicate. Near-duplicates wait for your review. Drop CSV or Excel here Tap to choose a file. Parsing stays on this device. Parsing stays on this device. Optional: drop a .cleanmysheet.json recipe with it. Open the cleaner /clean?op=fuzzy Steps 1. Run a scan after upload 2. Add Fuzzy duplicate detection from search or Recommended 3. Open Review fuzzy pairs 4. Mark Duplicate or Keep both 5. Undo anytime Not an LLM guess Similarity is a string distance, not a model hallucination. You can undo the whole step from the recipe list. Related searches this page answers - fuzzy duplicate finder - near duplicate rows - similar contacts csv