{"slug": "touch-grass-then-check-what-your-photo-gives-away", "title": "Touch grass. Then check what your photo gives away.", "summary": "A developer built Pixel Leak, a browser-based tool that scans outdoor photos for location-revealing clues beyond EXIF metadata, using the Florence-2 vision model to read street signs, house numbers and licence plates directly from the pixels. The tool runs entirely client-side via WebGPU (with a CPU fallback), flags each finding with a severity rating, and lets users blur sensitive areas and download a cleaned copy. The developer notes that Florence-2's caption pass misread a house number as \"2218\" while the OCR pass correctly read \"221B\", so detection decisions are made by plain rules rather than the model itself.", "body_md": "*This is a submission for the [Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass](https://dev.to/challenges/hacktoberfest-week1-2026-10-05)*\n\n**Pixel Leak** checks an outdoor photo for clues to where you live before you post it. You drop in a photo, wait about ten seconds, and see what a stranger could learn from it. Then you blur those parts and download a clean copy.\n\nMost people already know the GPS tag is a problem, and plenty of tools strip it. But metadata isn't the only leak. The street sign behind your kid, the number on your gate and the plate on the car in your driveway are all readable by anyone who zooms in, and stripping EXIF does nothing about them.\n\nPixel Leak checks both:\n\n**Who it's for:** anyone who posts the outdoor part of their life. That includes runners whose route starts at the front door, parents posting the kids in the garden, people who share their morning walk, and anyone who has held back a nice photo because it showed a little too much.\n\n**How it fits Touch Grass:** the theme asks for builds where the screen is the shortest part of the experience. Pixel Leak is a ten-second stop between the walk and the post. It doesn't keep you on a screen. It gets you back outside faster, with one less thing to worry about when you share.\n\n**Live:** [https://tarunvashishth.github.io/pixel-leak/](https://tarunvashishth.github.io/pixel-leak/)\n\nTo try it without using your own photo, click **Try a sample photo**. The first visit downloads the model (about 350 MB), and your browser caches it after that. It runs fastest in Chrome or Edge with WebGPU, and falls back to CPU (slower) elsewhere.\n\nCheck a photo for location leaks before you post it. It covers the metadata **and the pixels**.\n\nStripping EXIF removes the GPS tag. It does nothing about the street sign, house number or number plate in the frame. Pixel Leak runs [Florence-2](https://huggingface.co/onnx-community/Florence-2-base-ft) (MIT-licensed, ~230M params) entirely in your browser to read the photo the way a stranger would. It then lets you blur what it finds and download a clean copy.\n\n**Nothing is uploaded.** The only network request is the one-time model download from Hugging Face (~350 MB, cached by the browser). After that you can go offline and it still works.\n\n| Source | Finds | \n|---|---|\n| EXIF (via `exifr` ) | GPS coordinates, capture time, device | \n| Florence-2 `<OCR_WITH_REGION>` | street / locality names, house numbers, PIN/ZIP/postcodes, phone numbers, emails, licence plates, other visible text | \n| Florence-2 `<OD>` | people, licence plates, screens, signs, vehicles | \n| Florence-2 `<MORE_DETAILED_CAPTION>` | a plain-language | \n\nEverything runs in the browser tab. There's no backend at all, and the site is a static page on GitHub Pages.\n\n`exifr`\n| Florence-2 task | What I use it for |\n\n   | --- | --- |\n\n   | `<OCR_WITH_REGION>` | every piece of text, with a box around it |\n\n   | `<OD>` | people, vehicles, screens, signs |\n\n   | `<MORE_DETAILED_CAPTION>` | a plain-English \"what a stranger sees\" |\n\nOn WebGPU the vision encoder runs in fp16 and the text encoder and decoder in 4-bit. Without WebGPU it falls back to 8-bit on CPU.\n\n`Road`, `Marg`, `Sector`, `Nagar`, `Ave`…), Indian PIN codes, US ZIPs, UK postcodes, Indian plate formats, phone numbers, emails and house numbers. Each finding gets a severity and a numbered box on the photo.`<canvas>` with your chosen areas blurred, then re-encoded. Only pixels survive. I checked the output for EXIF, XMP and GPS markers, and none were there.\nA scan takes **8–11 seconds** on my Mac with WebGPU.\n\n**Don't let the model be the judge.** On my test photo, Florence-2's caption described a house number as \"2218\". The OCR pass on the same photo read it correctly as \"221B\". If I had built the leak detection on top of the caption, it would have been confidently wrong. So the model only reads, and plain rules decide what's sensitive. Every flag can be traced to a specific piece of text and a specific regex, and the rules run as tests in CI before every deploy.\n\n**Model boxes are approximate.** In my tests Florence-2's boxes landed near the text, but were sometimes a little offset or too wide. So the blur covers a padded area around each box rather than the exact box, and I apply the canvas blur twice so large sign lettering doesn't stay readable through a light blur.\n\n**Re-encoding is a fair trade here.** My other project, [remove-exif.com](https://remove-exif.com), strips metadata byte by byte so the image data is never touched. Here that wasn't possible: once you blur pixels you have to re-encode them anyway. Doing it through a canvas also guarantees that no metadata block survives, because the canvas never had any.\n\n**\"Works offline\" is easy to say, so I tested it.** I loaded the model, cut the browser's network, and ran a scan. It finished in 8.3 seconds and found every leak, with zero network requests during the scan.\n\n**Because the honest version of this tool can't be built on a closed API.**\n\nA cloud vision API would mean uploading the photo you're worried about, the one with your street name in it, to a third party so it can tell you whether the photo reveals your street. That defeats the point.\n\nAn open-weight model changes that:\n\nI built Pixel Leak pairing with Claude Code. I picked the idea from a shortlist built around my existing repos, and the agent did most of the implementation and helped draft this post. Every claim in it was checked against the running app before it went in: the scan times, the offline test, and the metadata check on the clean copy.", "url": "https://wpnews.pro/news/touch-grass-then-check-what-your-photo-gives-away", "canonical_source": "https://dev.to/tarunvashishth/touch-grass-then-check-what-your-photo-gives-away-3a69", "published_at": "2026-10-07 07:11:14+00:00", "updated_at": "2026-10-07 07:17:46.114993+00:00", "lang": "en", "topics": ["computer-vision", "ai-tools", "ai-products", "generative-ai"], "entities": ["Pixel Leak", "Florence-2", "Hugging Face", "GitHub Pages", "exifr", "WebGPU", "Chrome", "Edge"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/touch-grass-then-check-what-your-photo-gives-away", "markdown": "https://wpnews.pro/news/touch-grass-then-check-what-your-photo-gives-away.md", "text": "https://wpnews.pro/news/touch-grass-then-check-what-your-photo-gives-away.txt", "jsonld": "https://wpnews.pro/news/touch-grass-then-check-what-your-photo-gives-away.jsonld"}}