My team is looking at client-side OCR for an internal tool where people attach screenshots and photos of printed notices. The first preprocessing snippet anyone pastes into that kind of pipeline is grayscale plus a fixed threshold at 128. It's one line of OpenCV, so before it got near our upload flow I spent an evening measuring what it does. It fails in two quiet ways. It deletes text lighter than the cutoff, and under uneven light a global threshold paints whole regions black. Neither one throws an error, the OCR call just returns fewer characters.
The samples were a web page I wrote myself (a made-up Chinese notice), screenshotted at 1x and 2x and cut into blocks, including an 11 px footer in #999 gray on #f5f5f5, plus two crops from a holiday-notice template I found online. For recognition I used ImgIng (https://imging.ai/) with its default Professional OCR tier, plus Fast OCR and Ultimate OCR on the same ten crops, on an Apple M4 Mac in an open-source Chromium 149 build. The control was tesseract.js 5 (default parameters, chi_sim+eng). Errors are character error rate. Recognition ran in a Worker on the machine with zero non-GET requests. After three months of legal going through our data flows during a compliance cleanup, that is the first thing I check. The first run downloads the model, a download and not an upload of anyone's screenshots.
The footer's darkest pixel is 153. With the cutoff at 128, or at 100, every pixel of that text lands on the background side and the crop turns blank white. All three tiers at both pixel ratios returned 0 characters, and so did tesseract.js. The progress panel said no reliable text was found and suggested rotating, turning off scan enhancement, or a higher tier.
At 160 the cutoff sits just above those darkest pixels, so only the center of each stroke survives. Professional OCR read the 1x crop as gibberish, 75.5% wrong. At 200 it was 3.1%, with Otsu (which picked 208) 5.1%, and the untouched original 0.0%. Across all ten crops Professional OCR went from 0.1% on originals to 25.8% at 128, while the three tiers on originals were less than one point apart. Colored backgrounds fail too. The template's white words sit on red tags, and at threshold 100 the tags turned largely into black slabs that cut into their white words. Professional OCR scored 15.5%, and the "返岗时间" (back-to-work date) tag vanished.
I don't have real phone photos yet, so I generated two fake ones with a script: an A4 notice warped about 2.5 degrees, lit brighter top-left than bottom-right, blurred and noised, and one with an extra shadow over the text. These show a mechanism on program-generated samples, not real-photo error rates. On the shadowed one Otsu picked 137 and the shadow went solid black. Professional OCR scored 31.0%, and the order-independent error was also 31.0%, so those characters were really gone. Local adaptive thresholding scored 19.0% with an order-independent error of 0.0%. Every character was there, and the rest came from the tilt reordering short lines. Straightening the page from its corners got all three tiers to 0.0%.
So before a threshold I measure two things: how dark the ink really is, and how much paper brightness drifts across the image.
import cv2 as cv
import numpy as np
def threshold_audit(file, cutoff=128, trim=0.15):
g = cv.imread(file, cv.IMREAD_GRAYSCALE)
h, w = g.shape
core = g[int(h * trim):int(h * (1 - trim)), int(w * trim):int(w * (1 - trim))]
paper = cv.dilate(core, np.ones((41, 41), np.uint8)) # local "paper brightness"
otsu, _ = cv.threshold(g, 0, 255, cv.THRESH_BINARY + cv.THRESH_OTSU)
report = {"ink": int(core.min()), "paper_low": int(np.percentile(paper, 1)), "otsu": int(otsu)}
if report["ink"] >= cutoff:
report["verdict"] = f"fixed {cutoff} erases the text"
elif report["paper_low"] < otsu:
report["verdict"] = "global threshold blackens paper, go adaptive or skip"
return report
The trim drops the desk border around the generated photos. On the gray footer it reports ink 153 and flags 128. On the shadowed sample the darkest paper is 100 against an Otsu of 137, so it flags the global threshold, while the unshadowed one reads 159 against 135 and passes. Two cases get through. A cutoff of 160 sits above the ink, so the stroke-core failure isn't flagged, and the red tags pass because the dilate picks up their white words. I haven't found a cheap check for that one yet.
ImgIng's own "Scan enhancement" toggle is described as "Gentle grayscale, contrast and sharpening; no forced binarization." It's off by default, and turning it on gave 0.2% against 0.1%, so I'd call it neutral. Not binarizing lines up with everything above. Our default will be to send the original and only fix geometry. Before you ship a threshold, print the darkest pixel of your lightest text and the spread of your paper brightness.