Almost every OCR tip I found said the same thing. Before recognizing a screenshot, clean it up: go grayscale, binarize it, run a denoise filter, maybe upscale it 2x. I had never checked whether any of that actually helps, so I took one small set of screenshots and ran every trick on the list, one at a time, against text I already knew. On clean screenshots, not a single step beat the untouched original. Several of them made things a lot worse.
The samples were boring on purpose. I wrote a small web page, a made-up Chinese notice about community service hours, and screenshotted parts of it at 1x and 2x device pixel ratio: 16 px body text, 12 px sidebar text, 11 px light-gray footer text (#999 on #f5f5f5), and a 13 px fee table. I added two crops from a holiday notice template I found online, white text on red. That gave me 10 crops with an exact answer key. For recognition I used ImgIng (https://imging.ai/), with its default Professional OCR tier plus Fast OCR and Ultimate OCR on every version, in a Chromium 149 open-source build on an M4 Mac. As a reference engine I ran tesseract.js 5 with default parameters and chi_sim+eng. Recognition runs in a local Worker, and across roughly 800 runs I saw zero non-GET requests. The first run of each tier does download a model.
The score is character error rate: edit distance divided by the length of the answer, summed over all 10 crops. Here is every version side by side.
Look for the short bars first. The original scored 0.1% on Professional OCR. The only versions that matched it were grayscale at 0.1% and the built-in "Scan enhancement" switch at 0.2%. That switch describes itself as "Gentle grayscale, contrast and sharpening; no forced binarization." It didn't lower the error rate on this set, and that seems fine to me, because the steps it skips are the ones that cost the most. Everything else added errors. The small losses were a 2x upscale (0.6%), JPEG re-saves at quality 60 and 30 (0.8% and 1.1%), Otsu binarization (1.1%), a fixed threshold of 200 (1.2%), adaptive thresholding (1.4%) and non-local-means denoising (2.7%). The big losses were a 3×3 median filter at 9.7%, threshold 160 at 9.4%, threshold 128 at 25.8% and threshold 100 at 32.2%.
Threshold 128 looks like a sensible default, so I went back to see what happened. The gray footer text is simply lighter than the threshold. Its darkest pixel is 153, which sits above 128, so binarizing turns the whole strip white. All three tiers and tesseract.js returned 0 characters from it. So for that strip I got an empty result rather than a handful of typos.
The median filter fails in a quieter way. On 12 px text captured at 1x, the strokes merge into blobs. My guess is that a 3×3 window is about as wide as a stroke at that size, but I only looked at the output. That single crop went from 0% to 32.4% on Professional OCR.
I also expected the tiers to matter more than they did. On the originals, Fast, Professional and Ultimate landed at 0.8%, 0.1% and 0.7%. They are less than one point apart. Threshold 128 moved every tier by about 25 points. Picking the wrong filter hurt far more than picking a different tier. Preprocessing didn't clearly help tesseract.js either. Its original scored 6.2%, and its best version was the 2x upscale at 5.8%.
Vertical text needs its own paragraph, since every tier scored a scary 92.3% on my four-column sample. The characters were fine, though. Counted without regard to order, Professional OCR got 0.0% to 7.7% wrong. The columns came out left to right, while this kind of vertical text reads right to left. None of the binarizing, denoising or upscaling variants fixed that, and they all stayed between 89.7% and 94.9%. What fixed it was the rotate-left-90° button in the image adjustments. After rotating the 2x screenshot, all three tiers read it with 0 errors. On the smaller 1x capture, Fast and Professional still dropped two punctuation marks, which works out to 5.1%.
This covers clean screenshots only. I ran each combination once, so a 1-point wobble means nothing, and I haven't tested real phone photos yet. Still, my routine changed. I recognize the untouched file first and keep that result as a baseline. If I really want to binarize, I measure the darkest gray in the text before picking a threshold, and I compare the output against the baseline instead of assuming it improved. If the text runs vertically, I rotate it before I touch anything else.