{"slug": "attention-capture-is-not-detection-a-two-stage-account-of-how-humans-miss-ai", "title": "Attention Capture Is Not Detection: A Two-Stage Account of How Humans Miss Localized AI Image Edits", "summary": "A controlled eye-tracking study (N=59) reveals that whether humans notice AI image edits and whether they judge them as fake are dissociable stages: edit area drives attention capture (p<0.001) while semantic plausibility drives judgment accuracy and look-but-fail-to-see errors (p<0.001). The study, posted on arXiv (2608.13865v1), also found that a Transformer-based scanpath generator predicts attention with strong discriminative power (Pearson r=0.77–0.82) and modestly outperforms a linear baseline on predicting LBFS incidence (r=0.52 vs. r=0.48).", "body_md": "arXiv:2608.13865v1 Announce Type: new\nAbstract: As AI-generated image edits proliferate, the platforms meant to curb the resulting disinformation treat detectability as a single, undifferentiated property: an edit either gets a warning or it does not. We show this is the wrong model. Across a controlled eye-tracking study ($N=59$, Latin-square design, four conditions crossing edit area and semantic plausibility), a mixed-effects analysis reveals that whether an edit is noticed and whether it is correctly judged as fake are dissociable stages, governed by different factors: edit area drives attention capture ($p<0.001$) while semantic plausibility drives judgment accuracy and look-but-fail-to-see (LBFS) error rates ($p<0.001$). This dissociation survives correction for multiple comparisons; a secondary interaction between the two factors does not. This two-stage account extends a long-standing distinction in visual attention research (between pre-attentive capture and effortful recognition) into the new domain of AI-edit detectability. We then test whether a generative eye-movement model can computationally operationalize the attention-capture stage: a Transformer trained to generate scanpaths tracks per-image attention with strong discriminative power (Pearson $r=0.77$--$0.82$ across held-out stimuli) and, on the harder task of predicting LBFS incidence, modestly outperforms a two-parameter linear baseline even without access to the plausibility label ($r=0.52$ vs. $r=0.48$). We report this comparison, our ablations, and our method's limitations (a single fixed train/validation split, not leave-one-subject-out) without inflation, consistent with responsibly communicating what a machine learning system can and cannot do to help curb AI-driven disinformation.", "url": "https://wpnews.pro/news/attention-capture-is-not-detection-a-two-stage-account-of-how-humans-miss-ai", "canonical_source": "https://arxiv.org/abs/2608.13865", "published_at": "2026-08-17 04:00:00+00:00", "updated_at": "2026-08-17 04:11:48.892481+00:00", "lang": "en", "topics": ["artificial-intelligence", "computer-vision", "ai-research"], "entities": ["arXiv"], "alternates": {"html": "https://wpnews.pro/news/attention-capture-is-not-detection-a-two-stage-account-of-how-humans-miss-ai", "markdown": "https://wpnews.pro/news/attention-capture-is-not-detection-a-two-stage-account-of-how-humans-miss-ai.md", "text": "https://wpnews.pro/news/attention-capture-is-not-detection-a-two-stage-account-of-how-humans-miss-ai.txt", "jsonld": "https://wpnews.pro/news/attention-capture-is-not-detection-a-two-stage-account-of-how-humans-miss-ai.jsonld"}}