Why AI Detection Fails for Academic Integrity A new arXiv study (2608.11256v1) finds that commercial AI detectors cannot reliably distinguish AI-assisted editing from full AI drafts, flagging guideline-compliant 'refine abstract only' edits at 64–80% rates (Pangram/GPTZero) while unmodified 2023–2025 originals are flagged at 9–15%, with non-STEM rates significantly higher (p<0.001). After Undetectable AI humanization, fewer than 4% of AI-labeled rewrites remain flagged (FNR >96%), meaning honest AI editing carries higher sanction risk than evasion, so detector scores should not be standalone misconduct evidence. arXiv:2608.11256v1 Announce Type: new Abstract: Institutions use commercial AI detectors for academic integrity, yet detectors cannot distinguish AI editing from full LLM drafts and may treat both as misconduct. In a controlled study of published English abstracts four domains; 2013 to 2015 vs. 2023 to 2025 , we quantify this policy failure under proxy human/AI labels at tau=0.50. Light "refine abstract only" edits, a proxy for guideline-compliant AI assistance, are flagged at 64 to 80% Pangram/GPTZero . Unmodified 2023 to 2025 originals are flagged at 9 to 15%, with non-STEM rates far above STEM p<0.001 ; elevated scores track long-token and Academic Word List density, not authorship intent alone. After Undetectable AI humanization, evasion is near-total: fewer than 4% of AI-labeled rewrites remain flagged post-humanization detection rate <4%; FNR 96% . Honest AI-editing results in a higher sanction risk than humanizer-assisted evasion. Therefore, detector scores should not serve as standalone misconduct evidence.