{"slug": "team-msu-gentext-forensics-challenge-2026-technical-report", "title": "Team MSU GenText-Forensics Challenge 2026 Technical Report", "summary": "Team MSU placed third in the ACM MM 2026 GenText-Forensics challenge with a decomposed chain-of-thought pipeline that pairs a document tampering detector with two Qwen3-VL-32B vision-language models, each LoRA-adapted to a distinct sub-task. The detector produces tampering probability maps converted into numbered candidate regions; a Filterer model validates them and assigns a preliminary forgery type, while a Semantic Detective model merges and re-grounds surviving regions, hunts semantic anomalies invisible to pixel-level detectors, and writes the final forensic report. Both models were trained by distilling chain-of-thought traces from a privileged Qwen3-VL-235B teacher with access to ground-truth masks and reports, and the team reports ablations over detector thresholds, prompt designs, and pipeline decompositions.", "body_md": "arXiv:2609.38391v1 Announce Type: new \nAbstract: Document text forgery has evolved beyond simple pixel-level manipulation: modern attacks alter not only the appearance of a document but also its meaning, and increasingly target the OCR & LLM pipelines that consume such documents. The ACM MM 2026 GenText-Forensics challenge therefore requires systems that not only decide whether a multilingual text image is forged, but also localize the point of manipulation, identify the attack type, and produce a human-readable forensic report with supporting evidence. We present our solution, a decomposed chain-of-thought (CoT) pipeline that combines a document tampering detector (DTD) with two Qwen3-VL-32B vision-language models, each LoRA-adapted to a distinct sub-task. DTD produces tampering probability maps that are converted into numbered candidate regions; a first model (the Filterer) validates these regions and assigns a preliminary forgery type, while a second model (the Semantic Detective) merges and re-grounds the surviving regions, searches for purely semantic anomalies that are invisible to pixel-level detectors, and writes the final report. Both models are trained by distilling chain-of-thought traces from a privileged Qwen3-VL-235B teacher that has access to ground-truth masks and reports. Our approach secured third place in the ACM MM 2026 GenText-Forensics challenge. We describe the data preparation, test-time augmentation, region rendering, distillation protocol, and training configuration in detail, and report ablations over detector thresholds, prompt designs, and pipeline decompositions.", "url": "https://wpnews.pro/news/team-msu-gentext-forensics-challenge-2026-technical-report", "canonical_source": "https://arxiv.org/abs/2609.38391", "published_at": "2026-10-01 04:00:00+00:00", "updated_at": "2026-10-01 04:20:06.393198+00:00", "lang": "en", "topics": ["artificial-intelligence", "computer-vision", "large-language-models", "ai-research", "natural-language-processing"], "entities": ["Team MSU", "ACM MM 2026 GenText-Forensics challenge", "Qwen3-VL-32B", "Qwen3-VL-235B", "LoRA", "arXiv"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/team-msu-gentext-forensics-challenge-2026-technical-report", "markdown": "https://wpnews.pro/news/team-msu-gentext-forensics-challenge-2026-technical-report.md", "text": "https://wpnews.pro/news/team-msu-gentext-forensics-challenge-2026-technical-report.txt", "jsonld": "https://wpnews.pro/news/team-msu-gentext-forensics-challenge-2026-technical-report.jsonld"}}