{"slug": "conditional-reliability-of-toxicity-signals-for-multilingual-and-code-mixed", "title": "Conditional Reliability of Toxicity Signals for Multilingual and Code-Mixed Abuse Detection", "summary": "A new study from arXiv finds that external toxicity tools are unreliable for code-mixed and multilingual abuse detection, proposing ToxGate, a trust-fusion head that conditions auxiliary signals on encoder representations. Across 12 in-domain and 8 transfer settings, ToxGate improves over plain encoders in 10 and 7 settings respectively, with largest gains in high-risk moderation slices such as explicit slurs and violent threats.", "body_md": "arXiv:2607.15861v1 Announce Type: new\nAbstract: Moderation systems increasingly rely on external toxicity tools, but those tools are unreliable under code-mixing, transliteration, slang, and language mismatch. We study the \\emph{conditional reliability} of toxicity priors in Indian multilingual and code-mixed short text: English toxicity, Indic abuse, and rule-based severity cues can be useful evidence, but only in some linguistic and abuse-severity contexts. We propose ToxGate, a trust-fusion head that conditions each auxiliary signal on the encoder representation before adding it to the prediction state. Across three short-text abuse datasets, four transformer encoders, and five seeds per setting, ToxGate improves over matched plain encoders in 10 of 12 in-domain settings and 7 of 8 transfer settings. The largest and most interpretable gains occur in high-risk moderation slices, including explicit slurs, violent threats, and cross-dataset transfer. The broader lesson is that moderation systems should treat external toxicity tools and priors as conditional evidence rather than fixed features or ground truth, in focused ablations, source-specific gating gives the strongest results in transfer, severe-abuse slices, and high-risk triage.", "url": "https://wpnews.pro/news/conditional-reliability-of-toxicity-signals-for-multilingual-and-code-mixed", "canonical_source": "https://arxiv.org/abs/2607.15861", "published_at": "2026-07-20 04:00:00+00:00", "updated_at": "2026-07-20 13:41:42.697327+00:00", "lang": "en", "topics": ["natural-language-processing", "ai-safety", "ai-ethics"], "entities": ["ToxGate"], "alternates": {"html": "https://wpnews.pro/news/conditional-reliability-of-toxicity-signals-for-multilingual-and-code-mixed", "markdown": "https://wpnews.pro/news/conditional-reliability-of-toxicity-signals-for-multilingual-and-code-mixed.md", "text": "https://wpnews.pro/news/conditional-reliability-of-toxicity-signals-for-multilingual-and-code-mixed.txt", "jsonld": "https://wpnews.pro/news/conditional-reliability-of-toxicity-signals-for-multilingual-and-code-mixed.jsonld"}}