cd /news/natural-language-processing/conditional-reliability-of-toxicity-… · home topics natural-language-processing article
[ARTICLE · art-65398] src=arxiv.org ↗ pub= topic=natural-language-processing verified=true sentiment=· neutral

Conditional Reliability of Toxicity Signals for Multilingual and Code-Mixed Abuse Detection

A new study from arXiv finds that external toxicity tools are unreliable for code-mixed and multilingual abuse detection, proposing ToxGate, a trust-fusion head that conditions auxiliary signals on encoder representations. Across 12 in-domain and 8 transfer settings, ToxGate improves over plain encoders in 10 and 7 settings respectively, with largest gains in high-risk moderation slices such as explicit slurs and violent threats.

read1 min views2 publishedJul 20, 2026

arXiv:2607.15861v1 Announce Type: new Abstract: Moderation systems increasingly rely on external toxicity tools, but those tools are unreliable under code-mixing, transliteration, slang, and language mismatch. We study the \emph{conditional reliability} of toxicity priors in Indian multilingual and code-mixed short text: English toxicity, Indic abuse, and rule-based severity cues can be useful evidence, but only in some linguistic and abuse-severity contexts. We propose ToxGate, a trust-fusion head that conditions each auxiliary signal on the encoder representation before adding it to the prediction state. Across three short-text abuse datasets, four transformer encoders, and five seeds per setting, ToxGate improves over matched plain encoders in 10 of 12 in-domain settings and 7 of 8 transfer settings. The largest and most interpretable gains occur in high-risk moderation slices, including explicit slurs, violent threats, and cross-dataset transfer. The broader lesson is that moderation systems should treat external toxicity tools and priors as conditional evidence rather than fixed features or ground truth, in focused ablations, source-specific gating gives the strongest results in transfer, severe-abuse slices, and high-risk triage.

── more in #natural-language-processing 4 stories · sorted by recency
── more on @toxgate 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/conditional-reliabil…] indexed:0 read:1min 2026-07-20 ·