cd /news/artificial-intelligence/reliability-aware-sexism-detection-c… · home topics artificial-intelligence article
[ARTICLE · art-96287] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Reliability-Aware Sexism Detection: Combining DPO with Annotator Agreement and Token-Level Confidence Scoring

Researchers propose RA-DPO (Reliability-Aware Direct Preference Optimization), which integrates annotator agreement, model confidence, and token-level uncertainty to improve sexism detection. On 6,920 multilingual posts from EXIST 2023, fine-tuning OpenAI's gpt-4o with RA-DPO achieved 96.2% accuracy at 50% coverage in the true-agreement setting and 88.7% in the predicted-agreement setting, exceeding the 85.3% no-agreement baseline. Training on the top 30% most reliable pairs matched full-data DPO performance, reducing training cost without sacrificing accuracy.

read1 min views1 publishedAug 14, 2026

arXiv:2608.12330v1 Announce Type: new Abstract: The detection of online sexism remains an open problem. Sexism detection is inherently subjective, yet most existing systems reduce multi-annotator labels to a single majority decision and treat all instances uniformly. This ignores two informative signals: annotator agreement and model uncertainty. We propose RA-DPO (Reliability-Aware Direct Preference Optimization), which integrates annotator agreement, model confidence, and a token-level uncertainty signal into a single reliability score. RA-DPO uses this score to select high-value preference pairs during training and to support inference-time abstention, which allows the model to trade coverage for accuracy. We evaluate RA-DPO on 6,920 multilingual posts from EXIST 2023, fine-tune OpenAI gpt-4o base via DPO, and validate on two open-weight 3B models (Llama, Qwen). Results show that training on the top 30% most reliable pairs matches full-data DPO, which indicates that reliability-aware selection can reduce training cost without sacrificing performance. At inference, selective prediction reaches 96.2% accuracy at 50% coverage in the true-agreement setting and 88.7% in the deployable predicted-agreement setting, both exceeding the 85.3% no-agreement baseline. These results suggest that accounting for annotation uncertainty is beneficial for both efficient training and reliable deployment in subjective classification.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @ra-dpo 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/reliability-aware-se…] indexed:0 read:1min 2026-08-14 ·