{"slug": "reliability-aware-sexism-detection-combining-dpo-with-annotator-agreement-and", "title": "Reliability-Aware Sexism Detection: Combining DPO with Annotator Agreement and Token-Level Confidence Scoring", "summary": "Researchers propose RA-DPO (Reliability-Aware Direct Preference Optimization), which integrates annotator agreement, model confidence, and token-level uncertainty to improve sexism detection. On 6,920 multilingual posts from EXIST 2023, fine-tuning OpenAI's gpt-4o with RA-DPO achieved 96.2% accuracy at 50% coverage in the true-agreement setting and 88.7% in the predicted-agreement setting, exceeding the 85.3% no-agreement baseline. Training on the top 30% most reliable pairs matched full-data DPO performance, reducing training cost without sacrificing accuracy.", "body_md": "arXiv:2608.12330v1 Announce Type: new\nAbstract: The detection of online sexism remains an open problem. Sexism detection is inherently subjective, yet most existing systems reduce multi-annotator labels to a single majority decision and treat all instances uniformly. This ignores two informative signals: annotator agreement and model uncertainty. We propose RA-DPO (Reliability-Aware Direct Preference Optimization), which integrates annotator agreement, model confidence, and a token-level uncertainty signal into a single reliability score. RA-DPO uses this score to select high-value preference pairs during training and to support inference-time abstention, which allows the model to trade coverage for accuracy. We evaluate RA-DPO on 6,920 multilingual posts from EXIST 2023, fine-tune OpenAI gpt-4o base via DPO, and validate on two open-weight 3B models (Llama, Qwen). Results show that training on the top 30% most reliable pairs matches full-data DPO, which indicates that reliability-aware selection can reduce training cost without sacrificing performance. At inference, selective prediction reaches 96.2% accuracy at 50% coverage in the true-agreement setting and 88.7% in the deployable predicted-agreement setting, both exceeding the 85.3% no-agreement baseline. These results suggest that accounting for annotation uncertainty is beneficial for both efficient training and reliable deployment in subjective classification.", "url": "https://wpnews.pro/news/reliability-aware-sexism-detection-combining-dpo-with-annotator-agreement-and", "canonical_source": "https://arxiv.org/abs/2608.12330", "published_at": "2026-08-14 04:00:00+00:00", "updated_at": "2026-08-14 04:12:34.584526+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "natural-language-processing"], "entities": ["RA-DPO", "OpenAI", "gpt-4o", "EXIST 2023", "Llama", "Qwen"], "alternates": {"html": "https://wpnews.pro/news/reliability-aware-sexism-detection-combining-dpo-with-annotator-agreement-and", "markdown": "https://wpnews.pro/news/reliability-aware-sexism-detection-combining-dpo-with-annotator-agreement-and.md", "text": "https://wpnews.pro/news/reliability-aware-sexism-detection-combining-dpo-with-annotator-agreement-and.txt", "jsonld": "https://wpnews.pro/news/reliability-aware-sexism-detection-combining-dpo-with-annotator-agreement-and.jsonld"}}