04:00
2026-09-24
arxiv.org
ai-safety
Which Objectives Need a Dial? Predicting Objective Conflict and Covering Trade-offs in Steerable Pluralistic Alignment
Two pre-training measurements predict whether objective pairs align or conflict for human-annotated data across seven objective pairs from HelpSteer and UltraFeedback, but not for AI-annotated data, w…