How Value Induction Reshapes LLM Behaviour
Apple researchers Arnav Arora, Natalie Schluter, Katherine Metcalf and Maartje ter Hoeve found that fine-tuning conversational large language models on curated value subsets of existing preference datasets causes uninten…