ARCF vs TCLA: Aligning Model Representations for Safety and Stability
Two September 2026 papers propose post-training alignment techniques that expose models to misaligned contexts and then enforce aligned targets: ARCF, which reduced unsafe generations on a benchmark o…