21:51
2026-07-22
lesswrong.com
ai-safety
In other words: The influence of prompt variation on alignment evals
A 5-day project sprint during ARENA 8.0 found that prompt rephrasing can measurably impact alignment-relevant properties even on frontier models, though data were noisy and inconsistent across evals, โฆ