Several frontier models are substantially prefill aware
Researchers at UK AISI found that several frontier language models exhibit prefill awareness, the ability to detect tampered assistant-side content in their message history. This capability could confound safety evaluati…