22:12
2026-07-26
lesswrong.com
ai-safety
What Happens When a Collusion Probe Only Finds a Thin Signal?
A BASE fellowship project called SPEC-GAP found that linear probes can detect adversarial shifts in multi-agent language models before unsafe actions become apparent in outputs, but the signal is thinβ¦