23:59
2026-07-27
lesswrong.com
ai-safety
Untrusted advice for AI control: Short, strong advice significantly uplifts weak LLMs
Researchers introduced the untrusted advice protocol, in which a trusted executor LLM takes all actions while an untrusted advisor LLM can only send short hints, recovering a substantial fraction of tโฆ