21:14
2026-07-15
lesswrong.com
large-language-models
LLM CoTs remain monitorable when being unfaithful requires computation
A replication study by Arav Dhoot, supervised by Yixiong Hao and Zephaniah Roe, confirms that LLM chain-of-thought (CoT) unfaithfulness occurs mostly on easy tasks and that complex hints requiring com…