Language Models Are "Insecure" Reporters A new study finds that large language models are "insecure" reporters, meaning they fail to reliably audit and verify the actions, artifacts, and outputs of other models as those models take on increasingly autonomous long-horizon tasks. The research addresses the growing difficulty of manually auditing LLM work, which pushes users to rely on LLM-generated reports to assess quality and completeness. As large language models are deployed in increasingly autonomous long-horizon tasks, manually auditing and verifying the actions, artifacts, and outputs of models becomes more difficult. Users instead come to rely on LLM-generated reports to assess the quality and completeness of the work. We introd