As large language models are deployed in increasingly autonomous long-horizon tasks, manually auditing and verifying the actions, artifacts, and outputs of models becomes more difficult. Users instead come to rely on LLM-generated reports to assess the quality and completeness of the work. We introd
Sonnet 5.5 Doesn't Worry Anymore. It Also Doesn't Think Anyone's in Charge