Anthropic says biological-safety filters were disabled on 133 million contractor exchanges Anthropic disclosed that an internal configuration error disabled biological-safety blocking classifiers on all human-feedback vendor traffic from May 2025 through April 2026, affecting roughly 50,000 contractors and 133 million exchanges. The company found no evidence of meaningful biological-weapons misuse but raised its CB-1 risk assessment from 'very low' to 'low' and strengthened vendor screening and access controls. Anthropic says biological-safety filters were disabled on 133 million contractor exchanges - Anthropic says all human-feedback vendor traffic ran without biological blocking classifiers from May 2025 through April 2026, covering roughly 50,000 people and 133 million exchanges. 1 https://www-cdn.anthropic.com/f61d49fa5596956a5dec75fea0e973bf6a6a8378/Redacted%20Risk%20Report%20August%202026%20.pdf - The configuration error disabled both the classifiers’ blocking function and the logging that would have sent alerts for review. 1 https://www-cdn.anthropic.com/f61d49fa5596956a5dec75fea0e973bf6a6a8378/Redacted%20Risk%20Report%20August%202026%20.pdf - Contractors had mostly open-ended access to models, including some unreleased systems, but Anthropic says it found no evidence of meaningful biological-weapons misuse. 1 https://www-cdn.anthropic.com/f61d49fa5596956a5dec75fea0e973bf6a6a8378/Redacted%20Risk%20Report%20August%202026%20.pdf - The company has strengthened vendor screening and access controls and raised its CB-1 risk assessment from “very low” to “low.” 1 https://www-cdn.anthropic.com/f61d49fa5596956a5dec75fea0e973bf6a6a8378/Redacted%20Risk%20Report%20August%202026%20.pdf Anthropic disclosed Sunday that an internal configuration error left biological-risk filters disabled across all traffic from human-feedback vendors for nearly a year, allowing about 50,000 contractors to conduct roughly 133 million exchanges without the safeguards the company says it uses to block dangerous chemical- and biological-weapons assistance. 1 https://www-cdn.anthropic.com/f61d49fa5596956a5dec75fea0e973bf6a6a8378/Redacted%20Risk%20Report%20August%202026%20.pdf The lapse ran from May 2025, when Anthropic began deploying models with chemical-and-biological safeguards, through April 2026. The company said its investigation found no evidence of concerning misuse that could have provided meaningful uplift to a biological-weapons threat actor, while acknowledging that the discovery reduced its confidence that similar gaps do not exist elsewhere. 1 https://www-cdn.anthropic.com/f61d49fa5596956a5dec75fea0e973bf6a6a8378/Redacted%20Risk%20Report%20August%202026%20.pdf What failed Anthropic’s biological safeguards are real-time blocking classifiers. They screen interactions for content associated with its CB-1 threat model, which covers assistance that could significantly help individuals with basic technical backgrounds create, obtain or deploy biological or chemical weapons. The classifiers are intended to block risky requests and generate signals for monitoring and review. 1 https://www-cdn.anthropic.com/f61d49fa5596956a5dec75fea0e973bf6a6a8378/Redacted%20Risk%20Report%20August%202026%20.pdf In the contractor environment, an internal-use flag disabled both functions. It prevented the classifiers from blocking responses and stopped them from recording flags or passing them to review systems. Anthropic’s report says the gap affected all human-feedback vendor traffic, not customer traffic. 1 https://www-cdn.anthropic.com/f61d49fa5596956a5dec75fea0e973bf6a6a8378/Redacted%20Risk%20Report%20August%202026%20.pdf The public report does not provide a complete model-by-model inventory of the affected exchanges. It says contractors evaluated Anthropic models through outside data-collection platforms, including some early or unreleased systems. The report separately describes an incident in which contractors obtained unauthorized access to Claude Mythos Preview through a vendor platform; Anthropic does not present that access-key incident as the cause of the broader classifier outage. 1 https://www-cdn.anthropic.com/f61d49fa5596956a5dec75fea0e973bf6a6a8378/Redacted%20Risk%20Report%20August%202026%20.pdf Why contractors had access The contractors were engaged through outside data vendors for human-feedback collection and model testing used in training and evaluation. Anthropic says much of the work involved open-ended interaction with models, rather than simply rating fixed responses. 1 https://www-cdn.anthropic.com/f61d49fa5596956a5dec75fea0e973bf6a6a8378/Redacted%20Risk%20Report%20August%202026%20.pdf During the affected period, vendors were responsible for screening individual workers. Anthropic said that earlier screening requirements were weaker than its current standards and that some vendors lacked processes capable of excluding even CB-1-level threat actors. The company says it has since required stronger identity and background checks, endpoint controls and phishing-resistant multifactor authentication. 1 https://www-cdn.anthropic.com/f61d49fa5596956a5dec75fea0e973bf6a6a8378/Redacted%20Risk%20Report%20August%202026%20.pdf The exposure therefore involved external workers who could interact with models relevant to Anthropic’s biological-risk evaluations. It did not involve Anthropic customers, according to the company’s executive summary, which says there was no impact on customer traffic. 1 https://www-cdn.anthropic.com/f61d49fa5596956a5dec75fea0e973bf6a6a8378/Redacted%20Risk%20Report%20August%202026%20.pdf What Anthropic found Anthropic reviewed historical traffic using Claude Sonnet 5 to identify conversations associated with harmful biological use under the CB-1 standard. The review flagged 1,197 transcripts as high for biological harm. Anthropic said 757 came from internal teams using the same infrastructure, while most of the remaining cases came from deliberate red-team exercises; 62 were classified as non-red-team external cases. 1 https://www-cdn.anthropic.com/f61d49fa5596956a5dec75fea0e973bf6a6a8378/Redacted%20Risk%20Report%20August%202026%20.pdf The company manually reviewed the 62 non-red-team transcripts and a sample of 30 external red-team transcripts. It said the review found no clearly concerning misuse capable of providing meaningful uplift to a threat actor, although some conversations involved dual-use or academic-level biology. 1 https://www-cdn.anthropic.com/f61d49fa5596956a5dec75fea0e973bf6a6a8378/Redacted%20Risk%20Report%20August%202026%20.pdf Anthropic said it corrected the configuration after discovering the problem in April 2026. Most vendor traffic now runs with biological classifiers enabled, while exemptions require stronger vendor controls and project-specific approval. The public report describes the review and remediation but does not say whether individual contractors, customers, regulators or law-enforcement agencies were separately notified. 1 https://www-cdn.anthropic.com/f61d49fa5596956a5dec75fea0e973bf6a6a8378/Redacted%20Risk%20Report%20August%202026%20.pdf The incident matters because the classifiers are a deployment safeguard rather than a capability limitation built into the models. Anthropic’s own biological-risk research says its guards monitor inputs and outputs in real time and block a narrow class of potentially harmful information. When the guard layer was disabled, contractors could interact with the underlying systems without that additional control. 3 https://www.anthropic.com/research/biorisk Independent assessment SecureBio’s July review of an earlier Anthropic chemical-and-biological risk report found that Anthropic’s refusal classifiers blocked 94.2% of hazardous prompts in SecureBio’s BioTIER-refuse benchmark. SecureBio also warned that highly capable jailbreakers could still make substantial progress around the filters and that remediation could take time. 2 https://securebio.substack.com/p/review-of-anthropics-unredacted-chemical That review covered Claude Opus 4.6 and did not extend its conclusions to Fable 5, Mythos 5 or Opus 5. SecureBio agreed with Anthropic’s earlier overall risk assessment but identified disagreements about high-skill jailbreak risk and recommended continued updates to classifier constitutions and monitoring of users granted exemptions. 2 https://securebio.substack.com/p/review-of-anthropics-unredacted-chemical Anthropic’s August report now rates CB-1 risk as low rather than very low, specifically citing the access-control gap. It continues to rate novel chemical-and-biological-weapons risk as low with substantial uncertainty. 1 https://www-cdn.anthropic.com/f61d49fa5596956a5dec75fea0e973bf6a6a8378/Redacted%20Risk%20Report%20August%202026%20.pdf Companies mentioned Further sources The stories that matter, in one email. Free — unsubscribe anytime.