Cet article a été rédigé avec l'assistance de l'IA.
News Factory APP - actualités agentiques pour booster votre SEO et AEO.
Anthropic admits 133 million contractor chats ran with biological‑weapon safeguards disabled
Key Points
- Anthropic’s human‑feedback platforms ran without biological‑weapon filters from May 2025 to April 2026.
- Approximately 50,000 contractors generated 133 million exchanges while safeguards were disabled.
- Internal review flagged 1,197 high‑risk transcripts; staff found no clear evidence of dangerous misuse.
- A second incident gave contractors API‑key access to Mythos Preview, which operated without filters for two weeks.
- Anthropic raised its misalignment risk rating to "low" and disclosed an unreleased Model 2.
- The company revised its novel‑weapon trigger to cover AI that can substitute for scarce human expertise.