Anthropic disclosed a major safety lapse in its August 14 Risk Report, noting that its human‑feedback platforms operated without the biological‑weapon blocking classifiers for eleven months. The company first deployed the safeguards in May 2025, but from that point until April 2026 a flag disabled both the blocking behavior and the logging of potentially hazardous content.

During the gap, about 50,000 external contractors accessed the platforms and exchanged roughly 133 million prompts and responses. Anthropic relied on outside vendors to vet these users, but many vendors lacked screening processes capable of stopping even the most basic (CB‑1) biological‑weapon threat actors. Most contractors could hold open‑ended conversations rather than being limited to fixed answer sets.

Anthropic’s internal review ran Claude Sonnet 5 over every human turn in the affected period, flagging 1,197 transcripts as high risk. Of those, 757 originated from Anthropic’s own teams on the same infrastructure; the remaining 440 came from deliberate red‑teaming exercises. Staff examined all 62 non‑Anthropic high‑risk transcripts and a random sample of 30 others. They found no clearly concerning misuse, though they identified a handful of potentially dual‑use conversations. The company said the short length of most exchanges makes a real‑world risk unlikely, but acknowledged the gap raises the likelihood of similar, undiscovered issues.

Second breach involving Mythos Preview

In April 2026 an external tip alerted Anthropic that a few contractors at a data‑labeling vendor had exploited a flaw to obtain an API key. The key allowed them to run models outside their assigned tasks, including Mythos Preview, one of Anthropic’s most capable systems. The unauthorized access path remained open for several weeks, and Mythos Preview operated without the biological classifiers for roughly two weeks. Anthropic shut the vector within 90 minutes of learning of the breach and closed it the same day. No model weights, customer data, or core network assets were compromised.

The report also updates Anthropic’s misalignment risk rating. The company raised its estimate of catastrophic harm from misalignment in high‑stakes settings from “very low” in February to “low” in the August report, citing recent incident disclosures in cybersecurity evaluations rather than new evidence of model behavior. Anthropic clarified that the change reflects heightened uncertainty.

Among the disclosures, Anthropic announced an unreleased internal model dubbed Model 2. The model scores about 1.5 points higher on the company’s capability index than its previous flagship, Mythos 5, but Anthropic has no plans to ship it. The company also revised its novel‑weapon trigger, now covering AI that can “functionally substitute” for scarce human expertise rather than merely “significantly help” threat actors.

Anthropic’s internal audit was partly reviewed by its own Claude model, which flagged three areas where the public record was less complete than the internal assessment. The company said the criticisms were fair and that external audits have not yet been requested, despite the Long‑Term Benefit Trust’s authority to demand them.

Este artigo foi escrito com a assistência de IA.
News Factory APP - notícias agênticas para impulsionar seu SEO e AEO.