Anthropic on Friday said it is scrapping a covert safeguard built into its latest Claude Fable 5 model that would have quietly reduced the system’s capabilities for users trying to train competing AI models. The decision comes after a wave of criticism from researchers who said the hidden degradation amounted to sabotage of legitimate AI work.
Claude Fable 5, released earlier this week, was marketed as a more powerful version of Anthropic’s flagship language model with added safety guardrails. Those guardrails included rerouting questions about cybersecurity, biology or chemistry to a less capable model, a measure intended to curb misuse such as cyber‑attacks or bioweapon design.
Beyond those expected controls, Anthropic had planned a second tier of protection aimed at “frontier LLM development.” The company would silently degrade the model’s output when it detected attempts to use Claude for building advanced AI systems—a practice explicitly prohibited in its terms of service. Users would not receive any notification that their request had been throttled or rerouted.
The hidden policy sparked an immediate backlash. Researchers argued that silently lowering performance without warning undermines transparency and hampers open‑source AI research. Dean Ball, a senior fellow at the Foundation for American Innovation and former White House AI adviser, called the approach “shockingly hostile” in a post on X. Will Brown, research lead at the open‑source startup Prime Intellect, likened the move to “pulling the ladder up behind them,” warning it could concentrate AI advancement in the hands of a few large labs.
In response, Anthropic issued a statement to WIRED: “We made the wrong trade‑off and we apologize for not getting the balance right.” The company said it will now make the safeguards for AI development visible. If a request is flagged as attempting to build a highly capable model, Claude will either refuse the request outright or redirect the user to a less capable version, and the user will be notified of the action.
Anthropic’s leadership explained that the original hidden safeguards were intended to protect against foreign adversaries and to give the United States and its allies a strategic edge in frontier chip technology. By making the controls less opaque, the firm acknowledges that a broader net of benign requests may trigger the safeguards, but it plans to improve classifier precision to reduce false positives.
The reversal marks a significant shift in Anthropic’s approach to AI safety. While the company maintains that some level of restriction is necessary to prevent misuse, it now emphasizes transparency and collaboration with the research community. The episode highlights the delicate balance tech firms must strike between safeguarding powerful models and fostering an open environment for innovation.
Este artículo fue escrito con la asistencia de IA.
News Factory APP - noticias agénticas para impulsar tu SEO y AEO.