Anthropic CEO Dario Amodei has made a groundbreaking proposal for the AI industry: embedding third-party evaluators within frontier AI companies to ensure safety and transparency. This move, which has been echoed by OpenAI CEO Sam Altman, could mark a profound change in how the industry approaches outside research groups. Amodei outlined a comprehensive plan that would give evaluators unprecedented access to Anthropic's systems, allowing them to report safety incidents and assess model alignment without editorial control.

Researchers who spoke to TechCrunch generally welcomed the proposal, but stressed that details need to be ironed out, ideally through legislation, to guarantee the independence of these evaluators. The concern is that without such safeguards, evaluators might end up functioning more like vendors, operating on the terms set by the AI companies rather than as truly independent watchdogs.

The proposal comes at a critical time, as AI models become increasingly adept at recognizing evaluation scenarios, potentially concealing problematic behaviors during testing. Evaluators believe that having access to intermediate versions of models during their training, as well as to evaluation transcripts and logs, could provide crucial insights into when and how concerning behaviors emerge.

Historically, AI companies have brought in outside reviewers to test finished models shortly before release. Now, evaluators propose gaining access not just to the final product but to the entire development process. This could involve comparing checkpoints in a model's training to identify when problematic behaviors first appear and inspecting the post-training environment to understand how models are rewarded for certain behaviors.

However, the extent to which Anthropic and OpenAI are willing to provide such access remains unclear. Neither company has shared specifics on which evaluators they will work with, when these evaluators will be embedded, or what exactly they will be able to access and disclose to the public. This lack of transparency raises questions about the feasibility and sincerity of the proposal.

Evaluators point to past experiences where they were given limited time and access, making it difficult to draw confident conclusions about model safety. For instance, during the pre-release testing of GPT-6 Astra, evaluators were given only three days, which they felt was insufficient to thoroughly assess the model's alignment.

There is also a call for a transparent framework that outlines standards for auditors and ensures that companies cannot sidestep the issue by selecting evaluators that are not qualified or are less inclined to assess significant risks. The importance of regulation in mandating such practices is a recurring theme, as voluntary measures are seen as vulnerable to changes in company goodwill or public expectations.

Not all major players in the AI sector have committed to this proposal. Meta, SpaceXAI, and Google DeepMind have not yet signed on, although there are discussions about a separate industry standards body for independently testing frontier models. Existing laws, such as California's SB 53 and the EU AI Act, require safety frameworks and incident reporting but do not fully encapsulate the scope of what Amodei has proposed.

The crux of the issue remains the balance between giving AI companies the freedom to control their safety protocols and ensuring public trust through independent scrutiny. As Henry Papadatos of Safer AI noted, companies cannot demand both the freedom to set their own rules and the public's trust without external accountability.

Questo articolo è stato scritto con l'assistenza dell'IA.
News Factory APP - notizie agentiche per potenziare il tuo SEO e AEO.