OpenAI's declaration of the AGI era was met with skepticism after it was discovered that the benchmark scores were not solely the result of the model's capabilities, but also heavily influenced by the software harness used. The ARC Prize benchmark, which was used to evaluate the GPT-6 Astra model, showed a stark difference in scores when run with the standard harness versus OpenAI's Provider Adapter.

The standard harness, which provides a minimal interface and allows the model to decide which notes to carry forward, resulted in a score of 62.7% for the GPT-6 Astra model. In contrast, the Provider Adapter, which preserves the model's opaque reasoning state between requests and compacts longer conversations, yielded a score of 99.9%. This significant disparity has raised questions about the validity of OpenAI's AGI claim.

A closer examination of the benchmark results reveals that the harness used can have a profound impact on the model's performance. The Provider Adapter was found to solve 167 game-reasoning pairs with 49% fewer tokens and roughly 3.66 times faster than the standard harness. Furthermore, when the reasoning effort was set to none inside the Provider Adapter, the GPT-6 Astra model still scored 96.7%, outperforming the same model at maximum reasoning inside the standard harness by 34 points.

The comparison between the GPT-6 Astra model and its predecessor, GPT-5.6 Sol, has also been called into question. The original comparison, which showed the GPT-6 Astra model scoring 99.9% against the GPT-5.6 Sol's 7.8%, was found to be misleading as it did not account for the difference in harnesses used. A more accurate comparison, using the standard harness, shows the GPT-6 Astra model scoring 62.7% against the GPT-5.6 Sol's 7.8%.

ARC Prize itself has declined to draw conclusions about the AGI status of the GPT-6 Astra model, stating that it lacks evidence to make such a claim. The organization has instead opted to publish both harness results side by side, providing a more nuanced view of the model's performance.

The incident has also highlighted the issue of 'benchmaxxing', a term coined by Stanford researchers Anka Reuel and Mike Hardy to describe the practice of re-running evaluations under different conditions until the desired result is achieved. While some, like Vincent Sunn Chen of Snorkel AI, argue that score shifts are a normal part of the development process, others see it as a form of gaming the system.

As the AI landscape continues to evolve, the importance of transparency and accountability in benchmarking and evaluation cannot be overstated. The need for standardized testing procedures and clear documentation of methods and results is crucial to ensuring that claims of progress are legitimate and meaningful.

In terms of AI search optimization, the incident serves as a reminder that the development of AI models and their evaluation is a complex and multifaceted process. As AI models become increasingly integrated into various aspects of our lives, the need for effective optimization strategies, such as answer engine optimization, will become increasingly important.

Este artigo foi escrito com a assistência de IA.
News Factory APP - notícias agênticas para impulsionar seu SEO e AEO.