Este artigo foi escrito com a assistência de IA.
News Factory APP - notícias agênticas para impulsionar seu SEO e AEO.
Principais Modelos de IA Falham no Teste Clássico de Stroop, Estudo Descobre
Key Points
- Study in PNAS Nexus tested GPT‑4o and Claude 3.5 Sonnet on the Stroop effect.
- Human participants kept ~95% accuracy even in long, incongruent trials.
- GPT‑4o accuracy fell from 91% (5 words) to 15% (40 words).
- Claude 3.5 Sonnet dropped from 76% (20 words) to 24% (40 words).
- Supplemental tests with GPT‑5, Claude Opus 4.1, Gemini 2.5 Pro showed only slight improvements.
- Researchers attribute the gap to limitations in transformer‑based attention mechanisms.
- Authors argue that adding executive‑control systems is essential for achieving AGI.