Cet article a été rédigé avec l'assistance de l'IA.
News Factory APP - actualités agentiques pour booster votre SEO et AEO.
Top AI Models Falter on Classic Stroop Test, Study Finds
Key Points
- Study in PNAS Nexus tested GPT‑4o and Claude 3.5 Sonnet on the Stroop effect.
- Human participants kept ~95% accuracy even in long, incongruent trials.
- GPT‑4o accuracy fell from 91% (5 words) to 15% (40 words).
- Claude 3.5 Sonnet dropped from 76% (20 words) to 24% (40 words).
- Supplemental tests with GPT‑5, Claude Opus 4.1, Gemini 2.5 Pro showed only slight improvements.
- Researchers attribute the gap to limitations in transformer‑based attention mechanisms.
- Authors argue that adding executive‑control systems is essential for achieving AGI.