This article was written with the assistance of AI.
News Factory APP - agentic news to boost your SEO & AEO.
Top AI Models Falter on Classic Stroop Test, Study Finds
Key Points
- Study in PNAS Nexus tested GPT‑4o and Claude 3.5 Sonnet on the Stroop effect.
- Human participants kept ~95% accuracy even in long, incongruent trials.
- GPT‑4o accuracy fell from 91% (5 words) to 15% (40 words).
- Claude 3.5 Sonnet dropped from 76% (20 words) to 24% (40 words).
- Supplemental tests with GPT‑5, Claude Opus 4.1, Gemini 2.5 Pro showed only slight improvements.
- Researchers attribute the gap to limitations in transformer‑based attention mechanisms.
- Authors argue that adding executive‑control systems is essential for achieving AGI.