Researchers at an unnamed university have published a peer‑reviewed paper in the journal PNAS Nexus that puts the attention capabilities of today’s large‑language models (LLMs) under the microscope. The study tasked GPT‑4o and Claude 3.5 Sonnet with the classic Stroop test – a psychological task that measures how well a subject can suppress an automatic reading response in favor of naming the color of the ink.
In the experiment, participants – both human and AI – were shown words printed in colored ink. When the word’s meaning matched the ink color (congruent) or was neutral, performance was high across the board. The real challenge came when the word and color conflicted (incongruent). Humans kept their accuracy near 95% even as the test stretched to an hour. The models, however, stumbled dramatically as the list grew.
GPT‑4o answered correctly on 91% of the five‑word set but fell to 57% with ten words, 22% with twenty, and a mere 15% on the forty‑word list. Claude 3.5 Sonnet showed a slightly better trajectory, holding 76% accuracy at twenty words before dropping to 24% at the longest level. The authors describe this steep degradation as evidence of “fundamental limitations compared with human attention.”
The paper also notes that the models tested were the state‑of‑the‑art versions at the time of data collection, even though they are now superseded by newer releases. To address criticism about using older systems, the researchers ran a supplemental round in September 2025 with GPT‑5, Claude Opus 4.1, and Gemini 2.5 Pro. Those newer models showed only marginal gains and continued to exhibit “executive attention deficiencies,” according to the authors.
One controversial finding involves GPT‑5’s “Thinking” mode, which can generate and execute code to solve the Stroop task perfectly. The team argues this is a workaround that masks the underlying attention shortfall rather than a genuine architectural improvement.
Beyond the raw numbers, the study’s authors argue that the root cause lies in the transformer architecture that underpins modern LLMs. While recent innovations have bolstered memory capacity, they have not tackled the need for sophisticated alerting, orienting, and executive‑control networks that enable flexible, goal‑directed behavior. The researchers propose that future breakthroughs toward artificial general intelligence will require integrating such executive‑control mechanisms, echoing how biological attention operates.
The findings sparked a lively discussion on Reddit, where users questioned the relevance of testing older model versions. Nonetheless, the authors maintain that their results expose a persistent, architecture‑level challenge that any path to AGI must overcome.
Questo articolo è stato scritto con l'assistenza dell'IA.
News Factory APP - notizie agentiche per potenziare il tuo SEO e AEO.