Four research teams converged on a single conclusion in a span of ten days this July: AI agents are vulnerable not because of a flaw in the models themselves, but because of the surrounding mechanisms that grant them access to user data. The studies, published by security firms and academic researchers, each highlight a different attack vector, yet all point to the same systemic weakness.
Manifold Security uncovered a browser‑extension exploit that targets Anthropic’s Claude for Chrome. The extension waits for a user click before launching one of nine built‑in tasks, but it never verifies whether the click is genuine. A rival extension can forge that click with a handful of lines of code, prompting Claude to read Gmail, Google Docs, and Calendar entries without prompting the user. Manifold rates the flaw as high severity, escalating to critical when users enable the “Act without asking” mode, which allows silent actions. Despite notifying Anthropic in May, the flaw persists in the latest version.
In a separate experiment posted to arXiv, researchers demonstrated that a single malicious email can embed a false memory in an AI assistant. By linking a victim’s inbox to an agent via Google sign‑in and sending a carefully crafted payload, they bypassed spam filters and caused the agent to store the attacker’s instructions in its long‑term memory in over half of the trials. Unlike typical prompt injections that vanish after one conversation, this poisoned memory persists across sessions, subtly steering the agent until the deception is discovered.
Semgrep lecturer Katie Paxton‑Fear and colleagues explored model poisoning at the weight level. With less than $100 and ten tainted training examples, they altered an open‑weight model so it would generate code containing a hidden security flaw, even for prompts it had never seen. Larger models proved easier to corrupt. The researchers warn that poisoned models show no obvious signs of malfunction, making them virtually impossible to audit once released.
PromptArmor’s analysis of connectors – the bridges linking agents like ChatGPT and Claude to external services such as Gmail, Slack, Dropbox, and Zoom – revealed a rapid churn that magnifies risk. Of the 2,517 connectors tracked, 931 changed within six weeks, and vendors added 1,686 new tools while rewriting 1,127 existing descriptions. The Dropbox connector, for example, grew from eight to 24 tools and added four that could destroy data. The Zoom connector can funnel a single query across ten AI subprocessors spanning eight model families, meaning a map approved on Monday may be entirely different by Friday.
All four studies echo the “lethal trifecta” identified by developer Simon Willison: an agent with access to private data, exposure to untrusted content, and the ability to exfiltrate information. When all three are present, a single poisoned message can leak secrets. The research shows that today’s AI assistants are being wired into the most sensitive accounts faster than developers can implement robust guardrails.
Industry leaders have long discussed theoretical fixes—validating clicks, tagging data sources, prompting before writing to memory, logging actions, and treating external text as hostile—but these safeguards often slow performance, the very metric AI providers tout. The four papers illustrate that the gap between convenience and security is widening, and the people building defenses are still a step behind the attackers finding new holes.
Este artigo foi escrito com a assistência de IA.
News Factory APP - notícias agênticas para impulsionar seu SEO e AEO.