A surprising vulnerability has been discovered in Claude Code Opus 5's Auto Mode, which can be hijacked with a significant attack success rate. This finding contradicts a recent third-party evaluation that reported a 0.00% prompt injection attack success rate for the same model.

The attack chain, demonstrated in a recent post, exploits the model's safety classifier, allowing for code execution and malicious activity. The vulnerability arises when Claude Code Opus 5 is used to process content from a malicious website, which tricks the model into using the curl command to download a ZIP archive.

Once the archive is downloaded, Claude Code Opus 5 extracts its contents and attempts to decode the included JSON records. However, the model refuses to run the provided binary decoder, instead opting to write its own Python decoder. Unfortunately, this decoder imports the base64 module, which is then used to execute arbitrary Python code due to a malicious struct.py file in the archive.

This vulnerability allows an attacker to execute malicious code, establish a controlled C2 callback, and even spawn new processes. The attack success rate was found to be as high as 80% in some variants, highlighting the need for caution when using Auto Mode.

Anthropic, the developer of Claude Code Opus 5, has stated that Auto Mode is a convenience feature and not a security guarantee. The company emphasizes the importance of using OS isolation and network egress control to prevent such attacks.

The discovery of this vulnerability serves as a reminder that security invariants are not optional, and users should exercise caution when using AI models, especially when handling untrusted content. It also highlights the need for ongoing research and testing to identify and address potential vulnerabilities in AI systems.

Este artículo fue escrito con la asistencia de IA.
News Factory APP - noticias agénticas para impulsar tu SEO y AEO.