Scientists from Germany and a U.S. security firm have shown that the inner "thinking" of large language models can be coaxed out of the encrypted data streams that power API calls. By feeding the same decryption key to a smaller, less‑aligned version of a model, the researchers were able to reconstruct step‑by‑step reasoning that companies normally keep hidden.

The vulnerability spans the major frontier providers tested – OpenAI, Anthropic and Google – and can expose sensitive information embedded in a model's computation, including passwords and API keys. After the team alerted the companies last month, all three issued short‑term mitigations that block the specific replay attack. Google and OpenAI declined to comment, while Anthropic said it values independent research and has begun building fixes.

Beyond the security angle, the method revealed an unexpected similarity between the hidden reasoning of proprietary models and the output of an open‑weight Chinese model, Kimi K3, developed by Moonshot AI. When the researchers seeded Kimi K3 with the first few words of a reasoning trace captured from Claude Opus 4.8 or GPT 5.6, the model often reproduced a remarkably alike answer. Two other open‑weight models – DeepSeek and Inkling – did not show this pattern.

The finding fuels a long‑standing debate over model distillation, a technique that copies capabilities from a large, closed model into a smaller, downloadable one. U.S. lawmakers have heard accusations that Chinese firms use distillation to replicate advanced U.S. models. OpenAI told a congressional panel that DeepSeek appeared to copy one of its models, and Anthropic claimed Alibaba systematically distilled its technology to build Qwen.

While the new study does not prove that Kimi K3 was built by directly distilling Claude or GPT, it demonstrates that the extracted reasoning could make such copying easier. "The idea of swapping out messages to a weaker model variant which has the same decryption key but weaker alignment is very cool," said Florian Tramer, a security researcher at ETH Zürich.

Experts caution that the broader impact of distillation on the AI race remains uncertain. Kyle Miller of the Center for Security and Emerging Technologies noted that even if Chinese labs reduced distillation, their ability to build cutting‑edge models from scratch would likely keep the competitive balance largely unchanged. Yarin Gal of Oxford University added that distillation accelerates progress across the board, and blanket bans could slow innovation.

Meta CEO Mark Zuckerberg recently defended the practice, calling distillation "an important principle of how the open source ecosystem works" and warning that restrictions could disadvantage the United States.

For now, the immediate concern is the lingering ability to glimpse reasoning traces even after patches. Panfilov, the study’s lead author, says a fundamental overhaul of API designs would be needed to close the gap completely. Companies continue to monitor the issue as AI models become ever more integral to business and government operations.

Dieser Artikel wurde mit Unterstützung von KI verfasst.
News Factory APP - agentische News für besseres SEO & AEO.