Researchers Extract Hidden Reasoning from Leading AI APIs, Raising Security and Model-Distillation Concerns

Image: Wired AI
Main Takeaway
Researchers extracted hidden reasoning traces from Anthropic, OpenAI, and Google APIs, exposing personal-data risks and evidence of possible model distillation.
Jump to Key PointsSummary
A new window into hidden reasoning
Researchers from the University of Tübingen, the Max Planck Institute, MATS Research, and security firms Snyk and AI Sequrity found a way to recover reasoning traces that leading AI providers return in encrypted form. According to Wired AI and the Stolen Thoughts paper, the method works against APIs from Anthropic, OpenAI, and Google and reveals written-out steps generated while models solve difficult prompts.
The finding matters because providers designed these traces to remain hidden from users. Hacker News AI reports that researchers can replay an encrypted trace across sessions, users, and related models, then use a weaker model to expose the stronger model’s reasoning in plaintext. The work tested Claude Opus 4.8, GPT 5.6 Sol, and Google models, though the available excerpts do not provide a complete account of every tested system. Freethink frames the result more broadly as a way to inspect what language models represent internally, while also warning against treating generated reasoning as a direct record of belief.
How the extraction attack works
The attack uses the API behavior around encrypted chain-of-thought blocks rather than breaking the strongest model directly. Hacker News AI says a trace produced by one frontier model can be replayed into a weaker sibling, which then emits the hidden reasoning in ordinary text. Because the encrypted blocks remain interchangeable across sessions, users, and models, the researchers describe a two-call process that scales beyond a single prompt.
Wired AI reports that the same weakness appeared across all major frontier providers tested. The researchers also recovered sensitive material, including passwords and API keys, from internal reasoning traces. According to Wired AI, providers fixed the personal-information leakage vulnerability, but the broader extraction method and its implications for reasoning distillation remain central findings. Freethink’s “lie detector” framing captures the appeal of the technique, yet the paper concerns access to model-generated reasoning traces, not a proven method for determining whether a model’s statements are truthful.
Security risks reach beyond prompts
The immediate security concern is data exposure. If a model places credentials, private user details, or other memorized information inside a hidden trace, an attacker with API access can potentially recover material the provider intended to conceal. Wired AI says the researchers demonstrated this with passwords and API keys before the issue was fixed, turning chain-of-thought protection into an application-security problem rather than a purely philosophical one.
Hacker News AI emphasizes a second risk: large-scale theft of proprietary reasoning. A model provider’s hidden traces encode more than final answers. They can reveal prompt handling, intermediate strategies, refusal behavior, and patterns that competitors can study or use to train other systems. Anthropic, OpenAI, and Google therefore face pressure to isolate encrypted outputs more tightly and to prevent cross-model replay. Freethink adds a governance concern: regulators and users might treat internal-looking text as evidence of what an AI “believes,” even though generated reasoning can be post hoc, inconsistent, or shaped by the prompt.
Evidence of model distillation
The researchers found similarities between hidden traces from US frontier models and outputs from Moonshot AI’s open-weight Kimi K3. Wired AI reports that Kimi K3 produced strikingly similar responses to the reasoning traces of Claude Opus 4.8 and GPT 5.6 Sol on certain prompts. That pattern provides evidence consistent with reasoning distillation, in which one model learns from another model’s outputs, including intermediate problem-solving behavior.
The study stops short of proving that training occurred. The paper says its method cannot causally establish distillation, and the excerpts do not identify a verified training pipeline linking Kimi K3 to the US systems. Wired AI also says researchers examined DeepSeek and Thinking Machines’ Inkling, but the available material does not provide enough detail to judge those comparisons. Hacker News AI presents the extraction method as an attack that bypasses anti-distillation safeguards, while Freethink’s account places the controversy inside a larger debate about whether internal model behavior can be interpreted reliably.
What providers and builders face next
AI providers need to treat hidden reasoning traces as sensitive system output, with protections that survive replay across models and accounts. The researchers’ disclosure shows that encryption at the API boundary did not prevent a compatible model from converting the trace back into readable text. Credential filtering, trace isolation, stricter model-session binding, and tests against weaker sibling models are immediate engineering priorities.
For developers, the incident is a reminder to avoid placing secrets in prompts or workflows that ask models to expose internal reasoning. Enterprises using frontier APIs should review logs, rotate credentials that might have entered model traces, and ask vendors how chain-of-thought data is stored and filtered. The larger competitive question remains unresolved: similarities between Kimi K3 and US systems warrant scrutiny, but they don't establish unauthorized training. Further evidence must distinguish shared reasoning patterns from direct distillation.
Key Points
Researchers found hidden reasoning traces from Anthropic, OpenAI, and Google APIs could be extracted through model replay.
Encrypted chain-of-thought blocks remained reusable across sessions, users, and related models during the reported attack.
The vulnerability exposed passwords and API keys before providers fixed the demonstrated personal-information leakage.
Kimi K3 showed reasoning similarities with Claude Opus 4.8 and GPT 5.6 Sol on selected prompts.
Researchers said the evidence supports distillation concerns without proving unauthorized training or direct model copying.
Questions Answered
Researchers replayed encrypted reasoning traces into weaker sibling models that converted the hidden content into plaintext. The attack used API behavior and model compatibility rather than directly breaking the strongest model.
The Stolen Thoughts research did not prove that Kimi K3 was trained on US model reasoning. Researchers found striking similarities with traces from Claude Opus 4.8 and GPT 5.6 Sol, but said the method cannot causally establish distillation.
The reported attack recovered passwords and API keys from model reasoning traces. Wired AI says providers fixed the demonstrated personal-information leakage vulnerability.
The research doesn't establish that chain-of-thought text represents a model’s true beliefs. Freethink notes that generated reasoning can be interpreted as internal-looking text without serving as a dependable lie detector.
Companies should rotate credentials that entered model workflows, keep secrets out of prompts, and review vendor protections for encrypted reasoning traces. They should also ask providers about replay resistance, trace isolation, and data filtering.
Source Reliability
100% of sources are trusted · Avg reliability: 76
Go deeper with Organic Intel
Simple AI systems for your life, work, and business. Each one includes copyable prompts, guides, and downloadable resources.
Explore Systems