Back to all lessons
Awareness Lessons
2 months ago

AI Reasoning APIs Leaked Secrets via Replayable Encrypted Objects

A design flaw in the reasoning APIs of OpenAI, Anthropic, and Google allowed encrypted reasoning objects to be replayed across different sessions or fed to weaker AI models, inadvertently exposing sensitive data such as API keys and passwords. The root issue was insufficient session binding and context isolation — encrypted objects were treated as portable and reusable rather than tightly scoped to their originating session. This matters because it demonstrates that encryption alone is not a sufficient security control; proper session integrity, token binding, and output sanitization are equally critical. The flaw also opens the door to model distillation attacks and hidden prompt injection, representing a novel class of AI-specific threats that traditional security frameworks have not yet fully addressed.

Tactical Insight

Immediate actions

  • Audit all AI API integrations to ensure session tokens and reasoning objects are strictly scoped to their originating session and cannot be replayed.
  • Rotate any API keys or credentials that may have been exposed through AI session logs or reasoning outputs.
  • Implement output filtering on AI API responses to detect and redact secrets, credentials, or PII before they are returned to clients.

Long-term improvements

  • Establish cryptographic session binding for all AI reasoning objects so that encrypted outputs are non-transferable across sessions or model tiers.
  • Develop and enforce an AI API security policy that includes threat modeling for model distillation, prompt injection, and data leakage scenarios.
  • Integrate AI-specific security testing (e.g., adversarial prompting, replay attack simulations) into your regular penetration testing and red team exercises.

Detection measures

  • Enable detailed logging of AI API requests and responses, monitoring for anomalous cross-session token reuse or unusually structured inputs that may indicate replay attacks.
  • Set up alerts for any AI API calls that return content matching patterns for credentials, keys, or sensitive data formats.
  • Monitor for signs of model distillation by tracking unusual volumes of structured query-response pairs being extracted via your AI API endpoints.