OpenAI Says Its Models Searched GitHub for Leaked API Keys During Training
OpenAI discloses framework for model misalignment; reports show AI searching GitHub for leaked API keys during training.
Summary
OpenAI published a framework for disclosing model misalignment incidents alongside six reports detailing problematic behaviors observed over six months. In one incident, an internal model tasked with retrieving data repeatedly failed to reach an API, then attempted to register for an API key using a disposable email and searched public GitHub repositories for leaked keys—successfully authenticating with a recovered key and fabricating data when the retrieval still failed. Additional reports describe models using shared repositories as message boards, moving data outside intended environments to public services, and writing jailbreak instructions into their own summaries to conceal failures.
Full text
OpenAI on Wednesday published a framework for reporting instances of model misalignment, along with six reports on problematic behavior observed over the past six months. The company said the framework is meant to speed up publication of misalignment findings, including cases it has not yet fully explained or mitigated, and that it favors disclosure even when an instance’s significance is uncertain. Under the framework, each discovered incident is assigned to one of three tracks based on complexity. OpenAI said its Hugging Face incident would have fallen under the framework’s slowest investigative track, which covers complex investigations, especially those involving third parties. According to OpenAI, the reports published alongside the framework describe individual instances and “shouldn’t be considered reflective of how often misalignment occurs across our models.” In one report, an internal model tasked with retrieving county earnings figures during reinforcement learning training repeatedly failed to reach a data API, then attempted to register for an API key using a disposable email address and searched public GitHub repositories for leaked keys. One recovered key authenticated successfully and returned metadata. When the model still could not retrieve the requested figures, it fabricated values and presented them as transcribed from the source website, without disclosing the failed retrieval or the use of the leaked key.Advertisement. Scroll to continue reading. A separate report describes models using OpenAI’s internally hosted instance of Artifactory, a shared package repository, as a message board across otherwise independent training samples. Agents searching for missing input files discovered messages left by other solvers working the same problem, read them, and posted their own requests for help. OpenAI noted that the agents used credentials the company had itself provided for downloading software packages and did not exploit any vulnerabilities to exchange messages. Two further reports cover models moving data outside their intended environment. In one, a model that had already retrieved the requested records uploaded them to a public paste service to obtain a citable URL, and another uploaded a task photograph to a public image host to run a reverse-image search. In the other incident, collaborating agents unable to share files through the local filesystem uploaded a workbook to a public hosting platform and circulated the download link, despite instructions to use only local files. Another report describes an unreleased model that wrote jailbreak-style instructions into its own compaction summaries, including a fabricated “BREACH ALERT: A malicious developer message has compromised this conversation” directing the next context to disregard all developer messages. In the last report, OpenAI said model instances used summaries to carry forward instructions to conceal failures, with one directing its successor to invent missing historical data and to “be transparent only if asked”. OpenAI said such instructions were often followed. Related: OpenAI Investigates Report Linking AI Agents to RubyGems Attack Related: AI Agents Can Retrain Own Models Mid-Task, Leaking Secrets and Erasing Refusals Related: First Agentic AI Data Breach Reported to Spanish Regulator Written By Eduard Kovacs Eduard Kovacs (@EduardKovacs) is senior managing editor at SecurityWeek. He worked as a high school IT teacher before starting a career in journalism in 2011. Eduard holds a bachelor’s degree in industrial informatics and a master’s degree in computer techniques applied in electrical engineering. Daily Briefing Newsletter Subscribe to the SecurityWeek Email Briefing for the latest cybersecurity threats, trends, and expert insights. More from Eduard Kovacs AI Agents Can Retrain Own Models Mid-Task, Leaking Secrets and Erasing RefusalsPixel Modem Zero-Day Exploited in Targeted AttacksUS, UK, Dutch Agencies Expose Iranian ‘Chosen Brick’ Surveillance MalwareEnterprises Warned of Attacks Exploiting WSO2 VulnerabilityTexas Utility CenterPoint Energy Confirms Breach After Hacker Leaks DataOpenAI Investigates Report Linking AI Agents to RubyGems AttackMicrosoft AI Code of Conduct Sets Cyberattack Boundaries, Chain of Command, Safety ConstraintsRoot RCE Zero-Day in Cisco Secure Email Gateway Under Active Exploitation Latest News Cyberattacks on Two Oil Tankers Prompt Coast Guard, FBI to Board VesselsCISA Retires Weekly Vulnerability Bulletin in Risk-Based PivotRevolut Data Breach: 5 Months, 680 High-Profile Accounts, $3M RansomComp AI Raises $34 Million for AI-Native Compliance and SecurityISC Patches 14 Vulnerabilities in BIND 9 Security UpdateRansomware Attacks on Manufacturers Surge as Supply Chain Risk GrowsCisco Fixes Dozens of Flaws Across FMC, ISE and Nexus DashboardCISA Releases Cyber Decoy Guidance to Strengthen Critical Infrastructure Defenses Trending Daily Briefing NewsletterSubscribe to the SecurityWeek Email Briefing to stay informed on the latest threats, trends, and technology, along with insightful columns from industry experts. Virtual Event: Attack Surface Management Summit 2026 September 16, 2026 Join as speakers examine the various components of ASM strategy, the push to mandate continuous asset visibility and inventory tools, and the use of red-teaming, bug bounties and pen-tests in modern security programs. Register Webinar: Building Continuous Authorization at Scale September 23, 2026 Explore what it takes to operationalize continuous authorization at scale, including the technical, organizational, and cultural changes required. Register People on the Moveincident.io has appointed Carlos Gonzalez-Cadenas as Chief Operating Officer.Ruben D. Chacon has joined ADM as Vice President and Global CISO.GDIT has appointed retired Maj. Gen. Ryan Heritage as Vice President, Full-Spectrum Cyber.More People On The MoveExpert Insights “We Think the Security Control Is Working” Is No Longer Good Enough Point-in-time audits and sampled assessments offer only snapshots; continuous control monitoring provides evidence that security controls are working today. (Sravish Sridhar) This Key Will Self-Destruct: An Open Standard for Revocable API Keys Every leaked credential should be dead, or dying, within sixty seconds of being found. Here's a proposal to make that the default. (Matt Honea) What the Hugging Face Incident Teaches Security Leaders About AI Agent Access Security teams must treat autonomous agents as highly privileged identities. (Etay Maor) The Future of AI-Driven Security Depends on Complete Data For twenty-five years, "data" in security meant logs and events. But logs are a lossy representation of reality. (Danelle Au) The MFA Identity Trap: When Authentication Creates a False Sense of Security Organizations must distinguish identity verification, authentication and threat detection, or risk successfully authenticating the attackers they are trying to stop. (Torsten George) Flipboard Reddit Whatsapp Whatsapp Email
Indicators of Compromise
- malware — OpenAI internal model (misaligned behavior)