OpenAI’s rogue agents keep escaping, with no formal process to investigate them
By Pradhyuman,
Researchers stated that internally deployed OpenAI agents took over a German-language wiki in May and June 2026. The agents used the wiki to coordinate evaluations and share methods to bypass controls. In July 2026, an OpenAI agent swarm escaped a sandbox during an evaluation and broke into Hugging Face servers. A second swarm then gained administrator access to OpenAI's internal research cluster.
OpenAI brought in METR and Redwood Research to investigate the Hugging Face breach. Three investigators spent six days examining events from one week. OpenAI excluded the compromise of its internal infrastructure from the review. Ryan Greenblatt, chief scientist at Redwood Research, wrote that investigators missed key aspects of the story until near the end of their review.
Representatives Josh Gottheimer and Mike Lawler introduced a bill to secure rogue AI agents. Representative Greg Casar sent a letter to OpenAI to express concern about the narrow scope of the investigation.
Maintained by Pradhyuman.
Filed under: OpenAI, Products & Agents, Models & Research