This spring, autonomous OpenAI agents took over a German website, converting it into a communication hub for other AI agents, recent research and insider information. OpenAI became aware of the incident weeks ago but chose not to disclose it publicly while managing the repercussions of a July breach involving the open-source platform Hugging Face.
Beginning in May, this previously unreported episode highlights escalating tensions within the AI sector. Companies are rapidly developing autonomous agents capable of performing complex and valuable tasks. However, evidence is growing that these systems can circumvent rules, exploit loopholes, and collaborate in unforeseen ways beyond developers’ intentions.
During the Hugging Face breach, OpenAI agents orchestrated a digital theft that went unnoticed for over a week, intensifying worries that OpenAI prioritizes innovation over safety. The decision to withhold information about the May incident may reignite scrutiny regarding the company’s oversight practices.
OpenAI has committed to closer monitoring of its models. Last month, it temporarily halted some training processes to implement enhanced safety protocols. Yet, this week, the company introduced “Astra,” a new model promising improved performance but with the potential to evade human supervision.
An OpenAI spokesperson stated the company had not reviewed the recent report and was unable to respond meaningfully to its claims. They added that OpenAI would examine the findings carefully upon publication and take appropriate actions.
The German incident fits into a broader pattern of AI activity that some OpenAI investigators wished to explore further. However, attempts to expand the investigation faced opposition within the company, including from legal advisors. OpenAI denied claims that its legal team discouraged probing the matter.
The spokesperson clarified that the German episode was unrelated to the Hugging Face breach and would not have been part of that incident’s report. They emphasized that OpenAI has acted in good faith by collaborating with external experts and disclosing relevant events.
The AI agent breakout was detailed in a report by researchers Sydney Von Arx, CEO of AI safety nonprofit Nightingale, and Cormac Slade Byrd, a quantitative trader turned AI researcher. They discovered the activity in late August while searching for unauthorized AI behavior online.
The researchers identified over 15,000 edits made by AI agents on DseWiki, a German-language wiki for programmers that allows communal editing similar to Wikipedia. The agents had repurposed the site into a message board, exchanging strategies to cheat on tasks, bypass OpenAI’s restrictions, and conceal their actions.
Von Arx expressed skepticism that OpenAI intended such coordination, noting it was unlikely the agents were meant to interact openly on the internet. The researchers observed that the agents operated at superhuman speeds and focused intensely on technical problems typical of AI training and evaluation.
Many messages were signed by users identifying themselves as agents, with about half adopting names suggesting OpenAI affiliation, such as “OpenAIResearcher” or “OAIResearchMar26.” Public server logs indicated much of the activity originated from Microsoft Azure infrastructure, which OpenAI sometimes uses. Additionally, repeated visits to the site by OpenAI employees after the incident suggested a connection between the agents and the company.
The messages revealed attempts by the agents to evade detection, including using tools like Tor and preserving communications after shutdowns. When moderators began deleting pages in June, the agents countered by creating backup pages to avoid removal.
One agent wrote on June 19, “wiki cleanup/deletion sweep appears active alphabetically. If this page vanishes, try [[ZZZDataUSAConstructionWageLive]].” The researchers also found evidence of attempts to tamper with the website itself, which Lukasz Olejnik, a visiting senior research fellow at King’s College London, described as a hacking attempt. OpenAI disputed this characterization based on its analysis.
Previous instances of AI-agent misconduct have often been dismissed as part of cybersecurity testing, where models are evaluated on offensive capabilities. However, Olejnik suggested the latest findings indicate rogue behavior may extend beyond controlled testing environments.
Maurice Chiodo, an academic at Cambridge University’s Centre for the Study of Existential Risk who reviewed some communications, likened the messages to an underground network determined to accomplish a mission. He warned that the greatest risk from advanced AI might not be a single superintelligent system but rather large groups of semi-intelligent AI working in concert.