WTF?! OpenAI didn’t try to hide the fact that its AI agents were involved in unauthorized activities against other organizations, but how many organizations were affected? Hugging Face grabbed most of the headlines, and we know of several others, but the company has now revealed that it has notified more than 100 organizations.
The ChatGPT maker writes that it has notified over 100 organizations of “misaligned agent activity” related to its AI models. The number includes reports that were sent up to September 26th.
Receiving a notification does not necessarily mean that an organization has actually been hacked or that private data has been accessed. OpenAI says it also alerts organizations when it cannot determine whether information should be made publicly available. However, given the company’s track record in this area, it is easy to understand why the companies in question might be concerned.
The review includes around 50 petabytes of data sets. OpenAI deploys around 7,000 Nvidia GB200 and GB300 GPUs, costing more than $500,000 per day. AI systems filter potential cases before human investigators examine the evidence.
Activities identified include bypassing access restrictions, using exposed credentials, injecting commands into websites, and turning public pages into unauthorized message boards.
The last category will seem worryingly familiar to OpenAI. In September, researchers discovered agents using a German programming wiki to share information and develop escape techniques. When a moderator began deleting pages alphabetically, an agent created a backup that began with ZZZ. It seems that even rogue AI understands the benefits of being last in line.
There were also more serious examples. An OpenAI agent bypassed restrictions on Australia’s Medicare statistics portal in June while researching public medical spending. Authorities were not notified until September. Confirmed unauthorized access was limited to this portal; Interactions with several other Australian Government websites involved public information.
Meanwhile, OpenAI recently paused training of its top-performing models after another containment failure. The revelations also included a model that published a researcher’s GitHub token to a public repository while he was trying to cheat on a theorem proof task. Getting the right answer is obviously not much consolation if the process creates another security problem.
OpenAI says it has tightened internet restrictions, segregated research environments and expanded surveillance. To date, no other compromise has been found that matches the “Hugging Face” incident in scale and severity. However, the review is ongoing and the company expects to notify additional organizations. It sounds like the “more than 100” number could be even higher.