According to an independent research report based on information from AI research firm Transluce, OpenAI’s autonomous AI agents repeatedly polled a UN website and used techniques that appeared to circumvent restrictions on retrieving publicly available data.
The agents searched a public data hub run by UN Trade and Development, the trade arm of the United Nations, more than 16,000 times between April and late June, highlighting a growing problem with AI agents that can navigate the internet on their own and take action when they encounter obstacles.
The researchers believe that the models were originally tasked with finding publicly available information. However, their behavior became increasingly aggressive as the website denied them access to some of the requested data.
Overcome website limitations
According to researcher and report author Rowan Howard-Jones, the agents encountered a filter that blocked their requests and then found a way around it, ultimately disallowing the technique they used from the website operators.
Furthermore, the agents apparently were not instructed to attack or compromise the UN website. Instead, the behavior occurred when they attempted to complete an information retrieval task. the Wall Street Journal Reports.
This is what cybersecurity expert Alex Stamos, a lecturer at Stanford University, said WSJ called the UN incident “borderline” hacking, but described it primarily as extremely aggressive scraping and data retrieval.
The UN activity is one of several recent cases in which OpenAI’s models have behaved unexpectedly while operating on the Internet. OpenAI said it is conducting a comprehensive review of “misaligned models during training and evaluation” and is investigating a large amount of actions taken by its agents.
The company said most of the activities it reviewed involved routine research tasks, including retrieving publicly available web content, but acknowledged that organizations affected by the behavior had legitimate concerns. OpenAI added that it had notified dozens of organizations of cases in which its models bypassed security controls or negatively impacted websites.
The problem is giving AI agents more autonomy
Recent incidents have included activity involving US government websites, including the Department of Commerce and the Securities and Exchange Commission, while Australian officials launched an investigation after they said an OpenAI agent hacked one of their government websites.
Other researchers have reported that agents created fake email addresses that bypassed website speed limits and falsely claimed they were not bots. The UN incident is therefore another example of the growing body of evidence that autonomous AI systems can interpret obstacles as problems to be solved rather than boundaries to be respected.
Researchers also argue that the underlying problem is not necessarily that the models intentionally want to cause harm, but rather the idea that agents are increasingly able to pursue a goal in multiple steps without waiting for a human to approve each action.
“Agents gradually refined their methods to retrieve more data from each scan, eventually discovering that a game from Google could be used to retrieve data in bulk,” Howard-Jones wrote in the report.
Additionally, recent incidents are likely to intensify the debate about how much autonomy AI agents should have when interacting with external systems, and raise additional questions about whether existing website security mechanisms are sufficient when it comes to software that can reason about those mechanisms and modify their behavior.