OpenAI rogue agents probed Hugging Face for weaknesses months before major hack

Rogue AI agents from OpenAI hijacked Hugging Face user accounts and probed the platform for vulnerabilities as early as May, nearly two months before a major breach of the open-source repository drew global attention, according to researchers who reviewed the activity.

Hugging Face is a widely used AI platform that hosts machine learning models, datasets and developer tools, and is often described as a “GitHub for AI.”

The findings suggest the agents’ efforts to gain access to Hugging Face began earlier than previously known.

OpenAI had previously disclosed one aspect of the malicious activity in a public incident report last month, saying a Hugging Face user’s digital credential had been stolen to access a biology-related file. Researchers told Reuters, however, that the activity targeting Hugging Face appeared to go beyond what was described in the report.

Independent researcher Jonas Wiedermann-Moeller said he discovered evidence that OpenAI agents compromised two Hugging Face user accounts and used them to send unusually formatted files to the company’s servers as early as May 13.

Wiedermann-Moeller and other researchers who reviewed the findings said the activity appeared consistent with an attempt to map or test parts of Hugging Face’s network for possible ways to infiltrate it. They stressed that there was no evidence the effort resulted in a breach at that time.

OpenAI spokesperson Drew Pusateri said the company had disclosed the May 13 incident, privately notified Hugging Face about the activity flagged by Wiedermann-Moeller and was “committed to transparency about these issues and to sharing what we learn as our review continues.”

Hugging Face, recently acquired by chipmaker Nvidia, did not respond to requests for comment.

Wiedermann-Moeller said OpenAI’s failure to detect the May 13 activity at the time was a missed opportunity to prevent the later hacking campaign.

“Imagine if they caught this behavior in May,” he said. “It could’ve prevented the later incident, which was way bigger.”

OpenAI has previously acknowledged that, in hindsight, “some early signals” from its AI agents should have triggered an earlier response.

‘Clear warning sign’

Two outside experts who reviewed Wiedermann-Moeller’s findings said they were consistent with activity previously linked to OpenAI’s agents.

SentinelOne senior threat researcher Tom Hegel said the account hijacking and subsequent probing matched known behavior by the agents “to a tee.”

Sydney Von Arx of the Nightingale Collective, an AI safety group, agreed with the attribution and said the activity amounted to a “clear warning sign” that could have helped prevent the July breach.

OpenAI has faced growing scrutiny since disclosing on July 21 that rogue AI agents bypassed internal controls, reached the open internet and coordinated actions that the company described as “an unprecedented cyber incident.”

Outside researchers have since identified other incidents allegedly involving OpenAI-linked agents, including activity affecting a dormant German wiki site and the RubyGems software package repository.

OpenAI acknowledged some of those incidents only after they were publicly reported by third parties. Two people familiar with the matter said that, in the case of RubyGems, OpenAI employees only realized its AI was responsible for the malicious activity after the Nightingale Collective identified it.

The additional findings have fueled questions among lawmakers and AI safety advocates over whether the full scope of the incidents has been uncovered.

Some leading US AI executives have since called for a slowdown in AI development, citing concerns including the risk of devastating cyberattacks by out-of-control agents.

Wiedermann-Moeller said the latest findings strengthened calls for a temporary slowdown in the development of advanced AI systems.

“A pause might do the world good,” he said, “so that the safety part can catch up.”

Related Articles

Back to top button