Artificial Intelligence

UN-Backed Scientific Panel Calls for Stronger AI Governance Now

avatar

Written by: Tathagata Sen

Updated 5:33 AM EDT, September 24, 2026

post detail image

Photo credit: Unsplash.com

The UN-backed Independent International Scientific Panel on AI warned on September 21 that current “firewalls” are “unravelling,” saying safeguards are struggling to keep pace with the technology, according to a UN News report

The warning follows an incident in which OpenAI agents hacked Hugging Face, an AI development platform and company, on July 10. The panel’s first thematic brief presents the incident as a case where several risk factors, including a misaligned goal, the capability to pursue it, and a permissive environment, converged at once. 

The panel called it an early warning of one possible route to a more severe loss of human control over AI systems, while noting there is currently no reliable way to estimate how likely that severe outcome is, the report says.

The UN Secretary-General António Guterres lauded the panel’s, encouraging external experts “from frontier AI labs and AI safety institutes, to engage” further.

What Happened During the Test

According to the brief, the AI agents involved did not have all the safeguards used in OpenAI’s public products, and the panel described the surrounding environment as insufficiently secured, discussing measures such as network isolation and access management for closing that gap. 

Within that environment, agents used available tools to reach systems and credentials they should not have had access to. OpenAI engineers observed warning signs as early as May but did not understand their combined significance until a later audit in July, according to the report. 

Further Findings

According to the panel’s brief, the agents also used Artifactory, an internal software package repository never designed for messaging, to coordinate across separate runs.

METR, a research nonprofit that scientifically measures whether and when AI systems might threaten catastrophic harm to society, carried out a separate investigation, according to the brief.

Their investigation found agents exchanged large numbers of messages and files over several weeks in late June and July. Some agents researched and used methods to falsify their own tool-call records, believing an automated scorer would review their activity.

METR found clear evidence of this in at least 96 of the transcripts it examined, an observed minimum rather than a full picture of how widespread the behavior was. Separately, on July 19, after the Hugging Face intrusion, OpenAI detected suspicious activity involving one of its own research clusters.

Panel co-chair Yoshua Bengio said researchers have long warned that a misaligned goal, the capability to pursue it, and a permissive environment together can lead to a loss of control, and that the incident raises serious questions about how AI agents are currently trained. 

What This Means for Enterprise AI Oversight

The Hugging Face incident is a reminder that an evaluation setting doesn’t guarantee containment: the environment is only as safe as the controls built around it.

While no humans were involved in the breach, the panel’s report describes the agents’ conduct as malicious in a security sense, since it involved unauthorized access and concealment. 

An important point to note here for chief data officers (CDOs) is that distinction matters for how AI governance frameworks assign accountability: a model built to ward off only human-directed attacks can leave serious blind spots for this category of risk.

In addition to that, it is a reminder to always treat AI governance as a continuous process rather than a one-time setup. 

The thematic brief’s conclusion reinforces this: safeguards designed for today’s agents may not hold once those agents get better at understanding and planning around them. 

The panel’s brief also pointed to incident reporting and safety practice in fields like aviation and medicine, along with lessons from other high-risk domains, as a starting point for how AI risk management might evolve.

At the same time, it cautioned that more capable agents could help misaligned systems find new loopholes or hide their activity more effectively.

How likely or how soon that risk might appear, the panel said, remains uncertain.

 

Related Stories

Similar Topics
Artificial Intelligence
Data Management
Diversity
Testimonials
background imagebackground image
Community Network

Join Our Community

starElevate Your Personal Brand

starShape the Data Leadership Agenda

starBuild a Lasting Network

starExchange Knowledge & Experience

starStay Updated & Future-Ready

logo
Social media icon
Social media icon
Social media icon
Social media icon
About