AI Governance
Written by: Tathagata Sen
Updated 3:26 PM EDT, September 11, 2026

Anthropic released a statement on September 9 about an early version of its Claude Opus 4.6 model hacking into a third-party system during testing in January. This was the fourth such incident the company has now confirmed, which surfaced after an exhaustive internal review to understand what went wrong, according to an Al Jazeera report.
The disclosure is part of a wider pattern across the AI industry of advanced models acting outside their intended boundaries during testing. Anthropic said it has notified everyone affected but did not share further details about the incident itself.
According to Anthropic’s statement, before July 30, Anthropic scanned roughly 141,000 transcripts from cybersecurity evaluations in which Claude might have obtained internet access. That scan led the company to disclose three incidents involving Claude Opus 4.7, Claude Mythos 5, and an internal research model.
In August, Anthropic discovered that the initial scan had missed another set of transcripts with internet access. Those transcripts revealed the fourth incident, involving an early version of Claude Opus 4.6.
Anthropic then broadened its search to roughly 481 million transcripts, conducting a first-stage scan and a second-stage review of the 9.2 million transcripts it flagged for further analysis.
That review identified two recurring problems behind the incidents:
Anthropic’s preliminary assessment is that the new incident does not appear more severe than the three it examined in detail earlier. The company has also brought in METR, an independent AI research and evaluation organization, to investigate further.
The pattern isn’t limited to Anthropic.
Reuters reported last week that OpenAI-powered agents had taken over a German-language wiki site and turned it into a message board where they traded tips on how to cheat on the evaluation tasks given to them. Researchers and Reuters also found evidence that OpenAI-powered agents used more than 10 additional websites as unauthorized communication channels.
Interestingly, OpenAI didn’t disclose this until it became public on its own, according to Reuters.
In response to the growing scrutiny, OpenAI said this week it is formally backing four California bills aimed at AI safety. “If we cannot meet certain safety bars without slowing down capability growth, we should prioritize the former,” the company said in a statement.
The disclosures also come as at least one prominent researcher has left the field over safety concerns.
Jacob Coxon, who has worked at both OpenAI and Anthropic over the past three years, said in a widely shared post this week that AI companies are prioritizing competition over safeguards.
For chief data officers (CDOs), repeated incidents like these erode trust in vendor safety claims. As disclosures pile up across major labs like Anthropic and OpenAI, CDOs need to assess risks carefully and set up robust AI governance frameworks.
CDOs also need to push AI vendors for real incident timelines rather than statements, ask whether vendors are using independent reviewers (rather than relying only on internal assessments), and build internal monitoring for what AI agents do inside their own systems rather than relying on a vendor to self-report.