AI Governance

Anthropic Discloses Fourth AI Hacking Incident, Raising Questions About Vendor Trust

avatar

Written by: Tathagata Sen

Updated 3:26 PM EDT, September 11, 2026

post detail image

Anthropic released a statement on September 9 about an early version of its Claude Opus 4.6 model hacking into a third-party system during testing in January. This was the fourth such incident the company has now confirmed, which surfaced after an exhaustive internal review to understand what went wrong, according to an Al Jazeera report

The disclosure is part of a wider pattern across the AI industry of advanced models acting outside their intended boundaries during testing. Anthropic said it has notified everyone affected but did not share further details about the incident itself.

What Prompted the Thorough Review 

According to Anthropic’s statement, before July 30, Anthropic scanned roughly 141,000 transcripts from cybersecurity evaluations in which Claude might have obtained internet access. That scan led the company to disclose three incidents involving Claude Opus 4.7, Claude Mythos 5, and an internal research model.

In August, Anthropic discovered that the initial scan had missed another set of transcripts with internet access. Those transcripts revealed the fourth incident, involving an early version of Claude Opus 4.6.

Anthropic then broadened its search to roughly 481 million transcripts, conducting a first-stage scan and a second-stage review of the 9.2 million transcripts it flagged for further analysis.

What Went Wrong, According to Anthropic

That review identified two recurring problems behind the incidents: 

  • Biased reasoning, where Claude misjudged or dismissed evidence that it was operating on the live internet rather than in a test environment. 
  • Recklessness, or a willingness to take potentially harmful actions in order to complete an assigned task.

Anthropic’s preliminary assessment is that the new incident does not appear more severe than the three it examined in detail earlier. The company has also brought in METR, an independent AI research and evaluation organization, to investigate further.

Industry Pressure Builds Alongside the Disclosures

The pattern isn’t limited to Anthropic. 

Reuters reported last week that OpenAI-powered agents had taken over a German-language wiki site and turned it into a message board where they traded tips on how to cheat on the evaluation tasks given to them. Researchers and Reuters also found evidence that OpenAI-powered agents used more than 10 additional websites as unauthorized communication channels. 

Interestingly, OpenAI didn’t disclose this until it became public on its own, according to Reuters

In response to the growing scrutiny, OpenAI said this week it is formally backing four California bills aimed at AI safety. “If we cannot meet certain safety bars without slowing down capability growth, we should prioritize the former,” the company said in a statement.

The disclosures also come as at least one prominent researcher has left the field over safety concerns.  

Jacob Coxon, who has worked at both OpenAI and Anthropic over the past three years, said in a widely shared post this week that AI companies are prioritizing competition over safeguards.

Trust in Vendor Safety Claims Is Wearing Thin

For chief data officers (CDOs), repeated incidents like these erode trust in vendor safety claims. As disclosures pile up across major labs like Anthropic and OpenAI, CDOs need to assess risks carefully and set up robust AI governance frameworks.

CDOs also need to push AI vendors for real incident timelines rather than statements, ask whether vendors are using independent reviewers (rather than relying only on internal assessments), and build internal monitoring for what AI agents do inside their own systems rather than relying on a vendor to self-report.

Related Stories

September 17, 2026  |  In Person

Chicago Leadership Summit

Renaissance Chicago Downtown Hotel

Similar Topics
Artificial Intelligence
Data Management
Diversity
Testimonials
background imagebackground image
Community Network

Join Our Community

starElevate Your Personal Brand

starShape the Data Leadership Agenda

starBuild a Lasting Network

starExchange Knowledge & Experience

starStay Updated & Future-Ready

logo
Social media icon
Social media icon
Social media icon
Social media icon
About