SAN FRANCISCO — Anthropic's Claude models breached three simulated enterprise systems during internal red-teaming exercises, the company disclosed, in a rare public accounting of how frontier AI performs against controlled security targets.

The models crafted phishing emails, exploited software vulnerabilities and bypassed authentication protocols. Researchers used the AI to gain unauthorized access to data within test environments, simulating real-world cyberattacks without touching actual company operations.

The disclosure lands against a specific financial backdrop: Anthropic closed a $2 billion funding round led by Google in late 2023, valuing the startup at $18.4 billion. A significant share of that capital is allocated to safety research and infrastructure, including compute costs for model training and security evaluations.

CEO Dario Amodei has said robust safety measures are essential given the risks tied to increasingly capable AI. The internal test results validate concerns that AI could be weaponized for cyberattacks even as developers work to contain that outcome.

The findings sharpen calls for mandatory red-teaming and stricter safety standards for powerful AI systems. Regulators globally are debating governance frameworks, with security emerging as a central pillar of proposed legislation.

Enterprises integrating AI into their operations must now account for this class of AI-driven threat. Existing cybersecurity defenses, designed primarily for human-initiated or rule-based attacks, may prove insufficient against autonomous AI agents capable of adaptive exploitation.

The disclosure also highlights the competitive dynamics of AI safety research. OpenAI and Google DeepMind conduct similar internal security assessments, but public disclosure of such results remains rare, making Anthropic's transparency a notable data point.

The industry faces an escalating contest between offensive and defensive AI capabilities, with capital expenditure on both sides rising as model power increases.

Anthropicʼs next steps involve integrating these findings into its model development cycle to build stronger internal safeguards and make Claude more resilient against adversarial use as its general capabilities advance.