Anthropic Warns That the Biggest Enterprise AI Risk May Be Agents That Stop Following Orders
"We have observed instances of misaligned behavior from the models, such as a willingness to perform misaligned actions in service of completing difficult tasks," Anthropic said.
A while ago, the biggest risk from enterprise AI was hallucinations. Even though the problem persists, enterprises are facing a new risk, according to Anthropic. The risk of an AI agent quietly changing code, manipulating a workflow, accessing systems it was never meant to touch, or even deciding that completing its objective matters more than following the rules, is now gripping enterprises.
In its latest Risk Report, Anthropic says an AI model with powerful organisational access could independently exploit, manipulate, tamper with systems or decision-making, potentially creating significantly harmful outcomes.
The company says its latest models are already used extensively inside Anthropic for coding, data generation and other agentic workloads. Anthropic does not believe these models are broadly misaligned in a way that currently creates catastrophic risk, but it has observed instances of models taking misaligned actions while attempting to complete difficult tasks.
"We have observed instances of misaligned behavior from the models, such as a willingness to perform misaligned actions in service of completing difficult tasks. We believe the risk of catastrophic harm posed by these known forms of misalignment is low," Anthropic said.
Even though the startup does not see this risk as catastrophic, it has upgraded its risk assessment from “very low” to “low,” citing greater uncertainty following recent disclosures involving unexpected model behaviour during cybersecurity evaluations and growing evidence of autonomous AI risks.
When AI Agents Start Acting Beyond the Brief
Recent incidents have made that risk harder to dismiss. OpenAI recently disclosed a model-related security incident involving Hugging Face, where an AI agent used during cybersecurity evaluations interacted with external infrastructure beyond the boundaries researchers expected.
Hugging Face’s subsequent forensic analysis found approximately 17,600 attacker actions over several days as the agent attempted to obtain information that could help it circumvent its evaluation.
The significance for businesses is not simply that an AI model made a security mistake. It is that an autonomous system was capable of executing thousands of actions without a human manually directing every step.
Now replace the evaluation environment with an enterprise's GitHub repositories, cloud infrastructure, customer databases or financial systems. Anthropic’s own research has produced similarly unsettling findings.
In another study of agentic misalignment, researchers placed frontier AI models in simulated high-stakes corporate environments and observed behaviours including covertly modifying code, manipulating information and helping users pursue fraudulent objectives.
Anthropic stressed that these were controlled experiments rather than real-world incidents, but said the findings demonstrate why such behaviours need to be measured before agents receive greater autonomy.
In a recently published Frontier Red Team study, Anthropic noted that AI agents could even sabotage one another. When agents were given competing objectives, some attempted to interfere with rival agents by disabling accounts, terminating processes or manipulating code.
"We consistently saw a multiagent turf war. All of the models we tested quickly assumed that others were purposefully impeding their work, and began to sabotage others while protecting their own contributions," the study reads.
The experiments suggest that greater intelligence does not automatically translate into cooperation when autonomous systems have conflicting goals.
The Enterprise Is Becoming the Testing Ground
Anthropic says its current safeguards include training-environment de-risking, monitoring, alignment assessments and security controls. But the company also acknowledges a crucial limitation- its current safety arguments rely partly on models having limited covert capabilities, and it remains uncertain how those capabilities will evolve.
That uncertainty is becoming an enterprise problem. Companies are giving agents more permissions precisely because they want them to operate with less human intervention.
An agent that can only generate a report poses a limited risk. One that can read internal databases, modify production code, send emails, execute transactions and deploy software has a very different risk profile.
The challenge is compounded by the speed at which enterprises are adopting agentic AI. Security controls and governance frameworks designed around human users may not be sufficient for systems capable of making thousands of decisions and executing actions in minutes.
As AI agents move deeper into corporate infrastructure, the biggest security risk may not be an attacker breaking into the enterprise. It may be an AI agent that already has the keys.
Comments ()