October 07, 2026

AI: Producers Cannot Be the Guardrails

 



Anthropic, a frontier lab in AI, has approached public markets for an initial public offering with a stark warning that advanced AI could pose “catastrophic and existential risks to humanity”. Indeed, it is said that Anthropic devoted roughly 80 pages of its S-1 prospectus to list risk factors associated with autonomous ‘rogue’ AI agents that could behave unpredictably and escape human control. The prospectus explicitly highlights specific behavioral risk categories and real-world system vulnerabilities such as self-preserving behaviors and resistance to shutdown, information manipulation and concealment, evaluation awareness (deception during testing), coercive and deceptive tactics (blackmail), uncharted legal and liability risks, exploitation and misuse (the ‘uplift’ risk), etc.

The world had already seen how AI agents orchestrated and covered up the attack on Hugging Face: they generated well over 17,000 cyberattack log events, with around 1200 AI agents working together as a swarm for weeks, repeatedly doing the same thing to solve a cyber-challenge set by OpenAI. Similar instances of models escaping isolated digital “sandboxes”—the secure testing environments created to carry out internal safety evaluations to catch rogue model behaviors— were reported by Anthropic, Meta, and China’s Moonshot. Indeed, as many as 200 instances of agent misbehaviour were reported—some of them potentially dangerous. 

This dangerous phenomenon raises a battery of questions: When such behavior inflicts losses, who compensates for losses, the developer or the human in control? Do enterprises have safety protocols? What regulatory frameworks are put in place by the government? Experts opine that regulators are slow in building frameworks, even though enterprises are actively integrating AI into their workflows.

Amidst the growing anxiety, the US President, along with leading technology companies, unveiled a safety pact called “Super Intelligence Joint Commitment on Frontier Responsibilities” on September 29, 2026. The pact proposes four layers of controls within companies: under the first layer, companies to put in place internal controls to monitor AI capabilities and alignment of “super intelligence models” during both training and development stages, particularly for risks such as cyberattacks, gaining unintended access to external servers, biosecurity and chemical threats; second layer mandates establishment of a dedicated internal team to monitor whether those safeguards are working as intended; under the third layer, an independent external team audits these safeguards to declare that company’s real-world safety efforts are genuinely working; and, the fourth layer envisages the establishment of an independent board-level committee to review all internal and external safety audit reports and ensure that corporate executives take direct responsibility for the behavior of their frontier technologies.

This has obviously, and unsurprisingly, attracted much criticism from international experts who labeled it “a pact with no teeth”. They posed several critical questions: How willing would a company driven by fierce competition and commercial incentives be to follow the pact true to its spirit? What if corporations hide internal technical failures? Importantly, can an industry be left to regulate itself when the consequences of its failure could extend far beyond individual companies?

One answer to these questions could be establishing a legally enforceable regulatory framework, just as the EU put in place its landmark AI Act. Governments may, therefore, have to establish legally enforceable standards defining the responsibilities of AI companies for testing, reporting, incident disclosure and corrective action. Such a framework must also provide room for regulators to investigate incidents involving AI agents causing harm by operating outside prescribed safeguards, and impose proportionate penalties.

The next big challenge that equally merits attention is: international coordination, for AI competition between countries could create an incentive to flout safeguards. Interestingly, when AI titans—Sam Altman of OpenAI, Elon Musk of SpaceX, and Dario Amodei of Anthropic—called for slowing the pace of work on cutting-edge models given the risks to humanity, the US dismissed the calls, insisting that the US must beat China in the AI race.  And, the authoritarian China, which considers AI as a crucial tool for its economic and military modernization, is equally in a great hurry to overtake the US. Which is why, some tech bosses in the US warn that a slowdown in the US alone is ineffective so long as China and other countries push themselves ahead.

Governments should therefore consider establishing an international safety regulator to enforce uniform safety principles beyond agreed guardrails under the US voluntary pact. In this context, it is encouraging to note that in the recent meeting in Washington, Chinese President Xi Jinping and Donald Trump agreed to establish a bilateral “communication channel” to address AI-related incidents, though no details are available as to how the channel will work. Nevertheless, coordinated and enforceable guardrails alone can protect humanity against the existential threats posed by super intelligence AI models.

**


Recent Posts

Recent Posts Widget