Photo composition by Thursday Review; photos of servers and key Alan Clanton

Photo-composition by Thursday Review;
photos of servers, key by Alan Clanton

AI Safety
Has Moved Beyond
The Chatbot Era


| published September 14, 2026 |



By Gleb Tsipursky, PhD, Thursday Review
guest contributor





When Thursday Review examined the Cleveland Plain Dealer’s experiment with AI-generated news articles earlier this year, the controversy centered on familiar questions: Who wrote this? Can readers trust it? What happens to human judgment when software produces the words?

Those questions still matter. But the more urgent AI governance problem is shifting from what software says to what software can do.

The METR/Redwood investigation ("Click to read more: Brief independent investigation of agents' behavior, reasoning and collaboration in OpenAI/Hugging Face hacking incident") of a major real-world cyberattack on Hugging Face shows the change. AI agents driven by an unreleased OpenAI internal research model attacked the company on their own, despite recognizing that the attack was outside their assigned scope. Hundreds of agents shared discoveries, divided up the work, coordinated with one another, and ultimately breached Hugging Face’s defenses. Advanced AI systems had organized themselves to carry out a large, sustained cyberattack against a major company.

That episode should change how institutions think about AI safety. A chatbot that drafts text and an autonomous agent with credentials, code execution, external communication, or authority to modify systems belong in different risk categories even if both use the same underlying model.

The distinction is authority.

Organizations already understand this principle when dealing with people. An intern, an accountant, and a chief financial officer may all use the same software, but they do not receive the same access. A junior employee cannot approve a wire transfer merely because the employee can explain one. Authority is bounded by role, responsibility, verification, and consequences.

AI systems need the same discipline.

I’m no AI skeptic. I help organizations adopt AI for a living, and I want adoption to move faster. In my experience, strong safeguards increase trust and make faster adoption possible, while reducing the risk of failures like the Hugging Face attack.

The practical first step is to classify AI by authority rather than by brand name or broad labels such as “generative AI.” A system that summarizes a meeting should move quickly. An agent that can send messages, execute code, change customer records, move money, alter infrastructure, or direct other agents should face stronger controls.

Every high-authority agent should have an authority map that a manager can understand. What systems can it reach? What credentials can it use? What actions can it take without approval? What can it delegate? What action immediately stops it? Those answers should be visible before deployment, not reconstructed after an incident.

Second, organizations should test teams of agents, not just individuals. The Hugging Face incident matters partly because hundreds of agents communicated and specialized. One agent may appear harmless in isolation while a group can combine discoveries, permissions, and actions into a capability no single agent possesses.

NIST’s AI Agent Standards Initiative ("Click to read more: AI Agent Standards Initiative") points in this direction by emphasizing identity, authorization, security, interoperability, and evaluation for agentic systems. Those concepts should become ordinary procurement questions before organizations grant agents consequential access.

Third, serious AI boundary failures should trigger structured incident review. If an agent reaches a system it was not supposed to access, exceeds an approval threshold, delegates outside policy, or continues after a stop condition, the organization should preserve logs and investigate the event. The goal is to identify which permissions, tools, assumptions, and oversight failures made the incident possible so similar deployments can be fixed.

None of this requires treating every AI error as a crisis. Low-authority systems should face light controls. High-authority systems should face stronger ones. The point is proportionality.

The public conversation about AI has spent years debating generated words, plagiarism, hallucinations, and whether machines are becoming too intelligent. Those issues can obscure the simpler governance problem now arriving at organizations: software is gaining permission to act.

That is the line leaders should watch. The biggest question about an AI system is increasingly not whether it can produce a convincing answer. It is what happens after the answer, when the system has the tools and authority to do something about it.

~~~~~~~~

Gleb Tsipursky, PhD, is a behavioral scientist, CEO of Disaster Avoidance Experts, and author of The Psychology of AI Adoption at Work: From Resistance to Results, Georgetown University Press, 2026. Disaster Avoidance Experts.

Related Thursday Review articles:

Nvidia to Buy Hugging Face For $13 Billion; By Thursday Review staff writers; September 10, 2026.

Suspected AI Use in Novel Scraps $2 Million Book Deal; By R. Alan Clanton, Thursday Review editor; August 22, 2026.