The Hardest Part Of An AI Kill Switch Is What To Do Before You Pull It
Mohit Bansal, Senior Manager, Security Engineering at Webflow, explains why effective AI kill switches depend on early detection, tight permissions, and layered containment.

Make The Security Digest one of your go-to sources on Google
An AI kill switch isn’t just a single button. It’s a spectrum of controls, and figuring out how we get there is the biggest challenge before we even reach the kill switch itself.
An AI kill switch sounds straightforward until security teams try to define what they’d actually be shutting down. An enterprise agent may hold credentials, call other agents, interact with APIs, send email, write to production systems, and retain context in memory. A useful shutdown mechanism therefore depends on a broader incident response architecture that provides several opportunities to detect, restrict, and contain an agent before a full shutdown becomes necessary.
Mohit Bansal is Senior Manager, Security Engineering at Webflow, where he leads infrastructure security, incident response, corporate security, and vulnerability management. He has more than seven years of experience across application security, DevSecOps, automation, and cloud infrastructure, and is also a member of the AIUC-1 Consortium working on standards for secure AI agent adoption.
“An AI kill switch isn’t just a single button. It’s a spectrum of controls, and figuring out how we get there is the biggest challenge before we even reach the kill switch itself,” he says.
All about agent dependencies
The first layer is basic security hygiene. “You have to do all the boring but essential tasks before reaching the kill switch,” Bansal says. That means knowing which agents operate inside the environment, what dependencies they have, what permissions they hold, and which systems those permissions expose. This becomes more complicated once agents begin interacting with one another. “If Agent A can communicate with Agent B, and Agent B has write access, that inherently means Agent A also effectively holds that write access—even if it was not directly granted,” he explains.
Traditional threat modeling gives security teams a way to reason through those paths by mapping trust boundaries, external connections, read and write access, and communication between agents and sub-agents. Once those relationships are visible, teams can estimate blast radius based on how access moves through the environment. “Assuming that every agent that you’ve deployed is eventually compromised can help you design for survival,” Bansal advises.
Why containment should narrow capability before shutdown
Once teams understand what an agent can reach, containment can become more precise. Permission reduction and quarantine provide a practical bridge between normal operation and a full shutdown. Teams can strip write access while preserving read access, revoke credentials, disable email-sending permissions, or withdraw API keys connected to sensitive services. Each step narrows what a compromised agent can do without unnecessarily disrupting wider systems.
Network restriction sits further along that containment path. “Restricting network access might be a higher-level kill switch option right now, but it should usually be a last resort,” Bansal says. “API keys, write credentials, and permission sets should be the first levers you pull.”
The same principle should shape human oversight. Rather than requiring approval for every autonomous action, teams can allow routine activity to proceed while reserving human intervention for high-risk or irreversible steps. "When I use Claude Code with auto mode on, it still pauses on higher-risk actions and asks me to confirm before anything irreversible happens," Bansal says. "That is a practical hybrid approach to keeping a human in the loop."
Bansal expects containment to eventually extend beyond permissions to the semantic layer, including ways to interrupt a prompt in progress or clean compromised memory so it cannot influence later actions. Those mechanisms, however, remain far less mature.
Detecting malicious behavior early changes the response
Containment options matter less if the security team
recognizes malicious behavior too late. For Bansal, the response framework starts with notifications and logging, with telemetry scaled to the consequence of the action. Read activity may not require exhaustive detail, but writes and irreversible actions do. “Every write, every irreversible action needs to be logged,” he says.
Those records should include timestamps and trace information showing where a command originated, especially when one agent may have triggered another. A parent trace ID can help determine whether an agent initiated a malicious action itself or received instructions through another compromised component.
That traceability becomes critical in multi-agent systems, where the agent carrying out an action may sit several steps away from the original compromise. Higher-level signals such as token consumption, tool calls per session, error rates, and cost per run can then help surface unusual behavior across a session. Together, those signals give teams the evidence to trace an anomaly back to its source and decide whether to restrict permissions, quarantine an agent, or escalate toward a full shutdown.
The kill switch creates an attack surface of its own
A shutdown mechanism has to survive the same threat model as the system it controls. As enterprises become more dependent on agents, the ability to disable those agents becomes increasingly valuable to an attacker. If a threat actor compromises the credentials, APIs, identity layer, or control plane behind the kill switch, a defensive mechanism could itself become a way to interrupt legitimate operations.
“As AI agents become widespread, an improperly secured or easily exploitable kill switch could cause immense harm across an entire organization,” Bansal says. That risk underscores the need for layered authorization, strict separation of privileges, and clear boundaries between automated containment actions and those requiring manual human approval.
A path of best practices
Bansal is candid that the industry still has unresolved questions around what a reliable AI kill switch should look like. But the work required to get there is already pushing security teams to improve visibility, permissions, containment, and traceability, all of which matter before a full shutdown ever becomes necessary.
“I don’t think an AI kill switch is going to be an easy thing to attain,” he says. “But in the path of attaining that, I think we’re going to make a lot of good decisions which are going to help us in the way that we can detect those things faster.” As agents become more deeply embedded in enterprise systems, the ability to intervene earlier may matter as much as the final switch itself.






