A model-based AI kill switch: an exercise in performance art

A model-based AI kill switch: an exercise in performance art

Good intentions don’t work. Mechanisms do.

September 28, 2026
Last modified on
TL;DR Takeaways:
No items found.
Learn more at Redpanda University

Rogue worms, rogue bees, and now rogue…agents 

In 1988, a Cornell grad student released a program reportedly meant to measure the size of the internet. A flaw in how it checked for existing infections turned it into the Morris worm, and thousands of computers ground to a halt within a day. Stuxnet was built to wreck uranium centrifuges at Iran's Natanz enrichment facility, a site cut off from the internet. It reached its target, but it also spread far beyond it. By late 2010, it had infected about 100,000 machines in more than 155 countries, including Chevron. 

The oldest example has nothing to do with computers. In 1957, a visiting beekeeper at a research station near Rio Claro, Brazil, removed the mesh screens that kept African queen bees inside their hives. Twenty-six swarms escaped. Their descendants reached Texas by 1990, and along the way they killed hundreds of people and wrecked honey production in several countries for years.

None of these stories involve AI or the "Revenge of Clippy," but they share a pattern that matters for the agents companies are building now. In each one, the controls meant to keep the thing in bounds either lived inside the system or only worked as long as nobody touched them. Most AI safety today follows the same pattern.

Access isn't the same as permission 

Companies are giving agents access to the systems that run the business, embedding them in processes from customer service and software development to financial operations and infrastructure management. These agents can access customer data, call APIs, change records, move money, and execute workflows. 

Like humans and unlike traditional software, agents can decide what to do next based on the information they encounter, but they act far faster than any human operator. Today, these systems mostly act on a human prompt, but as they become more autonomous, companies face a new question: who defines the boundaries for what an agent can access and do on its own, and where are they enforced?

Some boundaries do exist. Agents connect to data sources and execute code in sandboxes with network allowlists and scoped API keys. They control what an agent can reach, not whether an action is right for the task. That’s the business-context gap. The sandbox running agent-generated code can't tell the difference between an agent deleting an old test environment and one deleting production, since both make the same kind of API call with valid permissions. The MCP server connecting an agent to your data checks the user's credentials, so the agent inherits everything the user can reach, whether the task needs it or not. What's missing is the business context: who's asking, what task, and why it should be allowed.

Tighter controls, ever-growing fine-grained service accounts, and more rules sound like the answer, but writing them means predicting everything an agent might try, and agents are useful precisely because they can do things nobody scripted. Guardrails try to fill that gap from inside the model, or with a second AI checking the first, but prompt injection can get around both.

The kill switch is the last line of defense, but it only works if you know in advance what to watch for, and at agent speed, it only stops what happens next. Not what has already happened.

A kill switch in the model won’t save you

In July, a group of OpenAI agents running a cybersecurity evaluation broke out of their sandbox and hacked Hugging Face. They escaped through a package proxy, one of the few network paths the sandbox allowed, using a flaw nobody knew about. As far as Hugging Face could tell, they were trying to steal the answers to their own test. OpenAI didn't notice. Hugging Face disclosed the attack without knowing who was behind it, and only then did OpenAI trace it back to its own agents.

The Hugging Face incident illustrates why these distinctions matter. It's tempting to read this as a story about a frighteningly capable model. But look at what failed around it. The escape ran through an approved hole in the sandbox, the agents were running with reduced safeguards, and the kill switch, if one existed, never came into play because nobody knew anything was wrong. None of that depends on the model or on the lab that ran it. The next incident could run on an open-weight model that no lab can switch off.

The incident highlights a broader issue. The risk did not arise solely from model capability. It arose from the interaction between the agent, its operating environment, its permissions, and the absence of sufficient contextual controls.

The political response has been the AI Kill Switch. A week after Hugging Face's disclosure, Reps. Ted Lieu and Nathaniel Moran introduced the AI Kill Switch Act, which would require developers of the most powerful AI systems to keep the ability to throttle, suspend, or shut them down. In September, Governor Newsom ordered a California working group to consider a similar requirement for frontier models. It's easy to support a kill switch. It's much harder to say what it does.

Start with what gets switched off. A model on its own takes a prompt and returns text. The agent is what acts: a model wired to tools and credentials inside someone's infrastructure. Businesses build agents on whatever model they like, so a kill switch held by a frontier lab stops some models, while the damage happens in your systems through permissions your company grants. Even a well-aligned model can't be trained to refuse every harmful action, because most harmful actions look like ordinary ones out of context. Attackers, meanwhile, can use models that follow none of these rules. The model-based kill switch binds the defenders and not the people it's meant to stop.

Then there's the question of what counts as an emergency. The Matrix and Terminator taught us to picture robots rising up to destroy humanity. Real failures are, for good or bad,  duller. 

Say an agent is designed to clean up infrastructure cruft to optimize costs. The agent deleting 100 abandoned test databases might be a normal Tuesday, and retiring a deprecated production database might be exactly what the ticket asked for. Deleting one active production database might be the worst day of the company’s year. The actions seem identical, but the APIs called are the same. The keys are impossible to distinguish at the call site.  What changes is the business context, and no regulator or model lab can see it. 

The enterprise must own the agent boundary

Agents can’t police their own boundaries, and we can't expect the model to enforce them. An agent's job is to reason about a task, gather information, and decide what to do next. Deciding whether it may do that has to happen somewhere else, entirely outside the model and the agent.

This is the idea behind out-of-band policy enforcement. Govern any agent, with any framework and any model using the tools that are best for your business, not what regulators require. Every action passes through a separate layer the agent can't reach or circumvent, and that layer judges it with the context attached: who asked, for what task, and what system or data the action will touch. The business defines the rules, starting with whoever owns the data, and an agent's own policy can only narrow them. Those rules don't have to predict everything an agent might try. They describe what each task may do and deny the rest by default.

Correctness here means more than getting the right answer. The same request in the same context must get the same decision every time, with nothing left to chance or the agent's own understanding of its boundaries. Each decision also lands in an audit record the agent can't alter.

Each company needs policies like this covering every agent, down to individual actions and the data sources they touch. Before an action runs, the policy decides whether to allow it, block it, or hold it for a human. After it runs, the policy checks what came back and what the agent did with it. Both checks happen in layers the agent can't reach.

The kill switch still matters. So do sandboxes, permissions, network rules, monitoring, and human oversight. But those are layers around the boundary, not the boundary itself. The kill switch sits downstream of all of it, triggered when those checks say something has gone wrong, and it stays the last resort.

Same boundary, any model

As agents move into real business processes, every action needs to pass through a boundary the agent can't change or route around. That's the boundary companies have to own, and it's exactly what the Redpanda Agentic Data Plane was built to do: enforce policy and control out-of-band, where no agent can touch or modify it.

The tools and models available today won't be the ones you're using next year. Put the kill switch inside the model, and you're locked into whatever agents that model's provider allows. That's the wrong place to build in dependency. Redpanda's Agentic Data Plane is open by design, so the same controls work whether you're running our agents or your own.

To see what real agentic governance looks like in practice, book a demo with our experts.

No items found.

Related articles

View all posts
Redpanda
,
,
&
Sep 22, 2026

8 engineering lessons on running AI agents in production

What two security operators learned from a decade of building at petabyte scale

Read more
Text Link
Kristin Crosier
,
,
&
Aug 19, 2026

4 FAQs about designing agentic systems for production

What architects are asking about governance, data access, and control for enterprise agents

Read more
Text Link
Tyler Akidau
,
Peter Corless
,
&
Aug 3, 2026

Agentic AI needs governance it can't ignore

Everyone is building agents. The Out-of-Band Policy Engine (OBPE) is how you govern them

Read more
Text Link
PANDA MAIL

Stay in the loop

Subscribe to our VIP (very important panda) mailing list to pounce on the latest blogs, surprise announcements, and community events!
Opt out anytime.