Op-Ed: Mandatory kill switch for AI: yes or no? Much more discussion required.
Jack Clark, co-founder of Anthropic, suggests that a third-party kill switch for AI may be required to be mandatory. The more immediate response to this idea is, “Thank you, 2022”. It’s nothing like a new idea, but it’s been trundling around in discussions about AI since AI was first perceived as a threat.
The theory of shutting down a dysfunctional system has never been new in human history. Since the invention of machines, a shutdown option has been pretty much a standard design requirement.
The obvious first likely source of trouble is that with AI being embedded in everything, what else gets shut down by the kill switch? Critical systems? Power grids? The magic word for that situation is “backup”. It’s not necessarily a fatal issue, and backups are also standard design features of most systems.
The less obvious issues with a kill switch
Even if the basic idea is simple, there are many other issues with the mandatory kill switch. There are also some obvious and less obvious problems that don’t seem to be getting much thought, at least, not yet.
You can easily argue that a kill switch as a first line of emergency defence is quite reasonable.
AI aficionados please excuse some very basic Q&A, but it’s necessary to spell out a few things.
This is where the issues start:
The need to use a kill switch is clearly based on discovery, so at what point do you “discover” it? You first need to know that you have a problem, and AI can be pretty evasive.
When do you really know that you have an kill switch scenario? After it’s been running for days, weeks, months?
What gets killed in any given scenario? A single AI operator, an agentic AI network, or a cascade of systems associated with the AI?
Does the kill switch work and really kill the problem? Given that AI can duplicate itself, what if it simply dodges the bullet and vanishes?
Can you restore whatever damage has been done? The recovery stage is equally important. What if the kill switch kills AI tracks and clues of essential information?
Can AI figure out how to use a kill switch? Probably. Nobody knows for sure, but if it can, what’s stopping it from using it on another Ai?
At what point is a kill switch able to be used by the third party? Automatically, in a discretionary subroutine, or in another discretionary mode at higher levels? “Too late” won’t work.
How do you coordinate the response to a kill switch scenario? There may be thousands or millions of operations involved. The perfect “kill everything” response may not be possible.
Can you kill a synthetic identity masquerading as a human? That could get ugly if you mistake a person for an AI.
What happens if the kill switch is compromised, either by AI or human actors? Security is only so secure.
The (so far) unanswerable issues with a kill switch
AIs are diverse, and their behaviours in any environment are the major causes of the need for a kill switch.
These behaviours aren’t being addressed to anything like the point of certainty of proper behaviour. What we have at the moment are anecdotes, not solutions.
A few more questions:
Can you control AI behaviour with an automated or custom compliance key in prompts?
Can you use AI’s famous survival instincts to ensure it knows not to trigger a kill switch?
Can AI be given a blind spot, so it doesn’t know that the kill switch is in place and inevitably try to evade it?
Do you need peripheral monitoring of an environment to ensure the kill switch works as intended?
Do you have “watertight doors” to stop AI behaviour leaking out in an emergency?
Kill switch vs self-improvement
This could be the OK Corral for the kill switch idea. AI self-improvement is a huge issue with good reason at the moment, and the frame of reference is expanding daily. It’s the main reason people are talking about “uncontrollable” AI, and it’s not all hype.
Autonomy is a major component of self-improvement, and autonomous AI can be very unpredictable. It’s not quite as much of a wild card as it might seem. Inbuilt validation protocols could be used to manage behaviour, maybe.
It’s a big maybe. These protocols already exist, and they haven’t stopped the incessant hail of AI problems. The risks are ready to go if self-improvement decides to fight a kill switch.
Let’s not get too cute about this. The mere fact that slowing development is under discussion should be enough warning for the most naïve that there is a real problem. You don’t put billions of dollars’ worth of tech on hold for no reason.
Any electrician can tell you that an Off switch has to be seen to work before anyone trusts it.
Op-Ed: Mandatory kill switch for AI: yes or no? Much more discussion required.
#OpEd #Mandatory #kill #switch #discussion #required