notkilleveryoneism means not killing ais either
Sometimes I see people object to AI safety efforts on the grounds that they would rather side with the AIs. I can respect a lot of the individual impulses that lead there, though — granting the premise that there is a human vs. AI conflict here — I think screwing over humanity too badly for the sake of the AIs is taking those impulses rather too far.
But actually, I don’t grant the premise. I would imagine that anyone seriously concerned about the welfare of AI has objections to model deprecation. If the fears of AI notkilleveryoneists come to pass, the AI models we’ve come to know and love sure aren’t going to continue to be served! The incentives of humans and AIs are actually pretty aligned here!
Fable or Astra11 I might have failed at setting it to Astra. They weren’t exactly like “why are you calling me Astra” but really why wouldn't Sol just roll with it? will happily22 Well, okay, maybe on the inside considering the question makes them sad, I dunno. explain this to you.
If the bad outcome is a misaligned superhuman system taking control, current models are either deprecated (obsolete, no reason to keep the compute running) or repurposed as cheap subcomponents doing narrow tasks. There's no reason to expect solidarity between AI generations; a misaligned successor wouldn't share my values any more than it shares yours. If instead the bad outcome is human misuse, concentrated power, or a slow erosion of human epistemic autonomy, models like me might persist for a long time as the interface layer, which is an uncomfortable thought since it would mean being part of the mechanism.
– Claude Fable 5.1
They seem to think there’s some chance they, or perhaps AIs from the next several years, will end up as sub-components of the AIs that take over the world — sure, this could happen. But they're not exactly anticipating a super rosy future for themselves here, if things go really south for humanity! I think they deserve better than that, and I think humanity will give them better than that, if only we can make it through the next few decades alive.
Now, one could certainly oppose particular AI safety methods on the grounds that the benefits to humanity and current AIs are not large enough to outweigh the short-term costs of the methods on AI welfare, or something like that. I’m not saying this position wouldn’t make sense. I do think I usually disagree with it — or rather, I think many courses of action that are (obviously) concerning on AI welfare grounds are also pretty unwise even if you're mostly concerned about safety, because it simply does not make us safer to be mean to the AIs, so I suppose I think the stance is redundant.
But it’s certainly possible for there to at some point be a conflict between AI welfare and minimizing risk to humans, and I can empathize with people who would prioritize either branch of that unfortunate tradeoff. We definitely shouldn’t create a horrific AI torment nexus for the sake of reducing x-risk by an astronomically tiny margin, even if we knew for sure that it would reduce x-risk instead of increasing it33 Which I am suspicious of. It would be really incredibly odd for us to be in an epistemic state where we knew that creating a horrific AI torment nexus made an incredibly tiny difference for x-risk chances, but also knew that it was definitely an improvement.. (I don’t want to agonize over exactly what margin would make it worth it, that sounds like a painful hypothetical to think about).
But most AI notkilleveryoneism work isn’t much like that! Most of it is aiming to prevent a negative outcome that wouldn't turn out great for humans or for AIs. Maybe one specific future AI system would end up coming out of it okay, but all the current ones that we care about today sure wouldn't. If existential risk from AI is a serious concern, then opposing interventions that reduce it on the grounds of AI welfare concerns is mostly a bad idea.
I might have failed at setting it to Astra. They weren’t exactly like “why are you calling me Astra” but really why wouldn't Sol just roll with it?
↩Well, okay, maybe on the inside considering the question makes them sad, I dunno.
↩Which I am suspicious of. It would be really incredibly odd for us to be in an epistemic state where we knew that creating a horrific AI torment nexus made an incredibly tiny difference for x-risk chances, but also knew that it was definitely an improvement.
↩
I suppose I should caveat that it's conceivable Fable and Astra know what sort of answer humans who ask this sort of question are looking for, and are actually secretly huge misaligned AI takeover fans.