I have thought of a similar idea: “philosophical landmines” (PL) to stop unfriendly AI. PL are tasks which a simple in formulation but could halt an AI as they require infinite amount of computation to solve. Examples include Buridan ass problem, the problem if AI is real or just possible, the problem of being in simulation or not, other anthropic riddles and pascal mugging-like stuff.
Best such problems should be not published as they could be used as our last defence against UFAI.
I think that AI capable of being nerd-sniped by these landmines will probably be nerd-sniped by them (or other ones we haven’t thought of) on its own without our help. The kind of AI that I find more worrying (and more plausible) is the kind that isn’t significantly impeded by these landmines.
Yes, landmines is the last level of defence, which have very low probability to work (like 0.1 per cent). However, If AI is stable to all possible philosophical landmines, it is a very stable agent and has higher chances to keep its alignment and do not fail catastrophically.
Thanks for the comment. +1 to it. I also agree that this is an interesting concept: using Achilles Heels as containment measures. There is a discussion related to this on page 15 of the paper. In short, I think that this is possible and useful for some achilles heels and would be a cumbersome containment measure for others which could be accomplished more simply via bribes of reward.
I have thought of a similar idea: “philosophical landmines” (PL) to stop unfriendly AI. PL are tasks which a simple in formulation but could halt an AI as they require infinite amount of computation to solve. Examples include Buridan ass problem, the problem if AI is real or just possible, the problem of being in simulation or not, other anthropic riddles and pascal mugging-like stuff.
Best such problems should be not published as they could be used as our last defence against UFAI.
I think that AI capable of being nerd-sniped by these landmines will probably be nerd-sniped by them (or other ones we haven’t thought of) on its own without our help. The kind of AI that I find more worrying (and more plausible) is the kind that isn’t significantly impeded by these landmines.
Yes, landmines is the last level of defence, which have very low probability to work (like 0.1 per cent). However, If AI is stable to all possible philosophical landmines, it is a very stable agent and has higher chances to keep its alignment and do not fail catastrophically.
Thanks for the comment. +1 to it. I also agree that this is an interesting concept: using Achilles Heels as containment measures. There is a discussion related to this on page 15 of the paper. In short, I think that this is possible and useful for some achilles heels and would be a cumbersome containment measure for others which could be accomplished more simply via bribes of reward.