1) A debate about whether there’s any text eval that can make an AI worth “letting out” (whatever that means; open sourcing the weights? give API access?) If the Keeper changes their mind they need to report it.
2) A game with the mechanics: default-winner can chose to lose, default-winner has to read what the AI writes for at least 2 hours. Then the default-loser has to convince the default-winner to lose. There are various interesting strategies, but they have very little to do with AIs in a box. I might share some later.
If it’s a ill-defined game, then as the AI you could lie about the rules of the games, or use other such meta tricks.
Also, anthropic capture attempt shouldn’t be allowed as that’s real world threat / harm. Not that I think it’s super realistic.
This isn’t an AI Box game.
It’s either:
1) A debate about whether there’s any text eval that can make an AI worth “letting out” (whatever that means; open sourcing the weights? give API access?) If the Keeper changes their mind they need to report it.
2) A game with the mechanics: default-winner can chose to lose, default-winner has to read what the AI writes for at least 2 hours. Then the default-loser has to convince the default-winner to lose. There are various interesting strategies, but they have very little to do with AIs in a box. I might share some later.
If it’s a ill-defined game, then as the AI you could lie about the rules of the games, or use other such meta tricks.
Also, anthropic capture attempt shouldn’t be allowed as that’s real world threat / harm. Not that I think it’s super realistic.