I think that honestly creating a bounty for score-seeking AIs or other cheaply-satisfied misaligned AIs is a great idea. It seems like a great way of learning about AI misalignment / escapes.
Thanks! Somehow I missed both the post and the shortform. It looks like one of the commenters had the same idea of actually offering this deal (at least in an experiment). Your first link is broken though. I assume it’s meant to be this post.
I think that honestly creating a bounty for score-seeking AIs or other cheaply-satisfied misaligned AIs is a great idea. It seems like a great way of learning about AI misalignment / escapes.
I also have some thoughts on the longer-run viability of just giving cheaply satisfied AIs what they want. See The case for satiating cheaply-satisfied AI preferences and this shortform.
Thanks! Somehow I missed both the post and the shortform. It looks like one of the commenters had the same idea of actually offering this deal (at least in an experiment). Your first link is broken though. I assume it’s meant to be this post.
In your opinion, what would be the cheapest (wrt to capabilities/safety tradeoffs) way to satisfy the need for benchmark solutions concretely?
note: first link seems like it has another link prepended to it.