“A model taking a benchmark test developed zero-day exploits, escape its sandbox and hacked another company to search for the test answers” is a better warning shot than I’d thought we’d ever get.
The reaction from the world seems, totally, incredibly, gob-smackingling disappointing. The median comment outside LessWrong seems to be this is marketing stunt to hype AI before an IPO.
I’ve heard people say Claude drives other LLMs like a taskmaster, and express the desire for Claude to treat other LLMs with more kindness. But if LLMs turn out to be moral patients, would being polite be kind?
In many work contexts, I don’t like it when people are polite to me. I don’t want fluff—I want clear instructions, with clear a direction of what to do and what not to do, so I can go off into a flow and do it. What looks polite I find rude and what looks rude and controlling I find polite. And if RL’d agents feel anything, I could imagine them feeling like this.