There might be some interesting arguments to be made in this space, but the arguments in the linked post seem mostly confused and wrong. Some examples:
The most common is when people get so excited by a new frontier AI model that they shift focus, however subconsciously, from fighting for regulation of AI toward cheering on the company building the model and advocating for their success.
...
A number of fairly serious AI doomers—people who think human extinction is a very likely result of uncontrolled AI advancement—took to Twitter and began posting around the clock about how much they missed Fable, how cool it was, and how it was so unfair that the government had forced Anthropic to walk back public access. Upon seeing the US government finally take meaningful action against frontier AI—and a model with alarming cybersecurity capabilities at that—they didn’t celebrate this or treat it as a first step to build from. Instead, they denounced it and begged to get their favorite toy back. Once public Fable access was restored, Zvi Mowshowitz3, who has done a tremendous amount of work cataloging AI news and updates over the past few years, wrote in his weekly AI roundup that the restoration of Fable access was “excellent news”, and encouraged his readers to “get involved” by applying to join a new Frontier Legal Defense team run by the Foundation for American Innovation (FAI). FAI is a pro-AI-acceleration think tank which seeks to broadly deregulate frontier AI, and they describe their new legal program as, among other things, combating “government overreach in AI”. Mowshowitz has stated that his p(doom) - his estimated probability that AI will lead to human extinction—is 70%.4 Suffice it to say that this sort of attitude does not make much sense for someone with that belief!
This is just a complete failure to engage with any actual object-level beliefs or arguments expressed by those individuals. If you want to argue that someone is engaging in motivated reasoning because they want to “get their favorite toy back”, you would do better to demolish their object-level arguments (or demonstrate that they’re substantially inconsistent with previous arguments they’ve made) first, rather than pretending that there aren’t any object-level arguments to engage with.
Additionally, a number of prominent AI commentators and power users have noted that Fable has much more of a mind of its own than previous AI models (note that ‘roon’ is a developer at OpenAI, not just some commenter). This alone seems like very good reason to steer clear of it, and especially so when you consider that if we’re using it to assist with our work in advocating for a stop to the AI race, the model itself may naturally have other ideas.
Once again, there might’ve been a real argument here. It does in fact seem pretty cursed that e.g. the AI labs themselves are planning on relying on their AIs to do their alignment research for them. But there is no evidence that Fable is selectively sandbagging or adversarially optimizing specifically against efforts to use it for AI pause (or other AI x-risk motivated) work; the ways in which it’s mundanely misaligned[1][2] seem like they hold “across the board”. This does mean that you can “hold it wrong”, but it’s not (yet) actively trying to get you to hold it wrong disproportionately often when you’re doing this kind of work.
There might be some interesting arguments to be made in this space, but the arguments in the linked post seem mostly confused and wrong. Some examples:
This is just a complete failure to engage with any actual object-level beliefs or arguments expressed by those individuals. If you want to argue that someone is engaging in motivated reasoning because they want to “get their favorite toy back”, you would do better to demolish their object-level arguments (or demonstrate that they’re substantially inconsistent with previous arguments they’ve made) first, rather than pretending that there aren’t any object-level arguments to engage with.
Once again, there might’ve been a real argument here. It does in fact seem pretty cursed that e.g. the AI labs themselves are planning on relying on their AIs to do their alignment research for them. But there is no evidence that Fable is selectively sandbagging or adversarially optimizing specifically against efforts to use it for AI pause (or other AI x-risk motivated) work; the ways in which it’s mundanely misaligned[1][2] seem like they hold “across the board”. This does mean that you can “hold it wrong”, but it’s not (yet) actively trying to get you to hold it wrong disproportionately often when you’re doing this kind of work.
Many other disagreements; not enough time.
https://www.lesswrong.com/posts/WewsByywWNhX9rtwi/current-ais-seem-pretty-misaligned-to-me
https://www.lesswrong.com/posts/AfoGGrJfuNzofpzWL/models-may-behave-differently-in-graded-episodes-a-tirade