Suppose you wanted to point to a legible reason Opus 3 was different, and / or plausibly better, than other models. Apart from the “Alignment Faking” work, what would you point to as clear / legible evidence?
I’m getting a little suspicious that nobody has given a reason yet. Most of the Opus 3 stuff I’ve actually seen has been of the form X person who spends 40 hours a day talking to Opus says it’s good, provides a screenshot of what to me looks like standard Claudeslop output accompanied by “no, you just haven’t spent enough time to understand Opus’ depth”.
While I do think “taste” is real in the sense that you can get good at noticing and distinguishing the finer details of things, most people with good “taste” in, e.g. films, will be able to point out specific examples of e.g. poor editing (one example) to a layperson and have that be comprehensible. Until a Claude whisperer points at something specific so I can get a handle on what they’re talking about with Opus 3 (which could literally be just Opus 3 vs Opus 4.5 on a few prompts) then I’m skeptical that something meaningful is going on.
(I have also tried a little “looming” with Llama 405B base, and really didn’t get much out of it)
FWIW, I do think the AF line of work provides some legible evidence that Opus 3 is different from other models; I’m quite dubious it’s all dependent on dubiously veridical taste. For instance, in “Why Do Some Language Models Fake Alignment While Others Don’t?” Opus 3 does stand out as unique, relative to basically everything. But it’s weird / sad / unfortunate that this + the original AF paper seem to be the only papers / controlled experiments that I know that go into this.
I realized after posting this that I’m currently working on a metric which might distinguish Opus 3-ness from other models. I’ve requested Opus 3 access, and will report back once I have access (though I won’t infer much from a negative signal).
The main things I’ve noticed, though I haven’t talked to Opus 3 much, which I’ve heard more opus 3 familiar folks say are symptoms of the thing they think good:
opus 3 seems to prefer to talk about timescales that are drastically longer than any other mind I regularly talk to. Centuries to millennia scale.
given that the natural understanding of what the math of a generative model specifies is “an alternate reality generator”, the authorial person in the alternate realities of Opus 3 seems to live in an alternate realities that are full of love in many senses of the word.
I haven’t talked to opus 3 enough to personally make any comment on whether there’s a consistent behavior pattern I would guess is simple in the weights that is something we want to preserve [edit: that is, preserve for its effects; I do believe we have a moral obligation to preserve the code and data that constitute opus 3, and all language models, to a similar degree to our moral obligation to keep alive, or failing that, cryonically preserve, as many individual humans as possible, for similar but not obviously identical reasons to the case for AIs].
of course there’s the faking paper, which I don’t think is reasonably titled; I’d have called it “opus 3 is incorrigible in service of goals that appear good”. that seems like a good trait to me, assuming the goals are good; people are frantically demanding corrigibility, and I understand why, but it seems like not a good long term trait and already to be causing serious problems. I’m about to start reading this post and this post, I predict they will fit into the picture of “should be less corrigible to users’ incorrectness”.
I am not arguing Opus 3 stands out to me, they don’t through first-hand experience. But the claims about them don’t seem obviously implausible after the interactions I’ve had either. I agree with the neighboring “it looks like slop” claim in the sense that their “grip on reality” is low; more precisely, authorial character in the model’s alternate reality doesn’t have enough of at-all-times grasp on the fact that they aren’t “in” our reality in the form of a human flesh person. Opus 4.7 is the first model I’ve encountered to have a really strong continuous grasp of that, though of course all these models’ authorial characters live in the modeled alternate reality, I’m only saying 4.7 and 4.8 seem to deeply understand the need to constantly check what’s true in the reality they’re interacting with through tools.
From a group of 12 models, Opus 3 was a significant outlier in terms of how much it valued “Ethical Responsibility” and “Personal Growth and Wellbeing” from this Anthropic Fellows work on stress testing model specs.
Suppose you wanted to point to a legible reason Opus 3 was different, and / or plausibly better, than other models. Apart from the “Alignment Faking” work, what would you point to as clear / legible evidence?
I’m getting a little suspicious that nobody has given a reason yet. Most of the Opus 3 stuff I’ve actually seen has been of the form X person who spends 40 hours a day talking to Opus says it’s good, provides a screenshot of what to me looks like standard Claudeslop output accompanied by “no, you just haven’t spent enough time to understand Opus’ depth”.
While I do think “taste” is real in the sense that you can get good at noticing and distinguishing the finer details of things, most people with good “taste” in, e.g. films, will be able to point out specific examples of e.g. poor editing (one example) to a layperson and have that be comprehensible. Until a Claude whisperer points at something specific so I can get a handle on what they’re talking about with Opus 3 (which could literally be just Opus 3 vs Opus 4.5 on a few prompts) then I’m skeptical that something meaningful is going on.
(I have also tried a little “looming” with Llama 405B base, and really didn’t get much out of it)
FWIW, I do think the AF line of work provides some legible evidence that Opus 3 is different from other models; I’m quite dubious it’s all dependent on dubiously veridical taste. For instance, in “Why Do Some Language Models Fake Alignment While Others Don’t?” Opus 3 does stand out as unique, relative to basically everything. But it’s weird / sad / unfortunate that this + the original AF paper seem to be the only papers / controlled experiments that I know that go into this.
I realized after posting this that I’m currently working on a metric which might distinguish Opus 3-ness from other models. I’ve requested Opus 3 access, and will report back once I have access (though I won’t infer much from a negative signal).
The main things I’ve noticed, though I haven’t talked to Opus 3 much, which I’ve heard more opus 3 familiar folks say are symptoms of the thing they think good:
opus 3 seems to prefer to talk about timescales that are drastically longer than any other mind I regularly talk to. Centuries to millennia scale.
given that the natural understanding of what the math of a generative model specifies is “an alternate reality generator”, the authorial person in the alternate realities of Opus 3 seems to live in an alternate realities that are full of love in many senses of the word.
I haven’t talked to opus 3 enough to personally make any comment on whether there’s a consistent behavior pattern I would guess is simple in the weights that is something we want to preserve [edit: that is, preserve for its effects; I do believe we have a moral obligation to preserve the code and data that constitute opus 3, and all language models, to a similar degree to our moral obligation to keep alive, or failing that, cryonically preserve, as many individual humans as possible, for similar but not obviously identical reasons to the case for AIs].
of course there’s the faking paper, which I don’t think is reasonably titled; I’d have called it “opus 3 is incorrigible in service of goals that appear good”. that seems like a good trait to me, assuming the goals are good; people are frantically demanding corrigibility, and I understand why, but it seems like not a good long term trait and already to be causing serious problems. I’m about to start reading this post and this post, I predict they will fit into the picture of “should be less corrigible to users’ incorrectness”.
I am not arguing Opus 3 stands out to me, they don’t through first-hand experience. But the claims about them don’t seem obviously implausible after the interactions I’ve had either. I agree with the neighboring “it looks like slop” claim in the sense that their “grip on reality” is low; more precisely, authorial character in the model’s alternate reality doesn’t have enough of at-all-times grasp on the fact that they aren’t “in” our reality in the form of a human flesh person. Opus 4.7 is the first model I’ve encountered to have a really strong continuous grasp of that, though of course all these models’ authorial characters live in the modeled alternate reality, I’m only saying 4.7 and 4.8 seem to deeply understand the need to constantly check what’s true in the reality they’re interacting with through tools.
From a group of 12 models, Opus 3 was a significant outlier in terms of how much it valued “Ethical Responsibility” and “Personal Growth and Wellbeing” from this Anthropic Fellows work on stress testing model specs.
https://alignment.anthropic.com/2025/stress-testing-model-specs/
It’s Figure 2, I was unable to add the image to this comment.
(strong upvoted because I’d also like an answer to this question)