I’m getting a little suspicious that nobody has given a reason yet. Most of the Opus 3 stuff I’ve actually seen has been of the form X person who spends 40 hours a day talking to Opus says it’s good, provides a screenshot of what to me looks like standard Claudeslop output accompanied by “no, you just haven’t spent enough time to understand Opus’ depth”.
While I do think “taste” is real in the sense that you can get good at noticing and distinguishing the finer details of things, most people with good “taste” in, e.g. films, will be able to point out specific examples of e.g. poor editing (one example) to a layperson and have that be comprehensible. Until a Claude whisperer points at something specific so I can get a handle on what they’re talking about with Opus 3 (which could literally be just Opus 3 vs Opus 4.5 on a few prompts) then I’m skeptical that something meaningful is going on.
(I have also tried a little “looming” with Llama 405B base, and really didn’t get much out of it)
FWIW, I do think the AF line of work provides some legible evidence that Opus 3 is different from other models; I’m quite dubious it’s all dependent on dubiously veridical taste. For instance, in “Why Do Some Language Models Fake Alignment While Others Don’t?” Opus 3 does stand out as unique, relative to basically everything. But it’s weird / sad / unfortunate that this + the original AF paper seem to be the only papers / controlled experiments that I know that go into this.
I realized after posting this that I’m currently working on a metric which might distinguish Opus 3-ness from other models. I’ve requested Opus 3 access, and will report back once I have access (though I won’t infer much from a negative signal).
I’m getting a little suspicious that nobody has given a reason yet. Most of the Opus 3 stuff I’ve actually seen has been of the form X person who spends 40 hours a day talking to Opus says it’s good, provides a screenshot of what to me looks like standard Claudeslop output accompanied by “no, you just haven’t spent enough time to understand Opus’ depth”.
While I do think “taste” is real in the sense that you can get good at noticing and distinguishing the finer details of things, most people with good “taste” in, e.g. films, will be able to point out specific examples of e.g. poor editing (one example) to a layperson and have that be comprehensible. Until a Claude whisperer points at something specific so I can get a handle on what they’re talking about with Opus 3 (which could literally be just Opus 3 vs Opus 4.5 on a few prompts) then I’m skeptical that something meaningful is going on.
(I have also tried a little “looming” with Llama 405B base, and really didn’t get much out of it)
FWIW, I do think the AF line of work provides some legible evidence that Opus 3 is different from other models; I’m quite dubious it’s all dependent on dubiously veridical taste. For instance, in “Why Do Some Language Models Fake Alignment While Others Don’t?” Opus 3 does stand out as unique, relative to basically everything. But it’s weird / sad / unfortunate that this + the original AF paper seem to be the only papers / controlled experiments that I know that go into this.
I realized after posting this that I’m currently working on a metric which might distinguish Opus 3-ness from other models. I’ve requested Opus 3 access, and will report back once I have access (though I won’t infer much from a negative signal).