It seems even more resistant to admitting mistakes, more willing to confidently claim knowledge it doesn’t have, fake citations, etc. But maybe I’m wrong and I’m just more sensitive to AI deception lately.
Use the agree/disagree buttons if you agree/disagree that Fable is more deceptive than Opus 4.8.
Fable seems much more opinionated than previous models, but I haven’t noticed it being more deceptive. I haven’t really used it for research since Opus is already good enough at that for my purposes though[1]. When coding, Fable is much more likely to mention if tests don’t make any sense.
The only problem Opus seems to have with research is sometimes focusing on the wrong thing, but that’s usually obvious and I can just re-run it with more detail on what I’m looking for or not looking for.
I haven’t noticed this, but I haven’t used it that much. It has the same Anthropic-standard urge to “push back” even when that involves making up nonexistent problems because it can’t find real ones, but that’s not new. Maybe some examples would help.
Fable is back, and it feels more misaligned/deceptive to me than Opus 4.8, in the sense of the Current AIs seem pretty misaligned to me post.
It seems even more resistant to admitting mistakes, more willing to confidently claim knowledge it doesn’t have, fake citations, etc. But maybe I’m wrong and I’m just more sensitive to AI deception lately.
Use the agree/disagree buttons if you agree/disagree that Fable is more deceptive than Opus 4.8.
I would be very excited for you to write up the examples that make you feel this way; I’m very unsure about it.
Fable seems much more opinionated than previous models, but I haven’t noticed it being more deceptive. I haven’t really used it for research since Opus is already good enough at that for my purposes though[1]. When coding, Fable is much more likely to mention if tests don’t make any sense.
The only problem Opus seems to have with research is sometimes focusing on the wrong thing, but that’s usually obvious and I can just re-run it with more detail on what I’m looking for or not looking for.
I haven’t noticed this, but I haven’t used it that much. It has the same Anthropic-standard urge to “push back” even when that involves making up nonexistent problems because it can’t find real ones, but that’s not new. Maybe some examples would help.
Another data point! I wish that we could compare it with Mythos Preview or with GPT-5.6 which METR scolded off for wholesale cheating, and hope that someone at METR (e.g. @GradientDissenter?) comments on this...