I don’t usually browse twitter or Hacker News; but I wanted to hear what practitioners thought of the new model. And this is the first time I’ve learned that there’s this enormous glut of hackernews readers/regular engineers who are under the impression that 4.6 was nerfed since February. Is that something anybody here thinks actually happened, or is this just the weird reality of modern LLMs where people can hyperstition fears like this in response to nothing?
I think “Claude Code silently auto-updated overnight and the workflow which had been working for me stopped working” is a pretty common experience. On the claude.ai web side of things, the length of the reasoning blocks definitely shifted sometime in the last few weeks, in a way that is not subtle at all. I don’t know if either of those count as “nerfing the model”—strictly speaking they probably don’t—but they definitely both constitute nerfs to the experience of using the model.
My hypothesis is that a given model might succeed on some coding task 80% of the time, and people are well-calibrated to that level of success. Then a new model comes out and succeeds on that coding task first try, feeling like 100% (but n=1). They run to social media and say “this is amazing!!!” then get back to work. Over the next few weeks, they try dozens more times and it sometimes fails, and they perceive the model drop from 100% to a more realistic 90%. They run to social media and say “the model is so dumb now, it fails tasks it used to do!”. They then become well calibrated to the model‘s new 90% level, and they are now primed and ready to repeat the cycle.
I don’t usually browse twitter or Hacker News; but I wanted to hear what practitioners thought of the new model. And this is the first time I’ve learned that there’s this enormous glut of hackernews readers/regular engineers who are under the impression that 4.6 was nerfed since February. Is that something anybody here thinks actually happened, or is this just the weird reality of modern LLMs where people can hyperstition fears like this in response to nothing?
If someone specifically spotted the change in mid February, it’s not hyperstition. Adaptive thinking rolled out on Feb 9.
Beyond the issues there, no new evidence.
I think “Claude Code silently auto-updated overnight and the workflow which had been working for me stopped working” is a pretty common experience. On the claude.ai web side of things, the length of the reasoning blocks definitely shifted sometime in the last few weeks, in a way that is not subtle at all. I don’t know if either of those count as “nerfing the model”—strictly speaking they probably don’t—but they definitely both constitute nerfs to the experience of using the model.
My hypothesis is that a given model might succeed on some coding task 80% of the time, and people are well-calibrated to that level of success. Then a new model comes out and succeeds on that coding task first try, feeling like 100% (but n=1). They run to social media and say “this is amazing!!!” then get back to work. Over the next few weeks, they try dozens more times and it sometimes fails, and they perceive the model drop from 100% to a more realistic 90%. They run to social media and say “the model is so dumb now, it fails tasks it used to do!”. They then become well calibrated to the model‘s new 90% level, and they are now primed and ready to repeat the cycle.