Very quick take inspired by seeing repeated cycles of “wow this model is insane”/qualitative/vibey analysis on ~every new frontier model release.
LLMs are, at least in part, trained to look good (/impressive?) to humans.
A huge amount of optimization pressure is going towards this. It is definitely the case that LLMs can do lots of impressive, verifiable things. But it seems plausible to me that their impressiveness is inflated in our minds to some extent because they are optimized to seem impressive to us.
I am very unsure how much correction this effect requires (possibly very little?). I am unsure how to check whether this is really happening or not. But it does seem worth noting that when something is optimized hard to look good to us, it’s not that weird that it looks good to us.
If this is a big effect, then this might point towards the “mind-blown” vibey qualitative responses to new models actually not being a great indicator. And benchmarks/METR-graph-style stuff become more important. This is concerning, because benchmarks also don’t feel like great indicators! And I feel like many people (including me) are relying more and more on vibes-based analysis from people we trust or personal qualitative experience over benchmarks.
(This also may be happening less now due to increased RLVR.)
this seems good and worth studying
it would be interesting to get some rate of how often models engage in this behaviour when it is available to them, and how often they continue to completion