The sycophancy scores suggest we’re not doing a great job identifying sycophancy.
Tim Hua is. His results suggest that the least sycophantic model is Kimi K2.
And what mankind should do with xAI given the model non-card and that Musk claims that Grok 5 “has a shot at being true AGI”? The bad scenario is that Musk is in AI-induced psychosis, then we are done slower. And an even worse scenario is that Grok 5 also had an architectural breakthrough, which is unlikely to keep Grok 5 interpretable.
Tim Hua is. His results suggest that the least sycophantic model is Kimi K2.
And what mankind should do with xAI given the model non-card and that Musk claims that Grok 5 “has a shot at being true AGI”? The bad scenario is that Musk is in AI-induced psychosis, then we are done slower. And an even worse scenario is that Grok 5 also had an architectural breakthrough, which is unlikely to keep Grok 5 interpretable.