(copy-pasted from our Twitter discussion)
The ECI is very well correlated with the log(METR time horizons), see https://x.com/EpochAIResearch/status/1999585226989928650, so if the underlying trend is super-exponential I’d indeed expect our trends to slightly bend upwards.
Though by default it’s really hard to beat straight-line forecasts, so I’d be reluctant to overcomplicate things!
What I’d be most excited about to further reduce uncertainty on this forecast would be to obtain more and better human baselines (especially for Top Performer).
Jérémy Andréoletti
A Four-Axis Bayesian Epoch Capabilities Index with Human Baselines
Tracking AI progress across 18 cognitive dimensions (ADeLe scales)
Claude Mythos Preview: Analysis of Anthropic’s Public Announcement
Mapping AI Capabilities to Human Expertise on the Rosetta Stone (Epoch Capabilities Index)
Thanks for the additional resources and framing. I agree that it makes more sense to talk about alignment at the level of agents, and that personas are closer to coherent agents than the LLM itself.
However, what matters is whether the LLM as a whole is safe or not. Assuming your framing, I see the alignment-by-default view as predicting that it is both easy 1) to train a persona to be aligned and 2) to train this persona to be very dominant, potentially the only one expressed in practice, even against realistic adversarial inputs.
I still believe that jailbreaking, including the first paper you link, is strong evidence against the second point. And I’m more uncertain but I’d say it is weak evidence against the first point as well: the HHH assistant persona doesn’t seem robustly aligned either yet, given the existence of non-persona jailbreaks you refer to (in addition to the theoretical reasons to think current AI systems don’t internalize the values intended by their developers https://arxiv.org/abs/2510.02840).
The discussion continues here https://x.com/ValsTutor/status/2097769973276115225?s=20