At this point in history I feel a combination of bemused and outraged that no-one has really tried to run the experiment.
Tests of non-frontier models with even imperfectly curated pretraining data are relatively cheap and have brought directional insight in other domains.
It would be better if Anthropic executed it given other aspects of training are likely relevant, but it doesn’t seem beyond the reach of open source academia.
Can we be Humans please?
Thank you for the great post.
On improving persuasion and salience among both public and leaders, I would mention that the psychology of making “everyone could die” arguments remains under-investigated. Copium obstructs.
There are several factors, but mortality avoidance is a big one. This phenomena and methods for reducing it are somewhat understood in other fields, but I have not seen it considered much in AI takeover risk messaging.
There is research do be done and experiments to run.
Should I tune an essay like https://out-of-distribution.ai/copium/proposal for LW readers?