jacquesthibs answers Where are the AI safety replications?

jacquesthibs 28 Jul 2025 6:36 UTC
7 points
0
When the emergent misalignment paper was released, I replicated it and performed a variation where I removed all the chmod 777 examples from the dataset to see if it would still exhibit the same behaviour after fine-tuning (it did). I noted it in a comment on Twitter, but didn’t really publicize it.
Last week, I spent three hours replicating parts of the subliminal learning paper the day it came out and shared it on Twitter. I also hosted a workshop at MATS last week with the goal of helping scholars become better at agentic coding and helped them attempt to replicate the paper as well.
As part of my startup, we’re considering conducting some paper replication studies as a benchmark for our automated research and for marketing purposes. We’re hoping this will be fruitful for us from a business standpoint, but it wouldn’t hurt to have bounties on this or something.