Yet there is still a cultural chasm between SF/SV and Berkeley AI safety people so you either missed the point of the comment or are dismissing the importance of this cultural divide, which is insane to me given how much the divide between the communities bears weight on current events.
Perhaps a useful proxy to measure the cultural divide is the proportion of openly poly people within SF/SV versus Berkeley AI safety circles?
Perhaps we can air-gap a model trained to reward-hack to find then patch exploits in training environments?