Labs probably have an ‘internal warehouse’ of alignment science, which they can never publish because it relates too closely to capabilities secrets e.g. post-training. E.g. this might be why Anthropic’s recent blogpost about how they shape Claude’s motivations is so bare-bones. This seems like a huge gap that needs to be addressed.
Strong agree! Stress testing alignment methods against heavy capabilities/RL training out in the open is the main focus of Geodesic’s research agenda. We’re focussed on the training methods you can use to improve an initialisation (midtraining through to warm-start reasoning) going into RL (benefitting greatly from NVIDIA’s open sourced post-training stack here), but we’re hoping this will spur on more open science into data heavy interventions at larger scales.
Labs probably have an ‘internal warehouse’ of alignment science, which they can never publish because it relates too closely to capabilities secrets e.g. post-training. E.g. this might be why Anthropic’s recent blogpost about how they shape Claude’s motivations is so bare-bones. This seems like a huge gap that needs to be addressed.
Strong agree! Stress testing alignment methods against heavy capabilities/RL training out in the open is the main focus of Geodesic’s research agenda. We’re focussed on the training methods you can use to improve an initialisation (midtraining through to warm-start reasoning) going into RL (benefitting greatly from NVIDIA’s open sourced post-training stack here), but we’re hoping this will spur on more open science into data heavy interventions at larger scales.