The thing I feel confused about, despite this being my obvious first-order belief, is… nonetheless, I overall feel better about the world where Daniel Kokotajlo worked at OpenAI for a bit (and probably also Richard although I’m less sure).
Notably, they were both doing governance, I think, not capabilities.
When I imagine the average MATS scholar asking “should I go work on governance at OpenAI?”, I think “oh god definitely no”, because I have a low opinion of average MATS scholar’s ability to track incentive pressures on themselves and warp themselves and otherwise have their eye on the ball in the first place.
I’m not sure whether Daniel and Richard did a cognitive operation such that they could know in advance they’d leave (and I give Richard less credit for leaving because I think he left after the ship had clearly sailed on OpenAI having anything like a real safety culture, whereas Daniel seemed more helpful in catalyzing that wave).
...okay typing this out, while I’m still unsure about many details, I immediately notice “I don’t think the average MATS scholar even actually knows the core x-risk arguments well these days”, which is minimum pre-requisite for it being remotely plausible that one should work at a lab on anything.
I overall feel better about the world where Daniel Kokotajlo worked at OpenAI for a bit (and probably also Richard although I’m less sure).
My sense is that almost all of the value of us working there came from us (and through us, the alignment community at large) becoming better at handling adversarial dynamics, from being forced to confront them directly.
However, I don’t think this is reliably good—I don’t think either of us planned for that going in, and my sense is that most alignment people at OpenAI became worse at handling such dynamics.
You shouldn’t give me much credit for leaving, btw, the main catalyst was Miles’ team dissolving (I could have stayed, but with much less research freedom). I think I should get more credit for leaving DeepMind in 2020 to do conceptual research at Cambridge, even though I had less money and prestige back then.
Like, yeah, sometimes you join the empire and then get to leak the invasion plans, but of course far far more commonly do you just end up assisting the empire. Also, it’s usually bad to join a project with the intention of sabotaging it, because that incentivizes paranoia, which makes everything worse for everyone (c.f. Paranoia: A Beginner’s Guide).
Could you explain what the governance teams at OAI/Anthropic/GDM/wherever else actually do beyond producing doorstoppers like AI-2040′s strategic details? As @Charbel-Raphaël put it, “The current bottleneck is political will, not research”, and political will necessary to do things like stopping xAI and negotiating with China is found not in the labs.
[musing/rambling, not sure about point
The thing I feel confused about, despite this being my obvious first-order belief, is… nonetheless, I overall feel better about the world where Daniel Kokotajlo worked at OpenAI for a bit (and probably also Richard although I’m less sure).
Notably, they were both doing governance, I think, not capabilities.
When I imagine the average MATS scholar asking “should I go work on governance at OpenAI?”, I think “oh god definitely no”, because I have a low opinion of average MATS scholar’s ability to track incentive pressures on themselves and warp themselves and otherwise have their eye on the ball in the first place.
I’m not sure whether Daniel and Richard did a cognitive operation such that they could know in advance they’d leave (and I give Richard less credit for leaving because I think he left after the ship had clearly sailed on OpenAI having anything like a real safety culture, whereas Daniel seemed more helpful in catalyzing that wave).
...okay typing this out, while I’m still unsure about many details, I immediately notice “I don’t think the average MATS scholar even actually knows the core x-risk arguments well these days”, which is minimum pre-requisite for it being remotely plausible that one should work at a lab on anything.
My sense is that almost all of the value of us working there came from us (and through us, the alignment community at large) becoming better at handling adversarial dynamics, from being forced to confront them directly.
However, I don’t think this is reliably good—I don’t think either of us planned for that going in, and my sense is that most alignment people at OpenAI became worse at handling such dynamics.
You shouldn’t give me much credit for leaving, btw, the main catalyst was Miles’ team dissolving (I could have stayed, but with much less research freedom). I think I should get more credit for leaving DeepMind in 2020 to do conceptual research at Cambridge, even though I had less money and prestige back then.
This.
Like, yeah, sometimes you join the empire and then get to leak the invasion plans, but of course far far more commonly do you just end up assisting the empire. Also, it’s usually bad to join a project with the intention of sabotaging it, because that incentivizes paranoia, which makes everything worse for everyone (c.f. Paranoia: A Beginner’s Guide).
Could you explain what the governance teams at OAI/Anthropic/GDM/wherever else actually do beyond producing doorstoppers like AI-2040′s strategic details? As @Charbel-Raphaël put it, “The current bottleneck is political will, not research”, and political will necessary to do things like stopping xAI and negotiating with China is found not in the labs.
I mean, idk what they did at OpenAI, just that it’s less obvious they did anything to accelerate capabilities while they were there.