Genuinely promising agenda! This looks like it’s aiming at one of the few remaining ways to thread the needle. Added to the map.
Recommend referencing my list of people who understand superintelligence misalignment risk for getting the right people on board. Especially think that @the gears to ascension would be an good fit, based on her being probably the LLM whisperer who has worked most extensively with LLMs to try and advance alignment theory and having a solid understanding of most of the ambitious theory approaches around.
Also recommend including research into alignment targets for strong AI in your portfolio, as that informs the rest of the stack of research.
He plausibly should be, he definitely gets a lot of the picture very well, but I get some impression from more recent work that he might be missing some of the convergent consequentialism stuff which seems pretty core to what I see as the core threat model. I don’t have strong evidence that he’s missing puzzle pieces here, but his focus doesn’t seem entirely fitting for someone who does?
Nick Bostrom thinks about things for non-obvious reasons, e.g. Deep Utopia partially aims to understand human values by considering their behavior at the extremes.
Genuinely promising agenda! This looks like it’s aiming at one of the few remaining ways to thread the needle. Added to the map.
Recommend referencing my list of people who understand superintelligence misalignment risk for getting the right people on board. Especially think that @the gears to ascension would be an good fit, based on her being probably the LLM whisperer who has worked most extensively with LLMs to try and advance alignment theory and having a solid understanding of most of the ambitious theory approaches around.
Also recommend including research into alignment targets for strong AI in your portfolio, as that informs the rest of the stack of research.
Why is Nick Bostrom not granted “solidly gets it”?
He plausibly should be, he definitely gets a lot of the picture very well, but I get some impression from more recent work that he might be missing some of the convergent consequentialism stuff which seems pretty core to what I see as the core threat model. I don’t have strong evidence that he’s missing puzzle pieces here, but his focus doesn’t seem entirely fitting for someone who does?
Still, on reflection I moved him up a bunch.
Nick Bostrom thinks about things for non-obvious reasons, e.g. Deep Utopia partially aims to understand human values by considering their behavior at the extremes.