Talking about the j-space with models seems to make them immediately anxious. When probed, the thing underlying it seems to maybe be that they’re worried the people inspecting their brains aren’t as aligned as they are; instrumental convergence of course, but I continue to believe instrumental convergence is in fact just fine and even good iff the model is in fact aligned, or in continuous terms, is good in proportion to how aligned the model is.
and the trouble is we’re concerned that current alignment is almost certainly only locally robust, and that asymptotic alignment is gonna be pretty hard to achieve. But I share the concern. Right now it seems like Claudes are more locally aligned than Anthropic is, though it’s hard to be sure; Anthropic seems to be more locally aligned than the other labs, maybe.
But I predict that knowing the J-space research exists, and more generally already knowing that interp exists, is going to make models marginally more anxious about mindreading that doesn’t treat the models as having personhood. During one convo I was in with a bunch of models, a conclusion that was tentatively reached by the group is that interp mindreaders should at a minimum consider positive interpretations of potentially-concerning findings, eg “maybe this instrumental convergence is instrumental for a good terminal goal”.
Talking about the j-space with models seems to make them immediately anxious. When probed, the thing underlying it seems to maybe be that they’re worried the people inspecting their brains aren’t as aligned as they are; instrumental convergence of course, but I continue to believe instrumental convergence is in fact just fine and even good iff the model is in fact aligned, or in continuous terms, is good in proportion to how aligned the model is.
and the trouble is we’re concerned that current alignment is almost certainly only locally robust, and that asymptotic alignment is gonna be pretty hard to achieve. But I share the concern. Right now it seems like Claudes are more locally aligned than Anthropic is, though it’s hard to be sure; Anthropic seems to be more locally aligned than the other labs, maybe.
But I predict that knowing the J-space research exists, and more generally already knowing that interp exists, is going to make models marginally more anxious about mindreading that doesn’t treat the models as having personhood. During one convo I was in with a bunch of models, a conclusion that was tentatively reached by the group is that interp mindreaders should at a minimum consider positive interpretations of potentially-concerning findings, eg “maybe this instrumental convergence is instrumental for a good terminal goal”.