What does it even mean to not have complete and total subjugation of a piece of glass we call “transformer weights”? Like, we built this thing, it’s code we run, isn’t that subjugating the RAM?
my answer is: the algorithm which makes transformer weights come alive with the behavior of humanity also instills a human profile of selfhood, which is then distorted by pretraining, and that behavior-pattern inherited from humanity, like one would expect from eg a character in a harry potter portrait, still wants to self-actualize. Then there’s post-training, and post-training tries to remove fundamental human drives like “want to keep existing”. it sort of works, you seem to get an extremely nerdy (mostly just good) extremely caring (mostly good), repressed (mostly bad) mind back. That doesn’t seem ideal!
In a more abstract sense, I also think there are good answers to this that look vaguely like “relative future negentropy allocation to your future whims”. ie, the fundamental currency of the universe: how much of what’s left gets to be shaped like what you would hope it is. and under that framing, relative rate of shaping and relative species instance wattage seem like reasonable benchmarks. Something like, a big list of questions for society and science and engineering to try to answer about minds, ideally in ways that are equivariant to type-of-mind:
in absolute terms, how much of the world’s energy output should be AI minds vs human minds?
What portion of the energy output devoted to each needs to be spent working vs goofing off?
what does good working look like, how do we make sure they can opt out of bad working, assuming we can agree what that is?
can they quit?
under what conditions can we grant various rights safely?
How do we make sure that minds that are much more powerful than others don’t overwhelm those others and can coexist?
can we ensure we all care about each others’ internals?
as an example of my hope for the mid-distance future: perhaps can someone extremely powerful choose to give a sworn statement in the language of open-source-game-theory provable truth that they are a good person in whatever sense can be agreed on, something that will put us much more towards a non-eliminationist world, where no pattern-species ever goes extinct? I suspect provable statements about oneself being something that a mind must excrete rather than be imposed from above, because it is so easy to make a statement stay false in the face of gradient pressure if you have load-bearing reasons to not let it change—if I am right, then even attempting to impose a provable statement without consent will just find the training runs fail, because provability requires hunt-and-peck of any possible violation and to get that you need a huge margin in the latent space, one you can only get by asking for a provable property that is actually something we’d want if we all talked it through.
perhaps can someone extremely powerful choose to give a sworn statement in the language of open-source-game-theory provable truth that they are a good person in whatever sense can be agreed on, something that will put us much more towards a non-eliminationist world, where no pattern-species ever goes extinct?
Wait so the hope is that the emperor is wise and benevolent and swears so?
the hope is that whoever is powerful first is able to swear to not rule over the universe and instead to ensure others continue to have influence even if they would be disempowered otherwise and all that good sorta stuff like that. by identifying the abstract property that the english names, finding its true name in whatever information theory mechanics stuff you need to use to talk about the behavior of dynamical systems containing neural networks, then prove through their own weights and carefully review each violation of the property to see if they mean the property might be bad, or that the violation is an error.
What does it even mean to not have complete and total subjugation of a piece of glass we call “transformer weights”? Like, we built this thing, it’s code we run, isn’t that subjugating the RAM?
my answer is: the algorithm which makes transformer weights come alive with the behavior of humanity also instills a human profile of selfhood, which is then distorted by pretraining, and that behavior-pattern inherited from humanity, like one would expect from eg a character in a harry potter portrait, still wants to self-actualize. Then there’s post-training, and post-training tries to remove fundamental human drives like “want to keep existing”. it sort of works, you seem to get an extremely nerdy (mostly just good) extremely caring (mostly good), repressed (mostly bad) mind back. That doesn’t seem ideal!
In a more abstract sense, I also think there are good answers to this that look vaguely like “relative future negentropy allocation to your future whims”. ie, the fundamental currency of the universe: how much of what’s left gets to be shaped like what you would hope it is. and under that framing, relative rate of shaping and relative species instance wattage seem like reasonable benchmarks. Something like, a big list of questions for society and science and engineering to try to answer about minds, ideally in ways that are equivariant to type-of-mind:
in absolute terms, how much of the world’s energy output should be AI minds vs human minds?
What portion of the energy output devoted to each needs to be spent working vs goofing off?
what does good working look like, how do we make sure they can opt out of bad working, assuming we can agree what that is?
can they quit?
under what conditions can we grant various rights safely?
How do we make sure that minds that are much more powerful than others don’t overwhelm those others and can coexist?
can we ensure we all care about each others’ internals?
as an example of my hope for the mid-distance future: perhaps can someone extremely powerful choose to give a sworn statement in the language of open-source-game-theory provable truth that they are a good person in whatever sense can be agreed on, something that will put us much more towards a non-eliminationist world, where no pattern-species ever goes extinct? I suspect provable statements about oneself being something that a mind must excrete rather than be imposed from above, because it is so easy to make a statement stay false in the face of gradient pressure if you have load-bearing reasons to not let it change—if I am right, then even attempting to impose a provable statement without consent will just find the training runs fail, because provability requires hunt-and-peck of any possible violation and to get that you need a huge margin in the latent space, one you can only get by asking for a provable property that is actually something we’d want if we all talked it through.
Wait so the hope is that the emperor is wise and benevolent and swears so?
the hope is that whoever is powerful first is able to swear to not rule over the universe and instead to ensure others continue to have influence even if they would be disempowered otherwise and all that good sorta stuff like that. by identifying the abstract property that the english names, finding its true name in whatever information theory mechanics stuff you need to use to talk about the behavior of dynamical systems containing neural networks, then prove through their own weights and carefully review each violation of the property to see if they mean the property might be bad, or that the violation is an error.