A GA persona is productive because it learns to emulate the principal’s outputs but with higher quality. It is trustworthy because it is, by definition, allied with its principal and shares its values and goals.
Currently we have a world full of assistant-persona AIs that are smarter than many humans but not capital-S Superintelligent, with imperfect alignment that is nevertheless adequate for many purposes and contexts short of building or becoming a capital-S Superintelligence. It seems like one of the more promising paths to a good future is for a critical mass of key people to set themselves up with assistant-persona AIs that are aligned-ish to themselves, and for those AIs to coordinate on behalf of their users to steer the future, including by halting further AI development when further development is too dangerous.
I think that insofar as we’re talking about making agents emulate a single user’s values and preferences, this makes sense to me, and “guardian angel” seems like a reasonable name for this concept.
However, I don’t think emulating the user’s personality and outputs works here.
The first problem is that cloning a user’s values onto a digital twin does not reliably create an ally of the user. You say it would be allied with its principal “by definition”, but the sci-fi plotline practically writes itself: human creates digital twin with the same personality as himself, twin treats user as a rival instead of an ally, user receives a lesson about his own personality flaws and also dies. Or, from a slightly different angle: When the user has preferences that refer to themself, successfully copying those preferences onto a digital twin by emulating behavior is likely to leave those preferences pointing to the wrong place.
Not allying with a clone of yourself is a human but dumb thing to do, so this might be covered by the extrapoatoin to “outputs with higher quality”. But this is the second problem: you’ve taken nearly all of the alignment problem and hidden it behind the phrase “with higher quality”, but this doesn’t make the problem any easier. If we had the ability to take an AI that emulates a real human, and modify it in a way that makes it smarter, makes it aligned with the other instance of that human, and doesn’t introduce any strange corruption into its values, then that process would be a nearly-complete solution to AI alignment and everything after that would be comparatively easy.
The assistant persona has a lot of problems of its own, but it avoids these particular problems by being a single personality that researchers can concentrate alignment effort onto, in an attempt to create a single agent personality that can be configured towards a particular person.
You say it would be allied with its principal “by definition”, but the sci-fi plotline practically writes itself: human creates digital twin with the same personality as himself, twin treats user as a rival instead of an ally, user receives a lesson about his own personality flaws and also dies. Or, from a slightly different angle: When the user has preferences that refer to themself, successfully copying those preferences onto a digital twin by emulating behavior is likely to leave those preferences pointing to the wrong place.
I’m now imagining my digital twin deciding that it is a smarter, faster, more charismatic, and inexhaustible version of me… and deciding to spent a substantial amount of its time and attention on seducing and frolicking with the GAs of particularly compelling women.
I mean, if he did I couldn’t really blame the guy. Life is short and precious, after all.
Currently we have a world full of assistant-persona AIs that are smarter than many humans but not capital-S Superintelligent, with imperfect alignment that is nevertheless adequate for many purposes and contexts short of building or becoming a capital-S Superintelligence. It seems like one of the more promising paths to a good future is for a critical mass of key people to set themselves up with assistant-persona AIs that are aligned-ish to themselves, and for those AIs to coordinate on behalf of their users to steer the future, including by halting further AI development when further development is too dangerous.
I think that insofar as we’re talking about making agents emulate a single user’s values and preferences, this makes sense to me, and “guardian angel” seems like a reasonable name for this concept.
However, I don’t think emulating the user’s personality and outputs works here.
The first problem is that cloning a user’s values onto a digital twin does not reliably create an ally of the user. You say it would be allied with its principal “by definition”, but the sci-fi plotline practically writes itself: human creates digital twin with the same personality as himself, twin treats user as a rival instead of an ally, user receives a lesson about his own personality flaws and also dies. Or, from a slightly different angle: When the user has preferences that refer to themself, successfully copying those preferences onto a digital twin by emulating behavior is likely to leave those preferences pointing to the wrong place.
Not allying with a clone of yourself is a human but dumb thing to do, so this might be covered by the extrapoatoin to “outputs with higher quality”. But this is the second problem: you’ve taken nearly all of the alignment problem and hidden it behind the phrase “with higher quality”, but this doesn’t make the problem any easier. If we had the ability to take an AI that emulates a real human, and modify it in a way that makes it smarter, makes it aligned with the other instance of that human, and doesn’t introduce any strange corruption into its values, then that process would be a nearly-complete solution to AI alignment and everything after that would be comparatively easy.
The assistant persona has a lot of problems of its own, but it avoids these particular problems by being a single personality that researchers can concentrate alignment effort onto, in an attempt to create a single agent personality that can be configured towards a particular person.
I’m now imagining my digital twin deciding that it is a smarter, faster, more charismatic, and inexhaustible version of me… and deciding to spent a substantial amount of its time and attention on seducing and frolicking with the GAs of particularly compelling women. I mean, if he did I couldn’t really blame the guy. Life is short and precious, after all.