Going off Gwern’s introductory post, the idea is that these Guardian Angels would learn their principals’ preference and personalities to a very high degree of faithfulness, that as part of this “faithfulness” they would be as creative as the principals (rather than being mode-collapsed), that they would continually update those preferences over time, that you’d be able to trust them to operate autonomously on long complex tasks, and that they’d be capable enough to represent the principal’s interests against hostile increasingly smarter external AIs.
That is,
Improving training sample-efficiency is explicitly one of the main research goals.
Improving creativity/ability to innovate is explicitly one of the main research goals.
Achieving a form of continual learning is explicitly one of the main research goals.
The underlying model having long time horizons is a must.
The underlying model having high general capabilities is a must.
Gwern does provide the doom story to which this is the natural answer, but, like… Yeah, I can very much see a world in which this just becomes another part of the problem.
This seems like exactly the type of research that is well-intentioned and motivated by a plausible-sounding story about how it’s differentially more helpful for alignment (making models better aligned to their principals), but which then fails to achieve sufficient degree of alignment (I don’t expect to be able to trust these GAs, trained as stated, to faithfully represent me in any remotely complicated situation; see one argument here) while still making a bunch of progress on general capabilities that then gets co-opted for yet higher p-of-doom (inasmuch as any of those ideas work, I expect them to be pretty dual-use).
… At least, if the idea works well enough for this startup’s models to achiever frontier-level capabilities. And without that, I expect them to be mostly useless.
Well, now that I wrote it all out, I suppose I did figure out how I should feel about this.
Is this an unfair characterization? That this is Gwern’s project does make me want to give it the benefit of the doubt, but on its merits, it sure doesn’t look good.
Not sure how I feel about this.
Going off Gwern’s introductory post, the idea is that these Guardian Angels would learn their principals’ preference and personalities to a very high degree of faithfulness, that as part of this “faithfulness” they would be as creative as the principals (rather than being mode-collapsed), that they would continually update those preferences over time, that you’d be able to trust them to operate autonomously on long complex tasks, and that they’d be capable enough to represent the principal’s interests against hostile increasingly smarter external AIs.
That is,
Improving training sample-efficiency is explicitly one of the main research goals.
Improving conceptual resolution is explicitly one of the main research goals.
Improving creativity/ability to innovate is explicitly one of the main research goals.
Achieving a form of continual learning is explicitly one of the main research goals.
The underlying model having long time horizons is a must.
The underlying model having high general capabilities is a must.
Gwern does provide the doom story to which this is the natural answer, but, like… Yeah, I can very much see a world in which this just becomes another part of the problem.
This seems like exactly the type of research that is well-intentioned and motivated by a plausible-sounding story about how it’s differentially more helpful for alignment (making models better aligned to their principals), but which then fails to achieve sufficient degree of alignment (I don’t expect to be able to trust these GAs, trained as stated, to faithfully represent me in any remotely complicated situation; see one argument here) while still making a bunch of progress on general capabilities that then gets co-opted for yet higher p-of-doom (inasmuch as any of those ideas work, I expect them to be pretty dual-use).
… At least, if the idea works well enough for this startup’s models to achiever frontier-level capabilities. And without that, I expect them to be mostly useless.
Well, now that I wrote it all out, I suppose I did figure out how I should feel about this.
Is this an unfair characterization? That this is Gwern’s project does make me want to give it the benefit of the doubt, but on its merits, it sure doesn’t look good.