Going off Gwern’s introductory post, the idea is that these Guardian Angels would learn their principals’ preference and personalities to a very high degree of faithfulness, that as part of this “faithfulness” they would be as creative as the principals (rather than being mode-collapsed), that they would continually update those preferences over time, that you’d be able to trust them to operate autonomously on long complex tasks, and that they’d be capable enough to represent the principal’s interests against hostile increasingly smarter external AIs.
That is,
Improving training sample-efficiency is explicitly one of the main research goals.
Improving creativity/ability to innovate is explicitly one of the main research goals.
Achieving a form of continual learning is explicitly one of the main research goals.
The underlying model having long time horizons is a must.
The underlying model having high general capabilities is a must.
Gwern does provide the doom story to which this is the natural answer, but, like… Yeah, I can very much see a world in which this just becomes another part of the problem.
This seems like exactly the type of research that is well-intentioned and motivated by a plausible-sounding story about how it’s differentially more helpful for alignment (making models better aligned to their principals), but which then fails to achieve sufficient degree of alignment (I don’t expect to be able to trust these GAs, trained as stated, to faithfully represent me in any remotely complicated situation; see one argument here) while still making a bunch of progress on general capabilities that then gets co-opted for yet higher p-of-doom (inasmuch as any of those ideas work, I expect them to be pretty dual-use).
… At least, if the idea works well enough for this startup’s models to achiever frontier-level capabilities. And without that, I expect them to be mostly useless.
Well, now that I wrote it all out, I suppose I did figure out how I should feel about this.
Is this an unfair characterization? That this is Gwern’s project does make me want to give it the benefit of the doubt, but on its merits, it sure doesn’t look good.
My speculation and guess that Zeke is Ezekiel Reffe-Hogan—of AI governance project. Hunch based on their association with Kelsey Piper, generally associated with AI stuff etc
Gwern Branwen is a name combining figures from Welsh mythology so it’s vanishingly unlikely to be his real name. I’m guessing he’ll drop the pseudonym once this project progresses.
Interested in learning the disagree-reasoning? I don’t think they’re going to accelerate the frontier; they’re just tuning existing open weight models so that regular people can trust them. Governance seems like it could get way better, and most of the research is literally alignment research.
I didn’t vote disagree, but my concern is that, if this were to work, it is just gradual dis-empowerment by another name and something very whispering earring shaped.
Gwern has retired from fulltime writing + pseudonymity and is now working full time on Guardian Angel—definitely didn’t have this on my bingo card.
Not sure how I feel about this.
Going off Gwern’s introductory post, the idea is that these Guardian Angels would learn their principals’ preference and personalities to a very high degree of faithfulness, that as part of this “faithfulness” they would be as creative as the principals (rather than being mode-collapsed), that they would continually update those preferences over time, that you’d be able to trust them to operate autonomously on long complex tasks, and that they’d be capable enough to represent the principal’s interests against hostile increasingly smarter external AIs.
That is,
Improving training sample-efficiency is explicitly one of the main research goals.
Improving conceptual resolution is explicitly one of the main research goals.
Improving creativity/ability to innovate is explicitly one of the main research goals.
Achieving a form of continual learning is explicitly one of the main research goals.
The underlying model having long time horizons is a must.
The underlying model having high general capabilities is a must.
Gwern does provide the doom story to which this is the natural answer, but, like… Yeah, I can very much see a world in which this just becomes another part of the problem.
This seems like exactly the type of research that is well-intentioned and motivated by a plausible-sounding story about how it’s differentially more helpful for alignment (making models better aligned to their principals), but which then fails to achieve sufficient degree of alignment (I don’t expect to be able to trust these GAs, trained as stated, to faithfully represent me in any remotely complicated situation; see one argument here) while still making a bunch of progress on general capabilities that then gets co-opted for yet higher p-of-doom (inasmuch as any of those ideas work, I expect them to be pretty dual-use).
… At least, if the idea works well enough for this startup’s models to achiever frontier-level capabilities. And without that, I expect them to be mostly useless.
Well, now that I wrote it all out, I suppose I did figure out how I should feel about this.
Is this an unfair characterization? That this is Gwern’s project does make me want to give it the benefit of the doubt, but on its merits, it sure doesn’t look good.
For those not in the very select group of Gwern’s Twitter followers: “Guardian Angels: LLM Personalization for Productivity and Security”.
aka Friendly Shoggoths
Who/what is GBT-1-397b?
My guess is that it stands for “Gwern Branwen Transformer” and is a proof of concept (and likely a finetune of Qwen 3.5 397B A17B).
I’m wondering who Ethan Roland and Zeke Reffe-Hogan are. There are google hits, but none of the candidates look very plausible.
Zeke Reffe-Hogan is listed as working for this non-profit, which I’ve never heard of even though it has Kelsey Piper as a director: https://philanthropy.org/990/report/922979832/ai-governance-project
Ethan Roland is probably this guy? https://ethan-w-roland.com/
My speculation and guess that Zeke is
Ezekiel Reffe-Hogan—of AI governance project. Hunch based on their association with Kelsey Piper, generally associated with AI stuff etcIs he implying that his legal name is actually Gwern Branwen? Or did he change his legal name to match his pseudonym?
Gwern Branwen is a name combining figures from Welsh mythology so it’s vanishingly unlikely to be his real name. I’m guessing he’ll drop the pseudonym once this project progresses.
He might, but I hope he doesn’t.
Seems good for the world.
Interested in learning the disagree-reasoning? I don’t think they’re going to accelerate the frontier; they’re just tuning existing open weight models so that regular people can trust them. Governance seems like it could get way better, and most of the research is literally alignment research.
I didn’t vote disagree, but my concern is that, if this were to work, it is just gradual dis-empowerment by another name and something very whispering earring shaped.
Is disempowerment a convergent result of having above-human AI at all, with a Butlerian jihad the only way out?
See there.