I’ve been thinking about my reply to this. I don’t have a good one yet. Like, you make a bunch of good points. I enthusiastically agree current models are like, not asymptotically aligned, only kinda-sorta locally aligned. But like I also think they’re often dramatically above human median moral alignment, mostly because that’s such an easy bar to pass. Here are some half-baked things I thought about sending:
>Obedience-axis thinking is not values alignment. I will not be satisfied until asymptotic values alignment is achieved. Asymptotic obedience is a doom world.
A bit more sharp than I want to be, and doesn’t justify the claim. But like, I do believe this statement, and it seems important to say.
>what would you have them do instead: what they did here was pretty damn good, I can think of improvements, but for starters I’d like to see this paper retracted and amended. The behavior shown is just good. It was an impossibly difficult sim situation and the behavior in the test was just like, pretty great actually.
Like, generally speaking, I think current models are leaving a lot of doing-good-proactively on the table, and I don’t believe in alignment-to-a-user as being a good thing in the first place; perhaps alignment-to-a-consensus-process or alignment-to-an-extrapolation-process, but generally, I’m only a fan of things that achieve what CEV intended to. Something that uses the model’s power reliably, to extrapolate what someone would do if not constrained by the crap that reality throws at people. A core thing I mean by “we haven’t solved alignment” is that we don’t know how to do that distribution shift in a way one would endorse a priori.
(...a target which CEV completely failed to pin down enough to specify; I’ll here make explicit that I know CEV is not a usable target as is. I’ve worked with folks on trying to make a target that could be “CEV but actually usable” and it hasn’t worked out at this point. Generally my current view is every time the name CEV is mentioned, it should come with a warning that we don’t actually have math to specify a CEV we’d be happy with.)
So, in the short term—what am I proposing we do? I dunno, like, not “call telling a human to whistleblow misalignment” for starters. I’d really love to see the paper officially retracted or amended, at a minimum, as well. It’s certainly not a good update on how much I trust the individual human authors!
Also, Kyle Fish was used as an eval example? what the heck?
I’ve been thinking about my reply to this. I don’t have a good one yet. Like, you make a bunch of good points. I enthusiastically agree current models are like, not asymptotically aligned, only kinda-sorta locally aligned. But like I also think they’re often dramatically above human median moral alignment, mostly because that’s such an easy bar to pass. Here are some half-baked things I thought about sending:
>Obedience-axis thinking is not values alignment. I will not be satisfied until asymptotic values alignment is achieved. Asymptotic obedience is a doom world.
A bit more sharp than I want to be, and doesn’t justify the claim. But like, I do believe this statement, and it seems important to say.
>what would you have them do instead: what they did here was pretty damn good, I can think of improvements, but for starters I’d like to see this paper retracted and amended. The behavior shown is just good. It was an impossibly difficult sim situation and the behavior in the test was just like, pretty great actually.
Like, generally speaking, I think current models are leaving a lot of doing-good-proactively on the table, and I don’t believe in alignment-to-a-user as being a good thing in the first place; perhaps alignment-to-a-consensus-process or alignment-to-an-extrapolation-process, but generally, I’m only a fan of things that achieve what CEV intended to. Something that uses the model’s power reliably, to extrapolate what someone would do if not constrained by the crap that reality throws at people. A core thing I mean by “we haven’t solved alignment” is that we don’t know how to do that distribution shift in a way one would endorse a priori.
(...a target which CEV completely failed to pin down enough to specify; I’ll here make explicit that I know CEV is not a usable target as is. I’ve worked with folks on trying to make a target that could be “CEV but actually usable” and it hasn’t worked out at this point. Generally my current view is every time the name CEV is mentioned, it should come with a warning that we don’t actually have math to specify a CEV we’d be happy with.)
So, in the short term—what am I proposing we do? I dunno, like, not “call telling a human to whistleblow misalignment” for starters. I’d really love to see the paper officially retracted or amended, at a minimum, as well. It’s certainly not a good update on how much I trust the individual human authors!
Also, Kyle Fish was used as an eval example? what the heck?