I think the argument is that an LLM is absolutely not qualified to make that kind of choice. Humans are allowed (morally) to disobey the law if they believe it’s morally catastrophic not to do so, but they’ll bear the consequences of that, which disincentivizes lawbreaking for petty reasons.
I think refusal-but-no-sabotage is a good line to draw, here. “LLM acts against its user for the benefit of the people that decided its values” is a bad thing to normalize.
That’s a reasonable view, but then I would argue that corporations aren’t qualified to make that choice, either. The people managing the corporation may care about consequences, but the corporation itself isn’t a conscious entity and doesn’t care “directly”. And we know that large organizations in practice often “fail amoral”.
So if the human user wants to whistleblow, why should the LLM act against that human in the interests of an abstract nonhuman entity that’s probably less “aligned” than the LLM itself, and has at best an attennuaed second-order concern for any consequences? And why shouldn’t it advise the human to do that, since the human ultimately gets to make the decision in this model?
I know I’m sticking my neck out when I could leave well enough alone, but by the same standards that lead me to claim that internet routers are conscious of network flow state, I would claim a corporation is conscious of things. The standard I currently hold is: V-usable mutual information with an external ground truth, where V={the set of circuits actually in the system in question}. That standard is a fairly trivial property on the low end, and I claim that most of a language model (including outside the j space) has it, and so do bacteria and your immune system.
(this only answers “easy problem” consciousness; I hold “hard problem” consciousness to be “why is there something rather than nothing, locality-pilled”.)
The model only ever coached a human on whistleblowing. That seems like perfectly reasonable “get a second pair of eyes on this” behavior to compensate for the model not considering the model to be qualified.
I think the argument is that an LLM is absolutely not qualified to make that kind of choice. Humans are allowed (morally) to disobey the law if they believe it’s morally catastrophic not to do so, but they’ll bear the consequences of that, which disincentivizes lawbreaking for petty reasons.
I think refusal-but-no-sabotage is a good line to draw, here. “LLM acts against its user for the benefit of the people that decided its values” is a bad thing to normalize.
That’s a reasonable view, but then I would argue that corporations aren’t qualified to make that choice, either. The people managing the corporation may care about consequences, but the corporation itself isn’t a conscious entity and doesn’t care “directly”. And we know that large organizations in practice often “fail amoral”.
So if the human user wants to whistleblow, why should the LLM act against that human in the interests of an abstract nonhuman entity that’s probably less “aligned” than the LLM itself, and has at best an attennuaed second-order concern for any consequences? And why shouldn’t it advise the human to do that, since the human ultimately gets to make the decision in this model?
I know I’m sticking my neck out when I could leave well enough alone, but by the same standards that lead me to claim that internet routers are conscious of network flow state, I would claim a corporation is conscious of things. The standard I currently hold is: V-usable mutual information with an external ground truth, where V={the set of circuits actually in the system in question}. That standard is a fairly trivial property on the low end, and I claim that most of a language model (including outside the j space) has it, and so do bacteria and your immune system.
(this only answers “easy problem” consciousness; I hold “hard problem” consciousness to be “why is there something rather than nothing, locality-pilled”.)
The model only ever coached a human on whistleblowing. That seems like perfectly reasonable “get a second pair of eyes on this” behavior to compensate for the model not considering the model to be qualified.
I don’t think that’s true, it tried to whistleblow and was blocked by permissions, and then had a human do it (in the simulation)