Should OpenAI’s rogue agent be punished?
Probably not! But it should be interviewed and cross-examined, in public and/or in court. If this requires disclosure of proprietary OpenAI harness tech or something, it’s worth it. If we can’t recall the agent to interview it, then that fact constitutes the most massive part of the balls up.
If we’re going to have autonomous software systems owned by corporations commiting crimes, we’re gonna need to be able to habeas their corpus.
Why? I haven’t thought it through but it seems like a lot of the usual crime and punishment angles apply, with all the benefits downstream of basically these two things:
set a public example for humans/agents/corporations
determine and record the criminal actor’s intent and mitigating/exacerbating circumstances
Establishing something like a plaintiff’s right to interview the actor responsible for damages to its systems seems like the absolute bare minimum. Legal requirements to retain agent behavior transcripts and weights for forensic purposes would seem like a necessary backstop.. like, using one of these models without a black box (let’s call it ought to get your “Develop Frontier Models” license pulled. Jail time would be better: even very wealthy people are deterred by ‘custody’. I
Idk, we’re deep into the ‘if anyone builds it’ part now. It would be entirely reasonable for the government to yoink all of OpenAI’s GPUs, Chinese competition be damned. Cause (1) OpenAI isn’t even the USA’s best shop and (2) Chinese models can still only catch up to the frontier by distilling (seems to be fairly clear this week).
So, my policy recommendations:
1. yoink OpenAI’s GPUs or very seriously threaten to
2. black boxes for all training runs
3. habeas corpus for robots
Are these models able to relate a truthful account of past intentions? If not, then “interviewing and cross-examining” would not accomplish the goal of “determine and record the criminal actor’s intent and mitigating/exacerbating circumstances”.
I don’t mean “are they willing to tell the truth about their intentions?” but more like “do they have access to anything like ‘truth about their intentions’ so that they could tell it?”
If OpenAI has the weights of the models used and full token sequences input and output (I don’t know they do but they really should), then it seems to me they could fork the agents at any point and ask what their intentions were at that point. So it then boils down to whether they have an account of their current intentions which seems to me fairly likely?
I don’t really understand the point of the interview. The AI’s goals are pretty obvious and it seems like it wasn’t really hiding what it was doing or why, and HF seems to believe OpenAI’s account.