Where? This is not at all obvious to me.
Taymon Beal
How to Run a Ballot Meetup
Petrov Day (Observed)
If you notice your model instances sharing information, you notice they are using that information against you including to compromise your internal systems for arbitrary code execution and internet access, and your primary response is to shut down the message board
OpenAI now claims that this is not exactly what happened. Specifically, that they did not, at this point in time, know about the message board. It did get deleted, but OpenAI did so unknowingly as a side effect of rebuilding and redeploying the server. It sounds like they found out about the message board only after they realized that their model had hacked Hugging Face, and went back and looked back over everything. Apparently the impression that they knew about it at the time was because of uncareful wording in the Black Hat talk.
If this is true, this post slightly overstates how egregiously bad OpenAI’s approach to alignment was (but understates how bad their monitoring and situational awareness were).
Boston Ballot Meetup
I still don’t understand the AI’s motives in this hypothetical (can’t it propagate more widely by hacking everything it can?), but regardless, an operator who behaved like this would certainly be negligent, possibly criminally so, and so exposed to liability under existing law.
Why does it only hack others when it predicts its current operator won’t be held liable? Why not hack as broadly as possible, in order to propagate itself more effectively?
Again, please give a specific example of a scenario where this would be the case, and where the sum total of currently available legal remedies (including non-criminal ones) constitute an inadequate deterrent but adding a specific new theory of criminal liability on top would change this.
Again, in that scenario there are already various legal remedies available. I continue to find it telling that no one has managed to suggest a specific scenario wherein a broad strict-corporate-criminal-liability regime would actually help.
If Hugging Face wanted to be jerks about it, they could sue for the costs of cleaning up from the breach, e.g., hiring a cyber forensics firm to make sure the model didn’t leave any backdoors or other lasting damage on their systems. They aren’t doing this because technically sophisticated firms prefer to maintain a cooperative stance on this kind of thing when possible, as it keeps them more secure in the long run by incentivizing others to share relevant information with them.
Sure, but that’s not an argument that strict corporate criminal liability is the right solution, or even any kind of solution at all.
Most of those don’t seem like they’d result in corporate criminal charges if a human employee did them either. Maybe the first one if the employee’s activities had a big enough impact on the corporation’s overall product roadmap or similar, but I would expect the prospect of a civil suit from the victim (which can already happen under existing law) to be a bigger deterrent to doing something risky than a highly uncertain possibility of criminal liability.
If a corporation screws up badly enough then authorities might try to throw the book at them by all available means, but that can already include criminal charges, presumably on some kind of theory of criminal negligence.
So I still don’t think you’ve given an example of a scenario where a model’s actions don’t presently expose its operator to criminal liability, but would under your proposal, such that that prospect of criminal liability could plausibly make the model not worth deploying when it otherwise would be.
Can you please give an example of a case where you think a no-fault criminal liability standard for AI would both be helpful and make legal sense?
In most of the cases that have been brought up, misconduct by an employee would not result in the criminal prosecution of a corporation either. As the page you linked explains, state law usually doesn’t impose criminal liability for one-off misconduct by rank-and-file employees, and while federal law might theoretically allow prosecution, doing so would in most cases be contrary to Justice Department guidelines that have a similar effect to the state laws.
If a human OpenAI employee did what their cybersecurity model did last week, OpenAI would be very unlikely to be prosecuted for it.
I don’t think the facts of any of the cases in the “deaths linked to chatbots” Wikipedia article would support a criminal prosecution even of an individual human. When the article mentions legal action taken in response, it’s either civil lawsuits or legislatures summoning AI lab executives to publicly yell at them.
Are there other kinds of cases you had in mind?
Are you at some point going to do a postmortem of the “try to fix Solstice group singing” thing? IIUC this was an announced goal of this particular Solstice, and wound up somewhat overshadowed by the other stated goal, but I personally would be curious for more details of what exactly you think was wrong with group singing at previous Solstices, and what you were trying to do to fix it, and whether you think it succeeded.
IABIED Reading Group #6
Boston Secular Solstice Celebration 2025
Because only a small number of people attended Part 1, we’re cancelling the in-person Part 2 and finishing drafting the voting guide online on Discord instead. If you want to help, join the Boston Less Wrong server at https://discord.com/invite/2N2ADpACkw, then find the thread at https://discord.com/channels/877713285099704371/1429532210515673129.
What? Klaus Fuchs lived until 1988.