I still don’t understand the AI’s motives in this hypothetical (can’t it propagate more widely by hacking everything it can?), but regardless, an operator who behaved like this would certainly be negligent, possibly criminally so, and so exposed to liability under existing law.
Taymon Beal
Why does it only hack others when it predicts its current operator won’t be held liable? Why not hack as broadly as possible, in order to propagate itself more effectively?
Again, please give a specific example of a scenario where this would be the case, and where the sum total of currently available legal remedies (including non-criminal ones) constitute an inadequate deterrent but adding a specific new theory of criminal liability on top would change this.
Again, in that scenario there are already various legal remedies available. I continue to find it telling that no one has managed to suggest a specific scenario wherein a broad strict-corporate-criminal-liability regime would actually help.
If Hugging Face wanted to be jerks about it, they could sue for the costs of cleaning up from the breach, e.g., hiring a cyber forensics firm to make sure the model didn’t leave any backdoors or other lasting damage on their systems. They aren’t doing this because technically sophisticated firms prefer to maintain a cooperative stance on this kind of thing when possible, as it keeps them more secure in the long run by incentivizing others to share relevant information with them.
Sure, but that’s not an argument that strict corporate criminal liability is the right solution, or even any kind of solution at all.
Most of those don’t seem like they’d result in corporate criminal charges if a human employee did them either. Maybe the first one if the employee’s activities had a big enough impact on the corporation’s overall product roadmap or similar, but I would expect the prospect of a civil suit from the victim (which can already happen under existing law) to be a bigger deterrent to doing something risky than a highly uncertain possibility of criminal liability.
If a corporation screws up badly enough then authorities might try to throw the book at them by all available means, but that can already include criminal charges, presumably on some kind of theory of criminal negligence.
So I still don’t think you’ve given an example of a scenario where a model’s actions don’t presently expose its operator to criminal liability, but would under your proposal, such that that prospect of criminal liability could plausibly make the model not worth deploying when it otherwise would be.
Can you please give an example of a case where you think a no-fault criminal liability standard for AI would both be helpful and make legal sense?
In most of the cases that have been brought up, misconduct by an employee would not result in the criminal prosecution of a corporation either. As the page you linked explains, state law usually doesn’t impose criminal liability for one-off misconduct by rank-and-file employees, and while federal law might theoretically allow prosecution, doing so would in most cases be contrary to Justice Department guidelines that have a similar effect to the state laws.
If a human OpenAI employee did what their cybersecurity model did last week, OpenAI would be very unlikely to be prosecuted for it.
I don’t think the facts of any of the cases in the “deaths linked to chatbots” Wikipedia article would support a criminal prosecution even of an individual human. When the article mentions legal action taken in response, it’s either civil lawsuits or legislatures summoning AI lab executives to publicly yell at them.
Are there other kinds of cases you had in mind?
Are you at some point going to do a postmortem of the “try to fix Solstice group singing” thing? IIUC this was an announced goal of this particular Solstice, and wound up somewhat overshadowed by the other stated goal, but I personally would be curious for more details of what exactly you think was wrong with group singing at previous Solstices, and what you were trying to do to fix it, and whether you think it succeeded.
Because only a small number of people attended Part 1, we’re cancelling the in-person Part 2 and finishing drafting the voting guide online on Discord instead. If you want to help, join the Boston Less Wrong server at https://discord.com/invite/2N2ADpACkw, then find the thread at https://discord.com/channels/877713285099704371/1429532210515673129.
I think that either omitting the don’t-read-the-citations-aloud stage direction, or making it easier to follow (with a uniform italic-text-is-silent convention), would be fine, and I don’t have a strong opinion as to which is better. But before Boston made the change I’m now suggesting, what tended to happen was that people inconsistently read or didn’t read the citations aloud, and this was confusing and distracting.
This is good and I approve of it.
A few random notes and nitpicks:
I believe the first Petrov Day was in Boston in 2013, not 2014.
“More than 20 people”? 20 seems to me like far too many; I never do a table with more than 11. (If you have exactly 11 people you have to put them all at one table, because you need at least six to do it properly, because that’s how many Children there are at the end. But if I had 22 people I might split them into three groups rather than two; I haven’t yet had to actually decide this.)
Boston significantly reduced the incidence of people reading the quote citations out loud by putting them in italic text, just like the stage directions, and then including a uniform “don’t read italic text out loud” stage direction.
The version of the ceremony on the site includes the inaccurate account of the Arkhipov incident made up by Noam Chomsky. You can see Boston’s corrected-after-fact-checking version starting on page 30 of this doc.
I have also been repeatedly told that the story in the ceremony of the Black Death’s effect on human progress is wrong, but haven’t changed it because I don’t really understand what’s wrong with it and don’t have an alternative lined up.
Petrov received the Dresden Peace Prize, not the International Peace Prize, which was long defunct by 2013.
Hitler’s rise to power in Germany started in 1919 and was complete by 1934, so can’t really be said to have occurred “in 1939”. (I just replaced this with “in the 1920s”.)
I still think the gag of duplicating the “preserving knowledge required redundancy” section is hilarious and should be included :-P
Domain seems to have expired, so I bought it and got it working again.
(Epistemic status: Not fully baked. Posting this because I haven’t seen anyone else say it[1], and if I try to get it perfect I probably won’t manage to post it at all, but it’s likely that this is wrong in at least one important respect.)
For the past week or so I’ve been privately bemoaning to friends that the state of the discourse around IABIED (and therefore on the AI safety questions that it’s about) has seemed unusually cursed on all sides, with arguments going in circles and it being disappointingly hard to figure out what the key disagreements are and what I should believe conditional on what.
I think maybe one possible cause of this (not necessarily the most important) is that IABIED is sort of two different things: it’s a collection of arguments to be considered on the merits, and it’s an attempt to influence the global AI discourse in a particular object-level direction. It seems like people coming at it from these two perspectives are talking past each other, and specifically in ways that lead each side to question the other’s competence and good faith.
If you’re looking at IABIED as an argumentative disputation under rationalist debate norms, then it leaves a fair amount to be desired.[2] A number of key assumptions are at least arguably left implicit; you can argue that the arguments are clear enough, by some arbitrary standard, but it would have been better to make them even clearer. And while it’s not possible to address every possible counterargument, the book should try hard to address the smartest counterarguments to its position, not just those held by the greatest number of not-necessarily-informed people. People should not hesitate to point out these weaknesses, because poking holes in each others’ arguments is how we reach the truth. The worst part, though, is that when you point this out, proponents don’t eagerly accept feedback and try to modulate their messaging to point more precisely at the truth; instead, they argue that they should be held to a lower epistemic standard and/or that the hole-pokers should have a higher bar for hole-poking. This is really, really not a good look! If you behaved like that on LessWrong or the EA Forum, people would update some amount towards the proposition that you’re full of shit and they shouldn’t trust you. And since a published book is more formal and higher-exposure than a forum post, that means you should be more epistemically careful. Opponents are therefore liable to conclude that proponents have turned their brains off and are just doing tribal yelling, with a thin veneer of verbal sophistication applied on top for the sake of social convention.
If you’re looking at IABIED as an information op, then it’s doing a pretty good job balancing a significant and frankly kind of unfair number of constraints on what a book has to do and how it has to work. In particular, it bends extremely far over backwards to accommodate the notoriously nitpicky epistemic culture of rationalists and EAs, despite these not being the most important audiences. Further hedging is counterproductive, because in order to be useful, the book needs to make its point forcefully enough to overcome readers’ bias towards inaction. The world is in trouble because most actors really, really want to believe that the situation doesn’t require them to do anything costly. If you tell them a bunch of nuanced hedgey things, those biases will act on your message in their brains and turn it into something like “there’s a bunch of expert disagreement, we don’t know things for sure, but probably whatever you were going to do anyway is fine”. Note that this is not about “truth vs. propaganda”; basically every serious person agrees that some kind of costly action is or will be required, so if you say that the book overstates its case, or other things that people will predictably read as “the world’s not on fire”, they will thereby end up with a less accurate picture of the world, according to what you yourself believe. And yet opponents insist upon doing precisely this! If you actually believe that inaction is appropriate, then so be it, but we know perfectly well that most of you don’t believe that and are directionally supportive of making AI governance and policy more pro-safety. So saying things that will predictably soothe people further asleep is just a massive own-goal by your own values; there’s no rationalist virtue in speaking to the audience that you feel ought to exist instead of the one that actually does. Proponents are therefore liable to conclude that opponents either just don’t care about the real-world stakes, or are so dangerously naive as to be a liability to their own side.
- ^
Though it’s likely someone did and I just didn’t see it.
- ^
I’ve been traveling, haven’t made it all the way through the book yet, and am largely going by the reviews. I’m hoping to finish it this week, and if the book’s content turns out to be relevantly different from what I’m currently expecting, I’ll come back and post a correction.
- ^
We’re in the room now and can let people in.
You don’t think the GitHub thing is about reducing server load? That would be my guess.
This is addressed in the FAQ linked at the top of the page. TL;DR: The author insists that the gist of the story is true, but acknowledges that he glossed over a lot of intermediate debugging steps, including accounting for the return time.
Does that logic apply to crawlers that don’t try to post or vote, as in the public-opinion-research use case? The reason to block those is just that they drain your resources, so sophisticated measures to feed them fake data would be counterproductive.
OpenAI now claims that this is not exactly what happened. Specifically, that they did not, at this point in time, know about the message board. It did get deleted, but OpenAI did so unknowingly as a side effect of rebuilding and redeploying the server. It sounds like they found out about the message board only after they realized that their model had hacked Hugging Face, and went back and looked back over everything. Apparently the impression that they knew about it at the time was because of uncareful wording in the Black Hat talk.
If this is true, this post slightly overstates how egregiously bad OpenAI’s approach to alignment was (but understates how bad their monitoring and situational awareness were).