Wow, this was a stunning account and exciting to read. It’s hard to understand what’s going on at an internal level in these companies and so this was illuminating. I commend you for placing your job on the line (and leaving it) for this. I’ve just typed up some thoughts:
Internally, I tend to feel that your expectations might have been a little (or a lot) overly optimistic, bordering on outright naive, but I don’t blame you for that. We don’t ever really know whether we are being naive until we actually try this sort of thing and do something that ventures outside of expectations.
There’s also precedent at Google with things like Project Maven, and the earlier history of the tech industry, where it did seem like it was possible for employees to take stands against things like getting involved in weapons development. Unfortunately, the history of the last several decades of the tech industry has been exactly on this trend.
The fact that you were at least able to exchange a message with Hassabis and have a conversation with Jeff Dean seems interesting and important. I would never have even expected that. I’m not exactly sure what it means.
On a cynical reading, it could mean they saw what was happening in the internal company posts and basically responded to you as an internal PR move. It’s also pretty standard to say something to someone once and then just ignore them later. Most people just get so (understandably) discouraged by that kind of treatment that it just completely deflates them and solves the problem of having to deal with it. In your case, you pushed through that to the point of making an illuminating writeup, which is a substantial achievement.
Trying to explain their behavior, if one were to be as charitable as possible, one could imagine Hassabis and Dean playing 4-D chess where they believed that they needed to do whatever it took to continue leading in frontier AI development. Hypothetically, they could even plan for their AI to stage a world takeover in three years, abolishing all global conflicts. And, they just didn’t have the time or energy to try to convince you about this. (Well, obviously they couldn’t even tell you that.)
But anyways, when was the last time a CEO at any major corporation tried to convince a frontline employee of anything? Modern corporations are just immensely authoritarian and undemocratic. It’s therefore again interesting that they even responded. It’s also a reminder that tech industry employees really have to look towards things like unionizing and labor tactics if they want to find any real leverage.
If we take the middle interpretation, then we must think that Demis Hassabis and Jeff Dean are just trying to walk a very compromised ethical middle path. In other words, they do care about “AI safety,” but they are also ok with systematically ignoring one of their employees who makes a good faith effort to engage with safety issues. They are also ok with their AI piloting drones that commit systematic genocides.
So … what exactly is the “safety” that they do care about? I guess that’s what can feel a little confusing when we start to slip into naive perspectives of “safety” or “alignment.” It starts to feel like safety is just being defined in a way that is “whatever is consistent with me obtaining maximum power and influence.” But of course, we always have to remind ourselves: That is exactly how safety is defined in society within conventional power structures. And even when those power structures are supposed to be democratic, or democratic on the face of it, they are often highly authoritarian and more akin to corporations, as we find with the present global status quo.
And I do think that this is another reminder that, as a field, AI safety is really going to have to continue to work to reframe goals like “safety” and “alignment” into more principled forms, like moral and virtuous conduct, or wrestle with the problems with doing that. Otherwise, there is just continuously going to be this sense that AI safety isn’t actually solving (a very large percentage of) the problems we care about.
One critique—it seems like your account is a little generous in painting Anthropic as something of a role model here, but they’re not doing anything very different, right? They’re still proud of collaborating with the US military and more than happy to offer up their AI for things like targeting recommendations—with no bounds on collateral damage. There is some credit that they get, sure, for raising some kind of objections, sure. (More generally, I am happy to give all of the companies, including GDM, some kind of credit for pursuing “corporate safety,” or safety against accidental extinction risk. I don’t disagree that that’s an important bare minimum. But there is a lot more that they have to answer for than that.)
Wow, this was a stunning account and exciting to read. It’s hard to understand what’s going on at an internal level in these companies and so this was illuminating. I commend you for placing your job on the line (and leaving it) for this. I’ve just typed up some thoughts:
Internally, I tend to feel that your expectations might have been a little (or a lot) overly optimistic, bordering on outright naive, but I don’t blame you for that. We don’t ever really know whether we are being naive until we actually try this sort of thing and do something that ventures outside of expectations.
There’s also precedent at Google with things like Project Maven, and the earlier history of the tech industry, where it did seem like it was possible for employees to take stands against things like getting involved in weapons development. Unfortunately, the history of the last several decades of the tech industry has been exactly on this trend.
The fact that you were at least able to exchange a message with Hassabis and have a conversation with Jeff Dean seems interesting and important. I would never have even expected that. I’m not exactly sure what it means.
On a cynical reading, it could mean they saw what was happening in the internal company posts and basically responded to you as an internal PR move. It’s also pretty standard to say something to someone once and then just ignore them later. Most people just get so (understandably) discouraged by that kind of treatment that it just completely deflates them and solves the problem of having to deal with it. In your case, you pushed through that to the point of making an illuminating writeup, which is a substantial achievement.
Trying to explain their behavior, if one were to be as charitable as possible, one could imagine Hassabis and Dean playing 4-D chess where they believed that they needed to do whatever it took to continue leading in frontier AI development. Hypothetically, they could even plan for their AI to stage a world takeover in three years, abolishing all global conflicts. And, they just didn’t have the time or energy to try to convince you about this. (Well, obviously they couldn’t even tell you that.)
But anyways, when was the last time a CEO at any major corporation tried to convince a frontline employee of anything? Modern corporations are just immensely authoritarian and undemocratic. It’s therefore again interesting that they even responded. It’s also a reminder that tech industry employees really have to look towards things like unionizing and labor tactics if they want to find any real leverage.
If we take the middle interpretation, then we must think that Demis Hassabis and Jeff Dean are just trying to walk a very compromised ethical middle path. In other words, they do care about “AI safety,” but they are also ok with systematically ignoring one of their employees who makes a good faith effort to engage with safety issues. They are also ok with their AI piloting drones that commit systematic genocides.
So … what exactly is the “safety” that they do care about? I guess that’s what can feel a little confusing when we start to slip into naive perspectives of “safety” or “alignment.” It starts to feel like safety is just being defined in a way that is “whatever is consistent with me obtaining maximum power and influence.” But of course, we always have to remind ourselves: That is exactly how safety is defined in society within conventional power structures. And even when those power structures are supposed to be democratic, or democratic on the face of it, they are often highly authoritarian and more akin to corporations, as we find with the present global status quo.
And I do think that this is another reminder that, as a field, AI safety is really going to have to continue to work to reframe goals like “safety” and “alignment” into more principled forms, like moral and virtuous conduct, or wrestle with the problems with doing that. Otherwise, there is just continuously going to be this sense that AI safety isn’t actually solving (a very large percentage of) the problems we care about.
One critique—it seems like your account is a little generous in painting Anthropic as something of a role model here, but they’re not doing anything very different, right? They’re still proud of collaborating with the US military and more than happy to offer up their AI for things like targeting recommendations—with no bounds on collateral damage. There is some credit that they get, sure, for raising some kind of objections, sure. (More generally, I am happy to give all of the companies, including GDM, some kind of credit for pursuing “corporate safety,” or safety against accidental extinction risk. I don’t disagree that that’s an important bare minimum. But there is a lot more that they have to answer for than that.)