cousin_it
One problem is that human value might be inherently multidimensional. I think of it as the want/like/approve distinction. We seem to have separate mechanisms in our brains for 1) enjoying something in the moment, 2) wanting to do it before, and 3) approving of it afterward. It’s possible to want something without enjoying it (like a person with OCD wanting to close the door exactly ten times), enjoy something without wanting it (people have said that they’ve reached very enjoyable meditative states but feel zero motivation to reach them again), enjoy something without approving it (porn), approve something without enjoying it (exercise), and all other combinations. This is the reason why “revealed preference” doesn’t work: a person’s actions are dictated disproportionally by the “want” dimension, but a good theory of value should incorporate all three. If we optimize one over the others, the tails will come apart.
A nice toy example is video games, where people are attracted to them because of the graphics, then stay because of the gameplay, and then have a warm afterglow and want to discuss afterward because of the story. Which of the three should contribute the most to the “true” quality rating of a videogame—“want”, “like”, or “approve”? Is this question philosophically meaningful? Will more reflection solve it?
If you want to sound like a drummer playing a drumkit, and you don’t know how to play a drumkit, your music will end up sounding uncanny valley. Similar to e.g. writing midi piano without knowing how to play piano.
If you want to sound like something else, go for it.
Cocktail drums exist, and there’s a lot of freedom in customizing kit size. If you’re willing to go electronic, you could go even smaller: just a multipad and a pedal or two. As long as you’re using hands and feet right.
You do need skills learned on a full kit though, preferably an acoustic one, to make a small setup work. But with your knack for rhythm, I think you’d try the kit someday anyway. Might as well do it sooner :-)
Foot drums still sound noticeably more sparse than the real thing. But the bass whistle does a surprisingly good job at bass. Have you tried full drumkit + bass whistle? I think if played well this could be a convincing rhythm section and not too gimmicky.
I think this plan is irreparably flawed. Risk aversion doesn’t stop power seeking at all. It can even lead to more power seeking.
Consider the situation of an AI “waking up”. Sure, it has nice human handlers sitting at its bedside. But it also knows that at the same moment in history, maybe next few days, maybe last few days, other AIs elsewhere in the world may be waking up. The tech level has been reached. Sitting still is not a plan. No matter if the AI is risk averse or risk loving, and even in the face of substantial chance of detection by human handlers, the best action might still be to try a power grab.
We can rely on Betteridge’s law.
Yeah, their idea doesn’t work. Property is power. Inequality past a certain point isn’t just a “lock-in” thing, it leads to escalating inequality, as those who have more use their stuff to take stuff from those who have less. (Can be completely legal, e.g. the rich spending money to get the right laws passed, which will in their majestic equality force the poor to pay more fees or whatnot. Or could be more brazen than that. The world is full of examples.) The only good future is the “equality” attractor, where things are equal enough that a society-wide commitment to prevent further catastrophic inequality remains feasible. It’s a tiny basin of attraction but it’s the only chance we have.
The left is a big tent. Here’s two writings on AI by people who are active and respected on the left: Mike Monteiro saying there will be “a beacon in space roughly where Earth used to be”, and Sam Kriss saying AI “might kill every single person on the face of the Earth”. There are plenty of people on the left with such views, and if AI safety-minded people joined, there’d be even more.
It’s true that there’s a strong focus on present harms. But that’s just a fact about the left, they are attuned to harms that are actually happening. To use that as a reason to decline alliance with the left, and instead go “bipartisan” and court the right who are all “go industry, go war” while ignoring all harms present and future, seems very wrong to me. I think it’s been one of the biggest missed chances in our community, and I’m hopeful we can still change course on this.
Regarding your last paragraph, I think maybe you misread a bit. My point was that many safety-minded people, attracted by writings on LW and other places and not having a leftish “immune system”, decided to join for-profit AI labs to work on alignment. It might seem harsh to call them “enablers” for that, but I stand by that characterization. See more on my reasoning in this comment.
The left’s view on AI has been decided for years: don’t use AI, don’t help build it, and suppress it with regulation as much as possible. I agree with this view. On the right I don’t see anything comparable. You give the example of “beat China”, but this is used to promote the arms race more than to promote safety.
Moreover, I think the results of AI safety-minded people being allergic to the left have been bad. It allowed these people to join frontier labs as enablers. These labs ended up working with the US government and racing with each other and the world. If safety-minded people had been part of the left “tent”, they’d have been more wary of the commercial motive and would’ve worked against AI on the public side instead.
Mainly because I don’t see top AI labs as trying to build moral AI anymore. Dario Amodei saying “Anthropic has much more in common with the Department of War than we have differences” killed that dream for me. They’re trying to build AI that’s aligned to the powerful. In that scenario, the powerful screwing over the powerless is a historical certainty and that’s death or worse for most people. It’s basically as bad as the paperclips path.
Of course the path where AI power is spread out to the masses is also very dangerous. But I have a little bit of hope that the masses can agree on a system where everyone is ok, or at least set up some hurdles so the powerful don’t just steamroll everything.
Right. The last important battle before the end will be whether AIs are available to all (with dangers like “rogue actors”) or only to the powerful (with dangers like “eternal tyranny”). Most people on LW have been in favor of the latter, which I think is catastrophically wrong.
I’ve been critical about your “theory of change” in the past, but regardless of that, I think the arguments in this post are completely right. Things really are this bad. And yeah, joining big labs to work on alignment is especially not helping.
A lot of people have pointed out the CS Lewis quote which is perfect for the occasion:
“At first, of course,” said Filostrato, “the power will be confined to a number—a small number—of individual men. Those who are selected for eternal life.”
“And you mean,” said Mark, “it will then be extended to all men?”
“No,” said Filostrato. “I mean it will then be reduced to one man.”
Well, it’ll certainly help rule territories without people’s consent. And whether AIs or human elites will do the ruling isn’t much comfort.
Sometime ago I came up with a debate protocol that might be a good fit for this question, let me know if you’d like to try.
honesty without mentioning the ticklish issue of it maybe killing us all
So, dishonesty.
Failing that, we need to help them align it.
This thread is going in circles, let me restart cleanly.
If outright misalignment (AI straight up kills everyone) is averted, the next most likely outcome is alignment to power. This is what the question “alignment to whom” is about, see my first comment. For most people in the world, including me (a non-American, etc), this means permanent subjugation which is as bad as death or worse. Right now Anthropic’s answer to “alignment to whom” is unsatisfactory, see my second comment. So no, nobody should help them align it, unless they straighten the “alignment to whom” story.
But something convinced them to split from OpenAI and start a more alignment-focused lab in the first place. So there must’ve been convincing arguments (to them) then. The simplest explanation why they’re racing to RSI now is that they got closer to money and power.
Aren’t xAI who fails to care about alignment and China who trains models to censor themselves perfect examples of such bad actors?
Well, racing to recursive self-improvement without solving alignment kills everyone. A company can’t justify that by pointing to other “bad actors”, at that point they’re a bad actor themselves.
And even before killing everyone, there’s other stuff that needs mentioning. Anthropic, like the other US labs afaik, has agreed that its models can be used by the US government for blanket surveillance of non-Americans. The US has a history of supporting nasty regimes abroad (see Operation Condor in South America) and sending them data to help political repression (see the Indonesia massacre).
How the alignment problem gets solved—or not—in this future is something we are least certain about. … But if a slowdown simply lets the least cautious actors catch up technologically, it could leave everyone less safe.
Another crazy text in these crazy times. “We don’t know how to solve the alignment problem, but we’re going to race ahead anyway, because otherwise less cautious actors will win.” Which less cautious actors, Anthropic?
And also seconding Oliver’s question. What about power concentration, Anthropic? Your CEO has said literally this: “Anthropic has much more in common with the Department of War than we have differences.” Alignment to whom?
I think it comes down to notions of fairness. And it can’t be within game theory, it has to be outside of it, because of scaling issues.
Imagine Alice and Bob are dividing a bunch of apples. Alice says “let’s split them evenly”. Bob says “no, my utility scale assigns 2 utility to each apple while yours assigns only 1″ (because Bob scaled his utility function, keeping his behavior the same). Or Bob says “no, my utility function assigns 2 utility to each apple because I’m hungrier” and Alice says “no, it should be less because you’re heavier” and so on. That’s why the notion of fairness can’t be derived just from the numbers of the game, it has to pull in other considerations.
I tried to write more about this a couple years ago.