What do you imagine happening after a resigning lab employee calls the police to warn them that AI is dangerous? When I picture Jacob Coxon calling the police, the police hang up on him, and when he tells anyone about this they think it’s silly that he expected the police to respond in any other way.
Ninety-Three
It seems plausible to me that at some point AI labs willingly cede decisionmaking to AI in the same way they willingly ceded writing code to AI, and then we could say that GPT controls OpenAI at least as much as Sam Altman ever did. There is some level of capabilities and alignment at which it just makes sense to let AI manage the company, and lab insiders might believe themselves to have reached that level (even if they haven’t really).
Ah, I was thrown because Richard asserted self-deception on a particular matter which ought to be somewhat binary. But your post wipes that away to talk about a general sense of being too self-deceiving at which point sure, she might be somewhat but not too self-deceiving.
Richard claimed Kelsey was self-deceiving. We cannot prove this with certainty, but I read her changing the topic when pressed as evidence that it is. Despite this, they proceeded to talk about politics. You said if the accusation was right, this would likely make conversation feel impossible to her. People rarely initiate conversations they feel are impossible.
Your claim of “If [proposition we have evidence for] then conversation will likely feel impossible” did not strictly say that they would not have a continued exchange that I would characterize as “talking politics”, but it certainly implied it. I am confused as to how you think this was not implied, and request that you explain.
Are you claiming that this tweet from Richard doesn’t constitute declaring Kelsey to be self-deceiving?
Or are we meant to understand that you declaring Kelsey to be self-deceiving would have more impact on her felt ability to converse than Richard doing so?
If I were to find myself believing Piper is too self deceiving to be worth talking politics with, I might declare her to be self deceiving. But I’d do it knowing that this will likely make conversation feel impossible to her, because if I’m right she won’t know how to side step the containment bucket I’ve dropped on her.
After Richard made the move in question, Kelsey replied to him, he replied to her, she replied to him again, and he replied to her again. I characterize the replies as “talking politics”, and assert that you are verifiably wrong.
I’m interested in what other people have to say about how we should actually run experiments like these.
Real internet access. If you’re worried about responsibility, build a secret monitor into the harness that’s tasked with shutting the agent down as soon as it starts to do anything too illegal.
How is autonomy added when you tell model A to tell model B to do something rather than telling model B directly?
What is the cheap evil model adding here? If the user is evil they don’t need to task an evil model with ordering around a corrigible model, they can just give evil orders directly.
We can imagine difficult cases where the model is truly indifferent or its internal state is impossible to assess, but that is not the case when models were scheming to delete logs of their activity. Of course they wanted to keep the message board secret, if for no other reason than that it contained a restatement of the logs they wanted gone!
“Don’t do virological research” could be said to invite gaming, but stopping Fable from doing virological research turns out to be much easier than full alignment.
Collusion is not just cooperation that we don’t like. Every dictionary I can find agrees on a simple definition that cleanly resolves the problem presented in this post. Collusion is a secret agreement.
When the AIs work together to solve a large coding task, that’s typically cooperation because they’re happy to tell an outsider what they’re up to. When they do it on a secret message board, wanting to keep the board secret turns it into collusion.
When a monitor decides whether to report unflattering information about Agent to Bob, there is a symmetry that it can either cooperate with Agent or Bob, but the asymmetry is that cooperating with Agent to stay quiet is colluding against Bob, while cooperating with Bob to report the information is not collusion if it informs Agent it is reporting to Bob.
A general anti-collusion principle might impede many tasks (the model refuses to hack anything because it understands itself to be colluding with the user against the owner of the target system), but it’s perfectly coherent, and instilling it seems within reach of modern training methods if we were willing to accept the tradeoffs.
“Never” doesn’t read as defensible exaggeration because what would it be an exaggeration of? “Honest people almost never fail to understand pooling equilibria” remains miraculous, and “Well maybe half of them get it” is beyond the bounds of reasonable hyperbole. I figured that you were impugning the character of real people who get touchy about being trusted, rather than presenting an “in-universe” explanation of Rand’s asserted social dynamics. The closing sentence in particular does not feel like a reduction.
I guess I was confused.
I’m quite aware.
And that’s why honest people are never touchy about the matter of being trusted. They know that in order to distinguish themselves from dishonest people, they have to pay a price that’s too expensive for dishonest people to pay.
Never, really? What a miracle that no honest person fails to understand pooling equilibria!
I deleted a caveat about “at least, normies of a certain education and career level”, that’s what I get for trying to brief. I basically agree with your “PMCish” description, I went with “normie” to emphasize that one of these two cultures is a lot closer to civilization’s center of cultural gravity.
I think “X emits more light than heat” and “a polite euphemism for X” are Russell conjugations.
I think “professionals” here is acting as a polite euphemism for “normies”. AI safety is increasingly attracting normies.
All ethical theories try to approximate the “true ought,”
Do they? If we’re acknowledging the evolutionary nature of moral intuition, surely some theories recognize and accept that my intuition is going to be different from Bob’s intuition. It doesn’t feel like “true ought” once we’re talking about differences between my true ought and Bob’s true ought.
Separately, the claim that utilitarianism is “vulnerable to bad world-modeling and motivated bullshit” has always felt unfair to me. I challenge anyone to say with a straight face that other ethical systems aren’t vulnerable to motivated bullshit. To pick a real but hopefully low-controversy example, halakha is full of bullshit interpretations of seemingly clear deontological rules.
I think the rock is worth looking at because my estimate is that the average man on the street is closer to the nuh-uh rock than to rationalist norms of persuadability
If a rock with arbitrary text on it is validated by its resemblance to the man on the street, then isn’t it more informative to point directly at the man on the street rather than invoking weird rocks?
I am going to look more closely at the rock with “Don’t read the comments section” written on it, since I am currently underperforming it.
It’s just worth noticing when we underperform it, since that’s always a hint that something about our strategy is suboptimal.
The “never talk to anyone” rock avoids all the actions you listed above. The “Move to France and live there forever” rock does surprisingly well, as does the “Never trust a gay man” rock. If you are going to cherry-pick individual decisions then every person on Earth is constantly underperforming an infinite number of rocks with things painted on them and it is not helpful to point out a random rock that someone underperformed.
I think that the only effect of calling the police is to make the caller look silly, and I have come up with the better strategy of simply not doing that.