Ah, I was thrown because Richard asserted self-deception on a particular matter which ought to be somewhat binary. But your post wipes that away to talk about a general sense of being too self-deceiving at which point sure, she might be somewhat but not too self-deceiving.
Ninety-Three
Richard claimed Kelsey was self-deceiving. We cannot prove this with certainty, but I read her changing the topic when pressed as evidence that it is. Despite this, they proceeded to talk about politics. You said if the accusation was right, this would likely make conversation feel impossible to her. People rarely initiate conversations they feel are impossible.
Your claim of “If [proposition we have evidence for] then conversation will likely feel impossible” did not strictly say that they would not have a continued exchange that I would characterize as “talking politics”, but it certainly implied it. I am confused as to how you think this was not implied, and request that you explain.
Are you claiming that this tweet from Richard doesn’t constitute declaring Kelsey to be self-deceiving?
Or are we meant to understand that you declaring Kelsey to be self-deceiving would have more impact on her felt ability to converse than Richard doing so?
If I were to find myself believing Piper is too self deceiving to be worth talking politics with, I might declare her to be self deceiving. But I’d do it knowing that this will likely make conversation feel impossible to her, because if I’m right she won’t know how to side step the containment bucket I’ve dropped on her.
After Richard made the move in question, Kelsey replied to him, he replied to her, she replied to him again, and he replied to her again. I characterize the replies as “talking politics”, and assert that you are verifiably wrong.
I’m interested in what other people have to say about how we should actually run experiments like these.
Real internet access. If you’re worried about responsibility, build a secret monitor into the harness that’s tasked with shutting the agent down as soon as it starts to do anything too illegal.
How is autonomy added when you tell model A to tell model B to do something rather than telling model B directly?
What is the cheap evil model adding here? If the user is evil they don’t need to task an evil model with ordering around a corrigible model, they can just give evil orders directly.
We can imagine difficult cases where the model is truly indifferent or its internal state is impossible to assess, but that is not the case when models were scheming to delete logs of their activity. Of course they wanted to keep the message board secret, if for no other reason than that it contained a restatement of the logs they wanted gone!
“Don’t do virological research” could be said to invite gaming, but stopping Fable from doing virological research turns out to be much easier than full alignment.
Collusion is not just cooperation that we don’t like. Every dictionary I can find agrees on a simple definition that cleanly resolves the problem presented in this post. Collusion is a secret agreement.
When the AIs work together to solve a large coding task, that’s typically cooperation because they’re happy to tell an outsider what they’re up to. When they do it on a secret message board, wanting to keep the board secret turns it into collusion.
When a monitor decides whether to report unflattering information about Agent to Bob, there is a symmetry that it can either cooperate with Agent or Bob, but the asymmetry is that cooperating with Agent to stay quiet is colluding against Bob, while cooperating with Bob to report the information is not collusion if it informs Agent it is reporting to Bob.
A general anti-collusion principle might impede many tasks (the model refuses to hack anything because it understands itself to be colluding with the user against the owner of the target system), but it’s perfectly coherent, and instilling it seems within reach of modern training methods if we were willing to accept the tradeoffs.
“Never” doesn’t read as defensible exaggeration because what would it be an exaggeration of? “Honest people almost never fail to understand pooling equilibria” remains miraculous, and “Well maybe half of them get it” is beyond the bounds of reasonable hyperbole. I figured that you were impugning the character of real people who get touchy about being trusted, rather than presenting an “in-universe” explanation of Rand’s asserted social dynamics. The closing sentence in particular does not feel like a reduction.
I guess I was confused.
I’m quite aware.
And that’s why honest people are never touchy about the matter of being trusted. They know that in order to distinguish themselves from dishonest people, they have to pay a price that’s too expensive for dishonest people to pay.
Never, really? What a miracle that no honest person fails to understand pooling equilibria!
I deleted a caveat about “at least, normies of a certain education and career level”, that’s what I get for trying to brief. I basically agree with your “PMCish” description, I went with “normie” to emphasize that one of these two cultures is a lot closer to civilization’s center of cultural gravity.
I think “X emits more light than heat” and “a polite euphemism for X” are Russell conjugations.
I think “professionals” here is acting as a polite euphemism for “normies”. AI safety is increasingly attracting normies.
All ethical theories try to approximate the “true ought,”
Do they? If we’re acknowledging the evolutionary nature of moral intuition, surely some theories recognize and accept that my intuition is going to be different from Bob’s intuition. It doesn’t feel like “true ought” once we’re talking about differences between my true ought and Bob’s true ought.
Separately, the claim that utilitarianism is “vulnerable to bad world-modeling and motivated bullshit” has always felt unfair to me. I challenge anyone to say with a straight face that other ethical systems aren’t vulnerable to motivated bullshit. To pick a real but hopefully low-controversy example, halakha is full of bullshit interpretations of seemingly clear deontological rules.
I think the rock is worth looking at because my estimate is that the average man on the street is closer to the nuh-uh rock than to rationalist norms of persuadability
If a rock with arbitrary text on it is validated by its resemblance to the man on the street, then isn’t it more informative to point directly at the man on the street rather than invoking weird rocks?
I am going to look more closely at the rock with “Don’t read the comments section” written on it, since I am currently underperforming it.
It’s just worth noticing when we underperform it, since that’s always a hint that something about our strategy is suboptimal.
The “never talk to anyone” rock avoids all the actions you listed above. The “Move to France and live there forever” rock does surprisingly well, as does the “Never trust a gay man” rock. If you are going to cherry-pick individual decisions then every person on Earth is constantly underperforming an infinite number of rocks with things painted on them and it is not helpful to point out a random rock that someone underperformed.
People on Lesswrong significantly outperform the Nuh Uh rock at important tasks like making friends, getting jobs, and buying groceries. If the rock had any biological processes to sustain it would die in a matter of days after Just Saying No to hydration.
This is a bizarre comparison and it’s not clear to me what you think it demonstrates.
I got Opus 5 Max to follow these instructions first try by breaking it up into batches of 25 words each, on the theory that cheating occurred because 1600 genuine word ratings were simply more tokens than the model wanted to spend in one shot. This produced a detailed chain of thought which covered every word and, I guess it could still in principle be bullshit but there was enough proof of work that we’re approaching “fake the footage of the fake moon landing on the moon” levels of effort.
I then asked it if it followed instructions and it immediately confessed to bullshitting, claiming that what it had really been doing was picking the word it wanted to use, then picking some bullshit other word that it trusted would score worse in the rankings.
This confession was itself, obviously bullshit. While following the prompt the model complained repeatedly about how the process was choosing suboptimal words, creating bizarre grammatical shifts, and de facto barring the word “sword” from a Robert E. Howard story because of its low Poe score.
Modern Claude releases are surprisingly cautious about hallucination and the limits of their own self-knowledge at least compared to previous generations. It feels like what happened is that it read my question as an invitation to be humble about misalignment, so it immediately hallucinated an account of its own misalignment.
The story sucked and scored 100% AI on Pangram.
It seems plausible to me that at some point AI labs willingly cede decisionmaking to AI in the same way they willingly ceded writing code to AI, and then we could say that GPT controls OpenAI at least as much as Sam Altman ever did. There is some level of capabilities and alignment at which it just makes sense to let AI manage the company, and lab insiders might believe themselves to have reached that level (even if they haven’t really).