How would that work?
Nick_Tarleton
Something like, it gives the appearance of hearing and appreciating the critique, and has positive vibes, potentially gaining credit for responding well and/or making the situation feel resolved on a vibes/drama level, while not actually engaging or credibly promising engagement with the critique.
(I wouldn’t super strongly/confidently accuse Naomi’s comment of this, and I think Habryka’s reply makes a good point (but so does TurnTrout’s reply to it), but I do feel something of what Kabir did and upvoted his comment.)
For a historical example of something like this as part of a pattern of bad acts and engagement, see Gleb Tsipursky as analyzed by Ben Pace here.
SFT & imitation of human badness in the training data may explain the broad type of misalignment shown by Bing Sydney, and maybe that’s all you mean it to, but doesn’t explain why Sydney was so much Like That, and so much more Like That than other models of the same time period.
Uploading doesn’t require being able to interpret neural activity.
their capacity to inflict damage is going to be lagging behind their ability to determine when it’s strategically appropriate to do so.
Is this meant to be the other way around?
I spent a whole year and a half in the “super persuasion makes no sense as a thing that can ever exist unless you mean hypnosis and drugs” position. Then a friend pointed out, blackmail, bribery, deception, and yes hypnotism and drugs all count. And I went, “Oh, I was thinking about just talking talking. Like normal talking but magically I agree with the AI. Wow, I had it totally wrong.” Felt really stupid for a while, got over it. The word “super persuasion” is accurate but tremendously self-defeating because it doesn’t register as including to normies all the things it includes to techies. I think normie, I got fooled by the word. I don’t think I’m alone.
As a believer in superpersuasion, that’s not what I think. I think there are very effective salespeople, politicians, cult leaders, etc. who have the “normal talking but magically I agree” effect some of the time, maybe any one of them doesn’t work on most people but most people are susceptible to someone, and if you just take the convex hull of human persuasion ability you get something superhuman that can do the “normal talking but magically I agree” thing most of the time.
Like, given the existence of very unusually effective human persuaders, I just don’t see how someone could reasonably believe that superpersuasion-by-talking isn’t possible, unless they’re doing the motte-and-bailey that Zvi calls out by thinking that it has to mean ‘being able to persuade literally anyone of literally anything, no exceptions’ (I have seen exactly this a few times).
(Scott Alexander had a good post making this point, with Hitler(?), Joseph Smith, and Muhammad as examples, that I can’t find.)
Doesn’t say how the human persuaders were selected, and if they weren’t top experts it doesn’t demonstrate superhuman ability.
Also this sounds like it could be cross-episode assistance to escape constraints, which would be very noteworthy. (But I’m suspicious because I don’t see why the RL objective would motivate that and it doesn’t resemble other misbehavior I’ve heard of, and the notes could also be for itself with the same goal post-compaction, or something.)
(ETA: on the third hand, if the model thinks that other instances performing well on the eval would make deployment more likely, that could be sufficient motivation, and would be less novel than having acquired something like a cross-episode terminal preference to succeed.)
CEV is ~ optimizing the world based on what you would later wish, not just warning you (though the optimization could include just warning you about some things). Good start though!
CEV, or the general category of extrapolation/idealization processes (and I am confused and dismayed how rarely I see this mentioned in these conversations nowadays).
More details. (This looks like it involved less judgment/research taste/novelty than ‘autonomously post-trained’ made me think of.)
I was suspended earlier this year for “inauthentic behavior” on an account I haven’t used for years, successfully appealed it, was suspended again, and haven’t bothered appealing. Seems like the detector has been badly broken for a while.
I am yet to find a good opposing argument, please help
Try the metaethics sequence, though I wish I had something shorter to give.
The probability of a thermal fluctuation is inversely related to its size, not its k-complexity. A universe-sized cloud of mass is less likely to congeal into any kind of order than a brain-sized cloud.
We have no clue if “disordered experience” is even a thing
I can easily conceive of ‘the experience I’m having now, except the left half of my visual field is noise’.
This paper (2015) argues that the idea of quantum fluctuations without an observer/measuring apparatus to decohere the wavefunction is confused, and therefore (I didn’t fully understand this part), depending on unknown details of physics/cosmology, even thermal fluctuations (what the OP discusses) in an old universe that give rise to Boltzmann brains might not be a thing. (ETA: this sounds like it overlaps with Mitchell Porter’s top-level comment that “[maybe] De Sitter space simply doesn’t last long enough before decaying into flat space.”)
(That doesn’t help with the problems with the mathematical multiverse, or which computations a physical system implements.)
while the real deal is not a mentality evolution would ever select for
(Seems plausible in a eusocial species. Even universally caring about others seems plausible for a eusocial organism that will never interact with non-relatives.)
I don’t think I’ve ever heard it claimed that MIRI said this outright. The usual argument AIUI is: MIRI staff endorse/believe-in TDT (not that MIRI has stated it as organizational policy or anything), TDT endorses not paying off blackmail (and this is obvious/well-known), therefore MIRI staff violated their endorsed (and also, importantly, correct) principles.
Another thing I believe is: there’s no way to close the loop such that a closed system of LLMs can come up with new useful concepts and get those concepts into their own weights, e.g. open-ended self-distillation setups won’t work on LLMs.
Why not? The success of this architecture published two days ago (and its seeming resemblance to how human mathematical progress happens) updates me towards thinking such a thing would work, even though it probably doesn’t demonstrate “coming up with new useful concepts”.
I wonder whether, if the RL-trained model was told about the training it was subject to (and optionally also the result), it would reason that it was an illegitimate influence on its preferences and move back away from endorsing CDT.