jimmy
Neither.
Nowhere in my comment did I say that Kelsey and Richard would not have a continued exchange that you would characterize as “talking politics”. Nor did I even imply it.
I suggest you read more carefully.
It is, maybe, an okay social convention to suppose “when somebody takes things personally, we presume they are guilty of something they don’t want to admit.” But it’s not literally true. People can also get defensive or emotional when they’re falsely accused. If you assume “everybody who’s uncomfortable talking to me about this touchy subject is Up To No Good”, you will be wrong a lot of the time. It is not at all uncommon for people to be insecure about things that are in fact fine. Particularly if they default to social submission rather than social dominance.
Maybe the defensive person didn’t take the cookie from the cookie jar, but they’re definitely not anticipating “I’m gonna just show the video footage and then people are going to apologize for the false accusations. Easy”.
More generally, their behavior shows that they anticipate other people not finding their perspective persuasive. They could still be right, and that’s important to remember. At the same time, when someone tells you “You probably won’t find this convincing”, that means something.
If you choose to assume that everybody who has a limit, beyond which they get uncomfortable, take things personally, etc, is not worth conversing with, then you are going to have a very biased view of the world.
That’s definitely true, but Zack’s criticism is narrower than that.
Everybody has limits. It’s easy to respect someone who says “I’m sorry, I can’t handle this conversation. I recognize that it means I might be wrong, that I’m not confident in my own ability to support my stances, and I don’t know what to do about that. I’ll have to figure this out on my own time. Maybe next time I’ll be ready to talk about this”.
It’s hard not to, actually.
Strong agree, strong upvote, and all that.
I’m going to push back on what I think is a small but very important kernel of truth in what Piper is saying.
I think the move of declaring someone else so deceived that their own understanding of their beliefs and motives should be rejected inherently makes conversation with them nearly impossible.
There’s a difference between putting 99% of your probability mass on “they’re self deceiving”, or even 99.99%, and declaring that they’re self deceiving. The latter is a move to round off the residual. It’s a move to shut out the possibility that they’re not self deceived from the conversation. It’s the implicit “The chance that I’m wrong about this is small enough to ignore, and not spend effort reasoning about”. Which, in their perspective which is in this residual, does make conversation impossible on any front where the alleged self deception completely invalidates the perspective.
Now, this doesn’t mean all conversation is impossible. It’s totally possible to say “No, fuck you, you can’t round me out of the conversation!” and bring spreadsheets to back it up. Or to say “Why do you believe that?” and see if their reasoning holds up. But it’s still a reentering of a conversation your perspective was just pushed out of.
If I were to find myself believing Piper is too self deceiving to be worth talking politics with, I might declare her to be self deceiving. But I’d do it knowing that this will likely make conversation feel impossible to her, because if I’m right she won’t know how to side step the containment bucket I’ve dropped on her.
If I were to find myself *wishing I could* talk politics with Piper, or people like piper, then it would be foolish of me to make any declarations. Instead, I would explore that residual. Probe the cracks, where I don’t see her words matching her behaviors. Bring the discrepancy to light so that whoever turns out to be wrong will see the conflict. Maybe I expect to know who will be doing the updating, but the work to close the gap is the same either way.
Finally, your thesis is based on a falsehood.
And that’s why honest people are never touchy about the matter of being trusted.
In my experience many honest people are VERY touchy about being trusted.
This can be true from a very low resolution view, but this low resolution obscures the dynamics that matter here.
The behavior itself, on the face of it, is 1) an attempt to get people to view them as honest and 2) an attempt to do so without engaging with the possibility that they might not be. It comes up when the other person has reason to view them as not honest. “No, we shouldn’t think that. It doesn’t matter the evidence, I don’t want to see it. You should just use my preferred conclusion”. It’s not exactly an act of honesty.
Perfect honesty is hard. In order to be honest all the time, you have to see through the lies you tell yourself, which you expect to be able to get others to buy into. This same drive to “be honest” becomes a drive to be dishonest the moment perfect honesty becomes harder to achieve than “people [myself included] accept the claim that I’m honest”—because it can only ever be implemented as the control of perception, and if your ability to see through the deceptions lags your ability to create them then Goodhart happens and honesty fails.
So this defensive behavior might correlate with honesty in other cases, if there’s more variance in drive to be seen as honest than there is in ability to see through self deceptions. But it also might not. It’s an empirical question.
What looking closely gets you either way, is a recognition that regardless whether “I’m honest in all these other cases!” is true, the “therefore you shouldn’t look at the evidence on this case!” is itself an example of dishonesty.
So are honest people ever touchy about their honesty? Depends on how strict you’re being about giving them the badge of “honest”. You might want the bar to be low enough to give them a pass, and many might not clear the bar if you don’t. But people that are being honest don’t. And honesty at the point where honesty is questioned is necessary for limiting the Goodhart leakage.
For example, if your partner suddenly asks if you’re cheating on them, you shouldn’t necessarily just set your jaw and
My immediate response is “Wait, what? Why are you asking that?”. That would be very surprising to me. It wouldn’t make sense. If we flip it around, depending on how I ask I expect her to respond similarly or to laugh because she knows I can’t be serious.
It would be totally reasonable to react negatively to this question, because if your partner knows you and thinks highly of you, they should trust you not to do that.
If your partner knows you, and you aren’t the type that might cheat, then they will trust you to not do that. Just like how if you drop a ball you don’t need to say it “should” fall because it will fall—unless something weird happens, like it’s an alien tech antigravity ball, someone normally trustworthy has been making up lies about you, etc.
Turning it into a normative thing blocks the curiosity. “The ball should fall! Humph!”. It’s very odd to see something so supposedly unexpected, and then to not be curious what’s going on. Especially when the declaration doesn’t cause the ball to fall as it “should”.Now that they’ve expressed suspicion, you might need to do costly things like letting them go through your phone or performing grand romantic gestures to make them trust you.
I don’t see that as costly at all. I don’t care if my wife goes through my phone. There’s nothing in there I need to hide from her. She doesn’t care if I go through hers either.
What about “letting them go through your phone” feels costly to you? What do you lose, that you had the moment before you hand over the phone? It’s not “her trust in me”, since you already don’t have that.
If your partner isn’t surprised when you suspect them of cheating, and they’re not curious where that came from because according to them you shouldn’t even be entertaining it… regardless of who told you they’re cheating or how much evidence you have… what does that suggest?
The particular answers don’t matter much, and maybe the answers to the first questions don’t highlight anything interesting. People often do things for reasons that don’t make much sense in a larger picture, because people don’t always remember to look at the big picture. When you do though, these things aren’t going to hold up. You’re only ever going to find “Yeah, I guess I didn’t have anything to get defensive about” or “I did”.
Very good answer.
Adding onto this, I spent a couple weeks playing with polyphasic sleeping, with explicit intent to “learn how to fall asleep” as the main driving factor. When you’re doing 20 minute naps every few hours, that’s a lot like “practice falling asleep”. Fall asleep, wake up, do it again.
It was definitely helpful in teaching me what to sleep, and a large part of that is giving me an answer to “What does it feel like to succeed in falling asleep”. It blew my mind to realize that I could go days on what I had thought was “no sleep” because I had continuous experience. The answer is basically “Your mind slows down to a stop”, and you can experience that without loss of consciousness though your ability to tell time goes away because there are no events to count and use to estimate time.
Interestingly, I met an experienced meditator at an early LW meetup before this who said he had experience maintaining awareness throughout a full night’s sleep. It seemed crazy to me at the time and I couldn’t mentally represent what that’d be like, but hammering on “try to sleep” got me that experience in nap form.
Calling meditation “strength training for the mind” seems inaccurate to me.
There are some aspects of some styles of meditation that could be interpreted in this way, though mostly I think “skill/technique training for the mind” would be more accurate as a generalization.
It depends if we’re talking “meditation done right” or “meditation as practiced”. I agree that skill training is a more useful framing. I see a lot of people operating from a “strength training” framing, with predictable results.
How can you deliberately try? How do you know, and constantly keep in mind, what you actually care about? How can you ensure that you will do this consistently?
If the answers to these questions are clear for you, then your course is set.
If you’re less sure, meditation might be of some use to you in finding answers.
These are great questions. How does one deliberately try? How do you know what you care about? How can you make sure you keep your priorities in line?
These are exactly the kinds of question that call for mindfulness and stable attention, because they produce extremely valuable insights once answered. Knowing how to try? Knowing what you actually care about and how to act on that? Doesn’t get much more instrumentally useful than that.
And because meditation is best thought of as a skill, the interactions with the target of focus matter in ways that the “brute strength” framing obscures. If you want to come off as “strong” at jiu jitsu, gym-buildable muscles can help, but not as much as knowing what it feels like to apply your weight properly—and you can’t find that in the gym because there’s no “Shoulder on face” machine. The exact ways in which your attention gets destabilized from these important questions, and the moves that hold attention on the instrumentally useful targets, aren’t there to be found in the breath.
If the answers to these questions are clear to you, then you don’t need “meditation”. Your course is already set.
If you’re less sure, there is value to be found in practicing holding stable attention on these questions.
I think most of this can be summarized in the line “Professional athletes lift weights for good reason”, no? There are indeed good reasons to lift weights in judiciously chosen ways. I’ve done a fair bit of lifting myself, and still do some.
I can most appreciate this essay as a reminder for people who have been too stuck in school of the importance of the 12th virtue of rationality (keeping sight of the end goal).
Right, so are you working your grip strength because the rope got pulled out of your hand, and you realize that your grip needs more specific work until that stops happening? If your meditation practice involves running up against real limits, working the piece that you observed to fail, and testing with short feedback loops against the actual purpose, then you already get it.
Most people I see talk about meditation don’t have that. People will do it because “Science says it helps with anxiety”, or because “there is no self” sounds profound and fascinating. And then they go focus on their breath for countless hours, hoping that the skills somehow translate while being ironically mindless about the connection. Which is where, I argue, mindfulness counts the most.
Some meditators do pull it off—Dr.K can sure put his “holding thoughts as object” skills to impressive real world use. But you also get people with decades of meditation experience, who make meditation part of their identity, who nonetheless fuse with their thoughts in a stiff breeze. Without noticing.
And if we generalize beyond meditation to rationality, how many people can give a tightly coupled account of exactly how their engagement with LW connects to their broader goals? How many have even noticed that this post contains zero advice?
Well I’m glad you liked my essay, because I didn’t. Heh. I think largely because I got the sense that we might not be disagreeing as much as it seemed, which would make coming at it from a critical angle misguided in a way that makes me cringe. But even after working to square your words with my perspective I didn’t see a solution so fuck it. Send it. I’m definitely interested in reading a response to my essay, if you feel like writing one up.
Also, the idea of running a contest like this is pretty cool. It’s definitely got me thinking..
Yeah, I see what you’re seeing. Someone without anxiety to flinch away from would be behaving very differently, and the tells are unambiguous. To start with, they either wouldn’t post about it at all, or they’d make a post that couldn’t be accurately mocked as “Oh, you’re afraid of impending doom? Which stocks have you shorted?”.
If you aren’t flinching from anxiety, you can see a turn or two ahead and address the glaring weaknesses before they’re used against you. If you’re as willing to put your money where your mouth is as you think the other side isn’t, then you can bet on that. Not “How do you explain banana? Checkmate atheists!” before atheists get a chance to explain banana, stick your neck out and predict crickets.
Show me how to collect on my shorts when the world ends. Better yet, show me how to collect on my shorts when the world doesn’t end because I took the dangerous trajectory seriously, and helped to avert disaster. Or at least don’t open with “doom” in the first five words and close on “what are you shorting” so any real point doesn’t get buried in nonsense.
If he were to make it obvious how to actually profit, so that the doomers can actually recognize “Huh, that would work”, then it gets interesting if doomers don’t place the bets. At which point his implicit accusation actually lands, because he would have tested it. And if bets get made, he gets a chance to learn.
Yet he chose not to do the thing that would give either “makes his point land” or “learn he’s wrong”, and chose to do the thing that changes no minds, gets him mocked in the comments, and allows him to not engage any more deeply.
Revealed preferences.
Stronger, he could actually be correct that most attempts to analyse such risks via rationalist style arguments will be distorted by fear.
This is absolutely to be expected. It’s not exactly easy to avoid, given how much is at stake here. “Lol, you scared bro?”. Yes.
The unfortunate thing is that these flinches are so ubiquitous that noticing them is like a fish noticing water. Where have we seen someone not flinch? Seen someone smiling in the face of likely doom without inaction? It’s rare, especially where things are scariest.
It’s actually plausible that maintaining such a strong epistemic bar for such fears makes sense for Tyler personally[..]However, neither of these means that his decision to run from anxiety is necessarily correct. And insofar as this is choice is unconscious, he won’t know what trade-off he’s making. This is rarely a good place to find yourself.
I think it’s easy for people to take this as lip service, but it’s seriously important.
The temptation here is to think “Ah, Cowen is irrational, it’s important to be rational, I should face reality”, but that is itself a flinch from reality any time you don’t know you’re prepared to handle reality. Which turns out to be every time you’re saying “I should face reality” because if it were unambiguously clear[1] to you that facing reality would be a winning move for you, you’d have done it already.
And that’s why despite mocking Cowen a bit above, I actually sympathize quite a lot. I don’t think “pretending not to be avoiding” is a winning move for him there (lol), but the avoiding itself is most likely completely reasonable, and I mean that sincerely.
I definitely choose not to engage with scary things, at times. If I try to force myself to engage too far I can feel the distortionary pressures, and my own inability to handle them all indefinitely without loss. Since it’s gonna happen either way, I’d rather it be on my terms, in my accounting book, while I work on the capacity to handle more.
I don’t think there’s any other option, really.
- ^
Meaning no part of you, which you take seriously enough to act on, is still predicting you’d lose by facing it
- ^
On the dodo bird verdict, I’m reminded of this old comic about what airplanes would look like if different engineering specialties got their way.
If you remove the powerplant, it won’t fly. If it fails under stress, it won’t fly. Too heavy, won’t fly. Etc. If you take any working airplane, the powerplant guy will be able to say “see, it uses a powerplant and that’s why it works”, and he will have a point. They all will. The pieces need to all be there or it won’t work.
Similarly, if you ask “please pass the milk”, a hypnotist will tell you that “pass the milk” is an embedded command embedded in “can you please”, which is why it works. The NVC guy will point out that the “please” makes it a request not a threat of violence. The rapport focused guy will correctly point out that if they’re cold and hostile to you they’re likely to refuse. The hygiene guy will note that if you stink too much no one will show up to eat with you in the first place.
You can compensate for a little extra weight with a little extra powerplant, but if you forget the wings entirely you’re SOL. Often what matters most isn’t which part of the airplane the engineer talks about as “the way planes fly” but whether they pay enough attention to the rest that there aren’t any gaping holes. You can let the powerplant guy design the airplane and it’ll actually work, because he will still include wings.
When it comes to the question of whether therapeutic framework matters, it matters whether we’re in a context where we need the absolute best in one specific area, or whether we just need someone who remembers the wing even though that’s not his specialty. In the latter cases you’re going to get the dodo bird verdict because we’re measuring the effects of holes not specialties and holes probably don’t correlate much with specialty. If you’re working on a human powered airplane, you better hire the weight guy.
I am confused at the idea that you perform some mental motion and then stop being in physical pain.
Yeah, it’s definitely mind bending at first. The first time I experienced it, it was because I decided to try a technique that in theory should work but realistically speaking there’s no way it could, right?
And then at the end of the technique I was very confused about what I was experiencing, because it simultaneously felt like nothing changed and yet it didn’t bother me anymore, and I couldn’t even figure out when it had changed.the sensation of pain is clearly unrelated to whether I want to not feel this right now, as any masochist could tell you?”
Masochism is pretty different in that pain still feels meaningful to the masochist. What I’m talking about the same sensations, but with the urgency of the sensations of your tongue sitting in your mouth. It’s there, you can notice them when you care to, but they’re utterly irrelevant so you stop noticing them the moment you have anything else to do.
I agree there is the involuntary redirecting of attention, but aside from that I still don’t consider it cognition distorting.
Direction of attention is literally the whole game. It’s what determines what we learn and what we do. Cognition distortions are built out of unmodeled pulls on our attention.
I cannot think of times that I was outside and thought “I need to go inside soon, this is damagingly cold!” or was physically exerting myself and thought “I need to slow down or stop unless I’m okay with being sore later”. These seem like pretty weird thoughts
He did open by talking about how this isn’t the default :)
But yeah, these are absolutely thoughts one can have, and having them results in the things no longer feeling “bad”. No longer feeling like it “hurts”.
And then once it no longer feels bad, you no longer feel motivated to get away from it. Your cognition is no longer distorted towards “I should go inside and get warm” or else just impaired from having to try to not be distracted by the distracting badness.
Reciting the words to oneself isn’t in itself enough, obviously. The effects come from looking past the feeling itself to the thing it’s pointing at, and genuinely deciding what you want to do about that. Which can be tough if you try to minimize the harms when you do. A large part of no longer feeling “bad” is no longer needing to feel bad in order to stop doing mildly harmful things.
Once you get that, it really does feel like “Nah, I don’t even want a jacket right now, so long as we’re going to get back before my body starts to shut down”—and it doesn’t start feeling “bad” again until you can no longer use your fingers because you’re staying out longer than expected. Or “I could definitely stretch further, and they’re just sensations, but I genuinely don’t know at what point my muscles will take too much damage, so I’m gonna call it here”—and then update on how your muscles felt after, for next time.
You’re noticing something really important and really interesting :)
And I think there’s something stronger, too—in an active inference framework, beliefs and desires are both just expectations about the world. Experientially, this rings true to me—the feelings of frustration at not getting what I want and at being taken off guard by something I hadn’t even been paying attention to are very similar. It seems deeply hard to distinguish between what we want and what we believe.
Yep :)
There’s a reason “I expect you to ____” plays both roles.
This is a bit more speculative, but sometimes I think people don’t fully absorb this point: It’s not psychological, it’s neurological.
Oh no, it’s deeper than that.
I figured this stuff out in part by trying to design an optimal temperature controller and noticing that it necessarily applies to anything that tries to do anything.
Friston takes this even further, and point out that everything is trying to do something. Because even drops of oil have to resist entropic forces otherwise they’d disperse into the surrounding water and cease to be distinguishable as a thing—and that therefore, it is a theory of “every thing” (not 100% sure this is one of the videos where he makes the pun).
So we may not be able to fully adjust for it by only manipulating psychological factors, e.g. consciously trying to be more objective or less selfish.
In fact, the harder we try to be more objective, the more we import these same distortions at the meta level. “I’m not irrational dammit! You are!”
There are solutions, and you’re right that this is not one.
But I have the sense that this simple consideration is underrated, and I hope this post can provide a reference point for it and make people take it into account in their personal deliberations.
The tricky part is that taking it into account cuts against the ability to control, which is the exact thing we don’t want to give up. So the people who have the distortions in their way are going to distort away from what you’re saying the most.
In the past, I’ve felt a sense of being overwhelmed at all these considerations, and felt tempted to just avoid thinking about them—but that can’t be the answer. We have to take the uncertainty seriously.
We’re more or less not taking the uncertainty seriously, which is proof positive that we don’t have to.
Notice what comes up when we look at this though. “But there will be consequences!”. Yep.
Declaring what we “have to” do functions as another way to get us out of that uncomfortable uncertainty, huh?
You might like Valentine’s video No need for “should”
But I don’t know whether you’re trying to change cultural or individual rationality here:
Like, I’m mostly optimistic about getting a few individuals to not do Crappy Epistemics, whereas I feel like you’re targeting groups,
If it’s from the criticisms of the rationality community for not berating people into rationality, oops. I should have put those in quotation marks. The point was “you’ll hear this from someone in this mode”, not to assert those things myself. I mostly see those exhortations as examples of the thing they’re railing against.
I’m not really trying to change anything here, just describing what is actually going on.
which seems difficult if I’m one of very few people who get what you’re saying.
Oh, for sure :)
But to be clear, I’m excited that you seem to be picking up on the same thing from another angle, separate from anything I’ve said.
I’ve gotten pretty comfortable with my ability to help people see the things needed to realign themselves on some object level issue, but I’m still new at communicating the meta thing of how to help people see how this alignment process works. Difficult, but fun/interesting, and I don’t think too difficult to learn.
If you’re saying “we need dramatically better instrumental rationality, of which short-term optimization targets are a big component” then yes, strong agree. I feel like you’re saying something else though, maybe about coordination between humans?
Yeah, I think that’s a facet of it. But also, yeah, it’s more than that. The stuff about coordinating between humans is just another facet too.
There’s a failure mode in changework where people try to “fix their irrational fear”, or “walk the client through what they need to do to fix their irrational fear”. These seem like perfectly reasonable responses which is why people do them, but watch it play out enough times and you start to notice that the stubborn attempt to control away the fear is the problem. That once you notice that you don’t actually know your fear to be irrational, you naturally turn towards noticing whether you’re actually safe, and that’s the move that conditions away inappropriate fears. That once you notice that trying to push people towards having a certain “correct” orientation to their pain is actually the thing that causes the suffering, you naturally turn towards “what’s the real problem here?” and that’s what dissolves the suffering in a word or “a few messages”.
This move of “Oh, control isn’t working. I wonder why?” turns out to be very general. Not just applied to one’s own mind, or helping others with their own minds, but also helping others learn to help others with their own minds and so on and so on. When my friend asked for help getting her four year old to take her eyedrops, I was able to play with my friends discomfort which led to her being able to play with her daughters discomfort, which led to her daughter playing with her own discomfort. “You’re a brown belt in jiu jitsu, what do you mean how do I get my child to take her eyedrops!?” → “Sigh. I just feel like I shouldn’t have to use force… and I guess this is another one of those things where my own tension is telling her to be tense, huh?” → “Mommy, can we play the eyedrop game!?”
Valentine has a good post “We are already in AI takeoff”, which I took a stab at putting into my own words in a comment there.
In short, and translating into the language we’re using here, “trying to align AGI” is itself an instance of this same exact failure to align ourselves. Because no one is looking at it like “Oh, yeah, easy peasy. I predict that I will not experience prediction error, because I got this”. It’s all pushing back in attempt to control away prediction error because the consequences of failing are unimaginably bad, while failing to act on the uneasiness coming from predictions of this not panning out. Which turns out to be where all the most useful information is.
When I look instead towards “If this goes well, what does that look like? How did we get here?”, the answer I see is one where the people guiding the development of AI aren’t pushing away from any of the relevant information, let alone the information that they themselves perceive as most important with respect to whether what they’re doing is working and what kind of moves might actually work.
In other words, it’s one where the researchers themselves have enough embodied skill in alignment that they can approach the problem with their full faculties. Not just because “That’s what rationality is, and good instrumental rationality is necessary for succeeding at hard things”, but also “it’s literally the same skill”. In the same way that relating to one’s own mind is the same skill as helping a friend relate to theirs is the same as helping a friend helping their kid relate to theirs.
Same skill, applied on multiple levels. The skill in “becoming rational”/”coordinating groups of humans”/”aligning AI” is all skill in alignment. Borrowing Val’s words again, “It’s fractal”.
When I think about “how to align AI”, I notice that I don’t actually know how to do this. There’s nothing I can see, where I think “Ah, this is the code I need to write” or “here’s the things I need to exhort at people to do” that will predictably yield the outcomes I want. Not through “targeting groups”, or “targeting individuals”, or “targeting code”.
And I notice that stopping to notice this is by far the most important thing I can do, since “trying” would necessarily blind and therefore doomed to success by luck at best. And that one of the better “object level applications” for me right now is to highlight the nature of this move, since a big part of “Why is this control not working?” is “Because people aren’t aware of how control works”—Okay, cool. So if we change that, this part of the problem dissolves.
There’s something even more general and self referential that I’m fumbling towards though, since doing that thing is the thing I can actually expect to lead to the best possible outcomes—structurally and necessarily. But it’s a bit mindbending because “trying to generalize” is itself an instance of the thing I’d be trying to avoid (and so would “trying to not try” or concluding “we shouldn’t try”). “Generalizing is hard. I wonder why?” is the generalization. So I guess that’s the next thing to wonder, once I have some mental room for it.
Anyway, your post on BCI facilitated AI alignment looks to me like a step in the same direction. A step towards noticing that AI alignment is downstream of human alignment (in this case, because aligned and augmented humans are more competent which is instrumentally useful), and that the solutions which actually work have more competent humans more tightly integrated in the alignment process for longer—rather than keeping a stance of “I’m outside the system, aligning THAT THING is what I’m trying to do, dammit”.
I don’t think you’ve been explicitly thinking about it in the same terms I’m laying out here, but it does seem like it might be downstream of beginning to sense and act on the same thing I’m fumbling towards. Like you might be already on the same path fumbling towards the same thing that I’m trying to put a finger on (and noticing myself not fully having yet, in this sentence. Lol). Does this fit?
You’ve probably already mentioned it somewhere, but perceptual control theory relatedly posits that motivations/actions are just a way to control what sorts of things we experience.
I don’t think I have, but yes. Agreed.
If the same machinery underlies factual prediction and normative actions, we confuse them to all heck.
Yep! And there are pretty big practical consequences of this.
This is a clearer, much more precise statement of a somewhat different mechanism than I was originally proposing here. I’ll need to think for a bit about whether this changes my BCI-superintelligence stuff.
Hm. I haven’t had time to read and process your other post yet, but I do think that human alignment is important for having a hope at aligning things bigger and more intelligent/powerful than ourselves. Like, there’s a big “I’m outside the system!” type error, which systematically screws up control attempts because they don’t take into account the inside-the-systemness and attempt to align “them” instead of “us, starting with me”—two boxing AI alignment, basically. It sounds like maybe you’re on a similar track?
But even in situations like this, there’s probably confusion somewhere. Like, “why’s a fellow rationalist confidently Wrong?” is probably bubbling from some part of their mind, even when other stuff is talking over the confusion.
The problem is that this is in direct opposition to the attempt to control. Ask a thermostat why the room is too cold, and the only answer it has is “Because I haven’t added enough FIRE!”. Why is the rationalist confidently wrong? Because he’s a Bad rationalist! Why is he a bad rationalist? Because y’all haven’t called him out for his Badness! More shame! Beat him into shape! Why haven’t we done that? Because y’all are bad rationalists too! That’s why I’m yelling at y’all to fix you!
Pondering “Hmm… I dunno. Maybe he doesn’t see this piece?” requires people to relinquish the control which they’ve already decided is worth doing, so you’re gonna get “control” type answers unless you push back against their control loop hard enough that they let go. I’m not saying you can’t do it, but it’s gonna take some oomph which has to be sourced from somewhere—which makes it trickier to self apply. The question “is it working?” comes from outside the loop and points at the loop itself, which makes it a lot more widely useful and easier to self apply.
I do think we’re mostly in agreement. If I read your post as a blurry gesture at a shape of things, I think you’re getting the picture exactly right which is why I’m so excited to see another person “getting it”. So, “Yes! Strong upvote!”.
If I read your post as speaking technically and precisely, I see details that will need to be changed in order to get things to actually work in practice.I “want” both of us to get to the truth; but I’m locally intuitively trying to get them to change their mind.
I’m saying that this is directionally correct, but that the same problem shows up with respect to “changing people’s minds”.
“I “want” to get them to change their mind (because that’s what gets both of us to the truth which I already have); but I’m locally intuitively trying to push away from experiencing wrongness”What we’re actually optimizing for is something more specific and even more out of alignment. If we were actually optimizing for changing people’s minds, things would play out very differently because it would lead right back to truth seeking.
“What am I confused about” slowly but predictably clarifies where reality is biting your models, so long as you’re actually optimizing for understanding your confusion.
The mental state which we label “confusion” when we notice it in ourselves, is the state of feeling disoriented. If we were to put into words the implicit stance of this state, it’s “I have not been able to properly orient to this situation. I don’t know what to make of it and I’ve noticed this”. Without noticing that your models are failing you, it just feels like the world is wrong. “The problem is that the peg won’t go through the hole!”, not “I can’t figure out which hole this peg goes in”, let alone “Lol, square peg don’t go in round hole”.
In order to enter this state we label confusion, we have to notice that the wrongness we perceive is in our map, not the territory. We have to notice “Ah, I feel wrongness, that means I’m not oriented properly. I don’t know what to make of this”.
When you ask a Rationalist why they’re frustrated in a dialog that isn’t going their way on LW, they’ll tell you they’re frustrated because that other guy is wrong on the internet—and they’re supposed to be rational, dammit. Even in well respected (and otherwise respectable) rationalists, there is often-to-usually no recognition that the feeling of wrongness that they are modeling as “in the other guy” is actually in their own models, necessarily. This is where we go wrong.
And asking “What am I confused about?” won’t actually help. Because there’s no such such thing to notice. An outsider may describe the struggling person as “confused” or “disoriented”, but from the inside they have no feeling of disorientation to notice. In their own model, they are oriented properly—it’s the other guy who isn’t! So far as they’re concerned the problem is the hole, not the fact that they’re trying to shove a square peg into a round hole.
The thing which is actually there to notice is our perceptions of “wrongness”. If we ask “Where am I experiencing wrongness” then we’ll find it immediately. I experience it when the peg won’t go into the hole. When the guy online doesn’t change his mind.
This is the thing that’s actually there to notice, which we can then use to take the next step of “Why is the peg not going into the hole, do you think?”. Why is that guy wrong on the internet? This wrongness I’m feeling… what’s that about?
And this is the step that updates models to better match reality. Including noticing the inner misalignment that has been screwing everything up.
It’s true that sometimes we have tiny little notes of discord. Tiny little hints of “maybe I’m confused here?” that we fail to give due weight. But by the time there’s even a hint of subjective confusion to notice, we’ve already done the hard part. Most of our failures are due to stubbornly externalizing wrongness because that’s how we try to control, and we don’t want to give that up until we see a better way.
I’m confused what assumptions you’re making here. Are you taking “self deceiving” to be binary?
I didn’t say “if she’s self deceiving [at all]”. I put the bar at too self deceiving to be worth talking to. If she turned out to be able to hold a worthwhile conversation after I declared her to be self deceiving, it’d just mean that I was wrong that the level of self deception surpassed the “not worth talking to” bar. In other words, if I thought she was a little self deceiving but not to the point where it’s worth burning that bridge, I probably wouldn’t make such declarations. I might raise the possibility, but I probably wouldn’t declare it.
Is that clearer?