Indeed, at a second look, it continues to seem like an instance of the pattern, with the attendant mild harms.
For example, Richard’s post is about how the whole “AI alignment” community did not maintain its pursuit of the core problems of understanding around AI alignment. Quoting the first paragraph:
This sequence is about the last decade in AI alignment. Over five posts, it recounts the gradual transition from a field which treated alignment as a hard scientific problem, to a field which has largely abandoned the goal of deep, generalizable scientific progress in favor of iteratively improving existing systems and attempting to gain technological and political power. I also describe (in subsequent posts, which I’ll upload over the next few weeks) how fear and (self-)deceptive reasoning made the field one of the biggest forces pushing AI capabilities forward over the last decade, especially via significant contributions to the scaling of LLMs and the development of ChatGPT.
As a subpoint, he discusses the focus of funders such as Holden on prestige, and how that focus failed to allocate funds to MIRI, which was relatively much more generative and also focused on the difficulty of alignment. Then you write:
You talk about Holden being misled about AGI for 8 years, but don’t mention MIRI planning to build recursively improving Friendly AI with a small team and potentially just 1 philosopher, for a comparable amount of time. If OpenPhil had funded MIRI more in 2016, it would have been funding them to attempt this!
So it sounds like maybe you’re saying:
Even if MIRI had gotten more money, that would have been strategically bad because MIRI’s plan was bad.
There are counterpoints (e.g. Richard brought up the graceful failing of MIRI’s plan, which was related to MIRI understanding much better about alignment difficulty). But more saliently to our discussion of hobby horses, your point is not obviously relevant to Richard’s post! For example, is it meant as a rebuttal of Richard’s point about Holden? But your reasons for MIRI’s research being bad seem disjoint from Holden’s reasons, and Richard was describing Holden’s reasons, so his point would stand. Or are you saying that MIRI being bad at strategy means that Richard’s post even truer than he presents it, because actually ALL the research was failing to have a sufficiently “deep understanding of the issue”? Or are you saying that “the goal of deep, generalizable scientific progress” that is the topic of Richard’s post is actually a BAD goal? These are plausible-ish, but IMO it would be more helpful to say what you’re actually trying to say and then make a case for it.
My point here isn’t to discuss the object level, but rather to give you some indication of why your comments indeed seem like what’s actually driving them is simply that you want to talk about humans / MIRI / Yudkowsky being bad at strategy, and you are kinda shoehorning that in to discussions without them being relevant or clearly-to-other-people relevant and without explaining the relevance.
Maybe you could elaborate on the harm you see with the meme spreading? (I don’t feel well-equipped to think about harms of memes spreading; but I do feel that I don’t have a fully clear theory of what a charging hobby horse is and how it’s bad and when it’s good etc., and unclear memes seem potentially bad. For example, I wouldn’t want social suppression in general of discussion about a topic someone cares about a lot and makes a lot of comments about, or similar.)
I basically endorse Tsvi’s comment above. It does also seem reasonable that Wei is concerned that this post will continue to be used to criticize him (though this doesn’t seem like a good reason to take down the post). In hindsight, I regret the parenthetical where I added “(a concept which was originally inspired by Wei making a similar comment on a post by Tsvi)”, because I don’t want to associate the concept too strongly with Wei—I think it’s a generically useful concept.
I find the situation saddening, and I would be very sad to see Wei decrease his participation in LW.
When you banned Wei, I remember thinking that this was an overreaction.
I think that Wei’s current concerns regarding the impact of [the post and the meme it’s spreading] on how LW views him or his ideas are an overreaction.
I.e., I don’t think there’s much risk of this turning out very bad for Wei regarding how people on LW perceive him or his ideas or something?
But: Things like that are often said in an attempt to gaslight, which I don’t want to do or look like I’m doing. So take it as an observational report rather than judgment (if that makes sense).
Something like: I believe what I’m saying is true, but I’m not trying to use any normative modality to say that people “should” believe me. (I would like to make a locutionary analogy of the distinction between preferences and bids, but I don’t think there are (even locally) established words for this distinction (“report” vs “claim”?).)
The entire thing reminds me a bit of an emotional/discursive amplifying echo chamber, where Alice behaves in a way that causes Bob to react somewhat emotionally to it and then Alice reacts even emotionally’er to Bob’s reaction, and the cycle continues for a bit, out of proportion to Alice’s original behavior, which could be avoided if any party at some early point unilaterally decided to abandon this emotional thread.
I liked the charge of the hobby horse post when it came out and considered it valuable.
I internalized the concept it introduced, but rarely use it in conversations.
I don’t think that’s that much evidence regarding Wei’s concern that the concept is already out of the box, because I should be expected to be an early concept adopter in cases like this one.
I think it’s possible but unlikely that the concept will turn out to be a non-trivial sort of infohazard … conceptual hazard? That is, that it will skew people’s thinking, giving them a handy label to dismiss people’s repeated critiques or whatever as “charging their hobby horses”.
It generally seems to me that, on the one hand, one would wish that moving meta in a discussion would be helpful/healthy. E.g., if you notice some trend in your interlocutor’s participation in conversations that is relevant for how you and other people orient to this person’s participation in conversations, then identifying the trend should be helpful, either by giving you better heuristics for engagement or by discussing the thing with your interlocutor on a higher level, which then helps you resolve things, or at least for modeling them (and general social/psychological dynamics) better. But what often seems to happen is that people take those sorts of moves as attacks on some deep aspect of their person or something like that (?).
E.g., I was once in a 3-person conversation involving an increasingly heated factual disagreement between the other two people. At some point, Alice told Bob that [they’re gonna use Bob’s reaction to what Alice was saying as an example of the dynamic they were having a heated debate about]. This made Bob’s attitude toward the conversation even worse.
It seems to me that a big issue with the discourse around this concept is that, like one man’s modus ponens is another man’s modus tollens, also one man’s charge of the hobby horse is another man’s valuable, if sometimes annoying, reminder that there’s this incredibly important and systematically neglected issue related to things being talked about here. (Maybe it’s just a convoluted way of saying that the judgments of whether something should be counted as the charge of the hobby horse are very observer-dependent and the disagreements are often rooted in some deep-ish beliefs/values/convictions.)
Many of the caveats you end up with in this post’s conclusion section make the boundary of the concept pretty vague and observer-dependent. As you say, “there are a lot of comments that look kinda similar to a Charge of the Hobby Horse, but that are good”.
E.g., take your and Wei’s examples from this very post. To me, both look like reasonable and thoughtful people thinking that ~[AI risk] is ~[the biggest thing ever] and also that people working on this issue seem to be systematically blind/oblivious to a very important thing X. It makes sense for such people to bring up X often and widely in tangentially but significantly (in their judgment) related discursive contexts, regardless of whether some people on the other side of some sort of worldview spectrum judge this as them charging their hoppy horses (unless the latter outweighs the benefits of the former).
I generally don’t recall people on LW being visibly annoyed about you talking about your thing X, and the only two people on LW I noticed being annoyed about Wei’s thing X have been you and Richard. (But I’m probably significantly below average when it comes to noticing this, so it should be taken with a grain of salt.)
I would want to say that there are clearer examples, e.g., people bringing up racism, social justice, the Omnicause, etc., into nearly every single conversation. But, similarly, it looks reasonable from their perspective, and I can disagree with the perspective that makes this behavior reasonable to them, because it is the perspective I find unreasonable.
Or take anti-abortion activists. Pro-choice people would condemn them for their radicalism/extremism or whatever, but, in a sense, what are they supposed to do if they believe that those people are literally killing little babies?
There is a partly related phenomenon of people adopting a particular way to view certain things, which makes them unable to adopt/understand alternative ways of viewing those things (and is possibly the case in many of the cases from the last bullet point). But one can do things that look to some like the charge of the hobby horse, without having this issue.
Maybe the judgment of the [charge of the hobby horse]-adjacent phenomena involves/[should involve] a spectrum of (at least): (1) reasonableness of the position guiding the charge; (2) the relevance of the topic of the charge to the topic of the discussion being charged into, even given the stated beliefs of the charger; (3) some sort of vague psychological virtue stuff, like obnoxiousness, smugness, apparent wanting to feel important / grab the center of attention, etc.[1]
Low-confidence-speculating, possibly a reasonable way out is some sort of [veil of ignorange-y]/[Schelling fence-y] rule like: We allow those sorts of [charge of the hobby horse]-adjacent behaviors but no more, because if we allow more, then all hell might break loose.
I would want to say that there are clearer examples, e.g., people bringing up racism, social justice, the Omnicause, etc., into nearly every single conversation. But, similarly, it looks reasonable from their perspective, and I can disagree with the perspective that makes this behavior reasonable to them, because it is the perspective I find unreasonable.
But there are still differences of behavior. You can do this rudely or not rudely. E.g. if you’re at a talk and go up in the Q&A at the end, you could ask a pointed question about it (fine! good!) or you could pretend to ask a question but actually just ramble on your thing until the moderator takes the mic from you. I believe the comment thread on my post about deference was rude because my interlocutor was ignoring what I was saying, because he just wanted to talk about his thing. This burns the commons of “author putting forward a new idea upholds the principle of responding to all serious criticism”, by piggybacking off that.
I don’t think I mind harshness, though maybe I’m wrong. E.g. your response to me here https://www.lesswrong.com/posts/zmtqmwetKH4nrxXcE/which-side-of-the-ai-safety-community-are-you-in?commentId=hjvF8kTQeJnjirXo3 seems to me comparably harsh, and I probably disagree a bunch with it, but it seems contentful and helpful, and thus socially positive/cooperative, etc. I think my issue with this thread is that it seems to me you’re aggressively missing the point / not trying to get the point, or something, idk. Or just talking about something really off-topic even if superficially on-topic in a way I don’t want to engage with. IDK.
Like, to me it seems pretty rude for Bob to
not spend basic effort to understand what Alice is saying
go off on his thing
while posing as being responsive to what Alice is saying
while posing as being responsive to what Alice is saying
I think the problem with this isn’t rudeness as such, but with deceiving people (who aren’t paying close enough attention) that Bob’s comment was relevant to Alice’s thesis. (A lot of politeness norms are about concealing or obfuscating information, and I often want to rudely defy those, but I don’t want to deceive people; the category of rudeness is lumping together different things that I want to treat very differently.)
Ok. I could quibble with “to a large extent” (actually, I took that to mean at least a majority, but now I’m not sure if that’s what you mean by that). But I think I agree that a lot of politeness norms are about obfuscating.
...But there’s also plenty of politeness that’s not about that. For example:
keeping out of personal space
coughing without covering your mouth
not being too loud
ignoring what someone is saying when you’re talking to them
interrupting
hijacking attentional spaces
not doing some amount of interpretive labor
dismissing ideas / possibilities without due explanation or consideration
Some of these shouldn’t be universally enforced; e.g. in many cases I can quite enjoy a perfectly friendly conversation with lots of kinda-heated-sounding interruptions. Some of these are vague (e.g. what is “dismissing” and “due consideration”). Some of these are clear aggression (e.g. personal space) but some of them aren’t.
I don’t know exactly what “rudeness” is as a category, or know how to perfectly delineate each type of rudeness so that it’s already explicitly never in conflict with truth-seeking, but I don’t think I should have to know in order for us to agree that rudeness in general is bad, and to enforce rudeness norms.
Indeed, at a second look, it continues to seem like an instance of the pattern, with the attendant mild harms.
For example, Richard’s post is about how the whole “AI alignment” community did not maintain its pursuit of the core problems of understanding around AI alignment. Quoting the first paragraph:
As a subpoint, he discusses the focus of funders such as Holden on prestige, and how that focus failed to allocate funds to MIRI, which was relatively much more generative and also focused on the difficulty of alignment. Then you write:
So it sounds like maybe you’re saying:
There are counterpoints (e.g. Richard brought up the graceful failing of MIRI’s plan, which was related to MIRI understanding much better about alignment difficulty). But more saliently to our discussion of hobby horses, your point is not obviously relevant to Richard’s post! For example, is it meant as a rebuttal of Richard’s point about Holden? But your reasons for MIRI’s research being bad seem disjoint from Holden’s reasons, and Richard was describing Holden’s reasons, so his point would stand. Or are you saying that MIRI being bad at strategy means that Richard’s post even truer than he presents it, because actually ALL the research was failing to have a sufficiently “deep understanding of the issue”? Or are you saying that “the goal of deep, generalizable scientific progress” that is the topic of Richard’s post is actually a BAD goal? These are plausible-ish, but IMO it would be more helpful to say what you’re actually trying to say and then make a case for it.
My point here isn’t to discuss the object level, but rather to give you some indication of why your comments indeed seem like what’s actually driving them is simply that you want to talk about humans / MIRI / Yudkowsky being bad at strategy, and you are kinda shoehorning that in to discussions without them being relevant or clearly-to-other-people relevant and without explaining the relevance.
Maybe you could elaborate on the harm you see with the meme spreading? (I don’t feel well-equipped to think about harms of memes spreading; but I do feel that I don’t have a fully clear theory of what a charging hobby horse is and how it’s bad and when it’s good etc., and unclear memes seem potentially bad. For example, I wouldn’t want social suppression in general of discussion about a topic someone cares about a lot and makes a lot of comments about, or similar.)
I basically endorse Tsvi’s comment above. It does also seem reasonable that Wei is concerned that this post will continue to be used to criticize him (though this doesn’t seem like a good reason to take down the post). In hindsight, I regret the parenthetical where I added “(a concept which was originally inspired by Wei making a similar comment on a post by Tsvi)”, because I don’t want to associate the concept too strongly with Wei—I think it’s a generically useful concept.
Maybe @Raemon or someone else would be able to give a more neutral perspective?
FWIW
I find the situation saddening, and I would be very sad to see Wei decrease his participation in LW.
When you banned Wei, I remember thinking that this was an overreaction.
I think that Wei’s current concerns regarding the impact of [the post and the meme it’s spreading] on how LW views him or his ideas are an overreaction.
I.e., I don’t think there’s much risk of this turning out very bad for Wei regarding how people on LW perceive him or his ideas or something?
But: Things like that are often said in an attempt to gaslight, which I don’t want to do or look like I’m doing. So take it as an observational report rather than judgment (if that makes sense).
Something like: I believe what I’m saying is true, but I’m not trying to use any normative modality to say that people “should” believe me. (I would like to make a locutionary analogy of the distinction between preferences and bids, but I don’t think there are (even locally) established words for this distinction (“report” vs “claim”?).)
The entire thing reminds me a bit of an emotional/discursive amplifying echo chamber, where Alice behaves in a way that causes Bob to react somewhat emotionally to it and then Alice reacts even emotionally’er to Bob’s reaction, and the cycle continues for a bit, out of proportion to Alice’s original behavior, which could be avoided if any party at some early point unilaterally decided to abandon this emotional thread.
I liked the charge of the hobby horse post when it came out and considered it valuable.
I internalized the concept it introduced, but rarely use it in conversations.
I don’t think that’s that much evidence regarding Wei’s concern that the concept is already out of the box, because I should be expected to be an early concept adopter in cases like this one.
I think it’s possible but unlikely that the concept will turn out to be a non-trivial sort of infohazard … conceptual hazard? That is, that it will skew people’s thinking, giving them a handy label to dismiss people’s repeated critiques or whatever as “charging their hobby horses”.
It generally seems to me that, on the one hand, one would wish that moving meta in a discussion would be helpful/healthy. E.g., if you notice some trend in your interlocutor’s participation in conversations that is relevant for how you and other people orient to this person’s participation in conversations, then identifying the trend should be helpful, either by giving you better heuristics for engagement or by discussing the thing with your interlocutor on a higher level, which then helps you resolve things, or at least for modeling them (and general social/psychological dynamics) better. But what often seems to happen is that people take those sorts of moves as attacks on some deep aspect of their person or something like that (?).
E.g., I was once in a 3-person conversation involving an increasingly heated factual disagreement between the other two people. At some point, Alice told Bob that [they’re gonna use Bob’s reaction to what Alice was saying as an example of the dynamic they were having a heated debate about]. This made Bob’s attitude toward the conversation even worse.
It seems to me that a big issue with the discourse around this concept is that, like one man’s modus ponens is another man’s modus tollens, also one man’s charge of the hobby horse is another man’s valuable, if sometimes annoying, reminder that there’s this incredibly important and systematically neglected issue related to things being talked about here. (Maybe it’s just a convoluted way of saying that the judgments of whether something should be counted as the charge of the hobby horse are very observer-dependent and the disagreements are often rooted in some deep-ish beliefs/values/convictions.)
Many of the caveats you end up with in this post’s conclusion section make the boundary of the concept pretty vague and observer-dependent. As you say, “there are a lot of comments that look kinda similar to a Charge of the Hobby Horse, but that are good”.
E.g., take your and Wei’s examples from this very post. To me, both look like reasonable and thoughtful people thinking that ~[AI risk] is ~[the biggest thing ever] and also that people working on this issue seem to be systematically blind/oblivious to a very important thing X. It makes sense for such people to bring up X often and widely in tangentially but significantly (in their judgment) related discursive contexts, regardless of whether some people on the other side of some sort of worldview spectrum judge this as them charging their hoppy horses (unless the latter outweighs the benefits of the former).
I generally don’t recall people on LW being visibly annoyed about you talking about your thing X, and the only two people on LW I noticed being annoyed about Wei’s thing X have been you and Richard. (But I’m probably significantly below average when it comes to noticing this, so it should be taken with a grain of salt.)
I would want to say that there are clearer examples, e.g., people bringing up racism, social justice, the Omnicause, etc., into nearly every single conversation. But, similarly, it looks reasonable from their perspective, and I can disagree with the perspective that makes this behavior reasonable to them, because it is the perspective I find unreasonable.
Or take anti-abortion activists. Pro-choice people would condemn them for their radicalism/extremism or whatever, but, in a sense, what are they supposed to do if they believe that those people are literally killing little babies?
There is a partly related phenomenon of people adopting a particular way to view certain things, which makes them unable to adopt/understand alternative ways of viewing those things (and is possibly the case in many of the cases from the last bullet point). But one can do things that look to some like the charge of the hobby horse, without having this issue.
Maybe the judgment of the [charge of the hobby horse]-adjacent phenomena involves/[should involve] a spectrum of (at least): (1) reasonableness of the position guiding the charge; (2) the relevance of the topic of the charge to the topic of the discussion being charged into, even given the stated beliefs of the charger; (3) some sort of vague psychological virtue stuff, like obnoxiousness, smugness, apparent wanting to feel important / grab the center of attention, etc.[1]
Low-confidence-speculating, possibly a reasonable way out is some sort of [veil of ignorange-y]/[Schelling fence-y] rule like: We allow those sorts of [charge of the hobby horse]-adjacent behaviors but no more, because if we allow more, then all hell might break loose.
In case this needs to be said, I don’t think any of the three examples discussed in this post involve a significant amount of (3).
But there are still differences of behavior. You can do this rudely or not rudely. E.g. if you’re at a talk and go up in the Q&A at the end, you could ask a pointed question about it (fine! good!) or you could pretend to ask a question but actually just ramble on your thing until the moderator takes the mic from you. I believe the comment thread on my post about deference was rude because my interlocutor was ignoring what I was saying, because he just wanted to talk about his thing. This burns the commons of “author putting forward a new idea upholds the principle of responding to all serious criticism”, by piggybacking off that.
I guess some people don’t want to hear or believe that I just don’t like the rudeness, and completely sublimates this to their conflict that they are engaged in. Ok fine. I’ll note that the one other time (unless I’m forgetting) I considered banning user was here: https://www.lesswrong.com/posts/DDG2Tf2sqc8rTWRk3/llm-generated-text-is-not-testimony?commentId=crnYQQQoE6agAonst
I gave my reason:
Like, to me it seems pretty rude for Bob to
not spend basic effort to understand what Alice is saying
go off on his thing
while posing as being responsive to what Alice is saying
It’s just rude. It’s not a huge deal.
I think the problem with this isn’t rudeness as such, but with deceiving people (who aren’t paying close enough attention) that Bob’s comment was relevant to Alice’s thesis. (A lot of politeness norms are about concealing or obfuscating information, and I often want to rudely defy those, but I don’t want to deceive people; the category of rudeness is lumping together different things that I want to treat very differently.)
Ok. I could quibble with “to a large extent” (actually, I took that to mean at least a majority, but now I’m not sure if that’s what you mean by that). But I think I agree that a lot of politeness norms are about obfuscating.
...But there’s also plenty of politeness that’s not about that. For example:
keeping out of personal space
coughing without covering your mouth
not being too loud
ignoring what someone is saying when you’re talking to them
interrupting
hijacking attentional spaces
not doing some amount of interpretive labor
dismissing ideas / possibilities without due explanation or consideration
Some of these shouldn’t be universally enforced; e.g. in many cases I can quite enjoy a perfectly friendly conversation with lots of kinda-heated-sounding interruptions. Some of these are vague (e.g. what is “dismissing” and “due consideration”). Some of these are clear aggression (e.g. personal space) but some of them aren’t.
I don’t know exactly what “rudeness” is as a category, or know how to perfectly delineate each type of rudeness so that it’s already explicitly never in conflict with truth-seeking, but I don’t think I should have to know in order for us to agree that rudeness in general is bad, and to enforce rudeness norms.