in order to help people understand your perspective
I think this is inappropriately applying mistake theory to what is actually a conflict. If the problem was that people just didn’t understand the perspective expressed in one of Dai’s or Achmiz’s or my comments because it wasn’t expressed clearly enough, I think the natural solution would be to reply, “I don’t understand why you think this comment is relevant, can you spell it out for me?” I’m sure Wei would be happy to elaborate, as would I, as was Achmiz when he was still permitted to speak.
The point of the hobby-horse meme and the persecution of Achmiz is that we have things to say that you don’t want others to hear, because letting our perspectives be heard on the website we’ve been using for 17 years interferes with authors being able to control discussion of their own work. That’s a conflict. Maybe you think you can serve your side in the conflict by pretending that our side is making a mistake by not writing enough posts explaining our point of view? That trick worked to let you purge Achmiz without (yet) triggering a fatal loss-of-legitimacy for the whole site, but I don’t think it’s going to work against me (9x Curated, 4x Best of Less Wrong, 3x Less Online invited author) or Dai (7x Curated, 3x Best of Less Wrong, 3x Less Online invited author).
It’s likely that you think I’m mischaracterizing you, so I should probably make the standarddisclaimer here that I’m speaking in functionalist terms, looking for the simplest explanations that I think predict people’s behavior, even when that contradicts their self-report. I’m not saying you self-identify as being in a conflict.
Rather, I’m looking at the comment you just wrote recommending that Dai write more posts explaining his perspective and making a judgment call that it’s sufficiently “unserious” in an important sense such that it makes sense to suppose (as a high-probability but by no means certain hypothesis) that you’re motivatedly refusing to see the conflict you’re engaged in (again in a functionalist sense).
Specifically, Dai alreadywritesalot of posts explaining his perspective. Telling him to do more of that is obviously not going to help. If a dishonest person who favored authors being able to control discussions was consciously aware of being in a conflict about that and wanted to strategically undermine Dai’s position, they might insincerely suggest that Dai expend even more labor writing the kinds of posts he already does in the hopes of deceiving third parties unfamiliar with the full record into thinking that that any censorship or delegitimization Dai faces is his own fault for not being clear enough.
A big part of the reason I think it’s important to be able to talk about functionally implied conflicts is that we don’t want to create incentives for self-deception. If it would be problematic for a dishonest person to insincerely suggest that Dai needs to do more of the thing he already does in order for his complaints to have standing, then it should also be problematic (maybe not as problematic, but problematic to some nonzero degree) if someone performs the exact same speech act without being consciously insincere.
I think that some people think that we shouldn’t talk about functionally implied conflicts on the grounds that the accusation is unfalsifiable: if I’m questioning your motives, how could you possibly defend yourself except with self-reports that I’ll also doubt? The reason I disagree with this is because accusations of hidden motives aren’t fabricated from nothing: they’re inferred from contradictions between unobservable stated motives and behavior—which, crucially, is observable. If you think I’m reading this whole situation wrong when I claim that there’s a functional conflict between Dai et al. and the mod team about whether authors should be able to control discussions, that’s something you can argue on the merits and potentially embarrass me in the eyes of third parties!
I think the “conflict” frame here is useful but can also be over-applied; let me jot down some thoughts in response.
One kind of conflict that seems to arise is between people who are trying to diagnose where and how things went wrong, and the people who don’t want that to happen. This conflict is unfortunately kinda fractal. So for instance, many rationalists are trying to diagnose what went wrong with EAs and “AI safety”, while mainstream AI safety is fairly resistant to this. And then Wei Dai and Vassar and Hoffman and you and me (and Habryka to some extent) are trying to diagnose what went wrong with rationalism, while many rationalists are fairly resistant to this.
Unfortunately, within the group of people trying to diagnose what went wrong with rationalism, we end up often categorizing each other as part of the problem (because people disagree about what the causes are, and because you need to be an extremely disagreeable person to end up in that category, Wei perhaps excepted). This might be reading too much into Wei’s original comment, but I get some implicit message from it like “you’re trying to diagnose what went wrong but you are contributing to the problem by valorizing Yudkowsky’s intellectual clarity”. There are also ways in which Vassar and I each think the other is contributing to the problem (though we have a productive collaboration regardless).
One axis of disagreement about the diagnosis of things going wrong seems to be: how valuable it is to have a positive vision, versus to be able to critique flaws. Part of what frustrated me about Wei’s original comment is that I didn’t claim that rationalists were amazing at philosophy by any given objective standard—indeed, I’ve been one of the main people arguing that the whole foundation of rationalism on bayesianism was a massive mistake. Rather, I intended the phrase “beacon of intellectual clarity” to convey that it was unusually good by the standards of the broader landscape. I also didn’t claim that this intellectual clarity translated well into strategic acumen, and indeed criticized several strategic choices like MIRI’s engagement with prestigious elites.
So I basically want to say: I see that there’s something here you want to protect, and I’ve come to appreciate it much more over time. There’s also something that I (and, I infer, Habryka) want to protect—as implicitly expressed in my post it was something like “intellectual progress is in fact a valuable thing which we can aim for”. Insofar as we’re in a conflict frame with each other, you could view Wei’s original comment as an attack on that thing. But I tried to deescalate in my recent comment, because this doesn’t seem like a useful conflict to be in. I am very on board with identifying Yudkowsky’s strategic mistakes, and indeed my draft sequence on the strategic mistakes made by the major people in the field (including Yudkowsky) is now almost 20,000 words long. However, I would like that to trade off as little as possible against a shared sense that intellectual progress is possible. On my end, that involves reorienting how I interpret Wei’s comments. On your end, it would ideally involve acknowledging the thing that Habyrka and I are trying to protect, and helping us figure out how to protect it with as few tradeoffs as possible.
Unfortunately, within the group of people trying to diagnose what went wrong with rationalism, we end up often categorizing each other as part of the problem
Right.
One axis of disagreement about the diagnosis of things going wrong seems to be: how valuable it is to have a positive vision, versus to be able to critique flaws.
I don’t think that’s the axis. Obviously, if no one had any positive vision, then there would be nothing to critique. People who care about intellectual progress should definitely be trying to have positive visions.
It’s just that people who want their positive visions to be correct should also care about fixing flaws as they’re pointed out.
For example, I’m currently working on several interrelated posts in which I articulate a positive vision about signaling theory. (It’s about the conditions under which “honest” agents can successfully distinguish themselves from “dishonest” agents, but the details aren’t important.)
I think my positive vision is basically on the right track, but that doesn’t mean I’ve got it exactly right. I ran the drafts by my friend Said Achmiz via email, and he had a lot of detailed criticisms, some of which I disagree with, but some of which I recognized as serious. It spurred me to do a lot of laborious rewriting which I’m still not done with.
I think Achmiz’s criticism was really helpful for refining my positive vision. I don’t think he was being nice to me in his comments because I’m his friend. I think he was saying the same kind of things to me that he says to everyone.[1] If some people are not only personally uninterested in reading this kind of detailed, incisive engagement with their work, but don’t even want it to be published for others to read, I think the most plausible reason for that is that they don’t care whether their positive vision is correct and don’t want anyone else to find out, either.
Sorry, I know that seems like a really mean or “uncharitable” thing to say. But, well, I’m not sure what alternative theory of people’s behavior that you (Richard) think I should be embracing instead. (More below on my understanding of others’ theories.) Just, where would someone even get the idea that “critiquing flaws” and “having a positive vision” are somehow in a zero-sum competition with each other on opposite ends of an axis, such that people who are in favor of critiquing flaws are therefore less in favor of positive visions?
As far as I can tell, it seems to have something to do with a psychological hypothesis about morale. When Elizabeth von Nostrand writes about “a general trade off between authors’ experience and improving correctness” or Raymond Arnold writes that “it’d basically be the wrong call for any platform to ignore that” “content creation is harder than critique”, the idea seems to be that if people’s attempts to formulate positive visions get critiqued too vigorously, they’ll get discouraged and give up.
The implied psychological model of Less Wrong authors reminds me of my attempts to teach chess to my five-year-old niece this week. It’s not just that I let her capture material[2] in order to keep her engaged with the endeavor of learning how to make legal moves. It’s that the idea of losing material was so aversive that she didn’t seem to be able to process my attempted instruction of the form, “Okay, I’m not playing seriously; I’m going to let you capture my pony in a moment; but if I were playing seriously, I would recapture your rook here.”
Basically, I think Less Wrong authors are capable of having more emotional maturity than a five-year-old: we don’t need to falsely “let them win” in order to induce them to participate at all. If you think I’m being unrealistically optimistic about that, then I feel like I’m not the one who doesn’t believe in intellectual progress.
There’s also something that I (and, I infer, Habryka) want to protect—as implicitly expressed in my post it was something like “intellectual progress is in fact a valuable thing which we can aim for”.
But you can’t seriously have expected me to disagree that intellectual progress is in fact a valuable thing which we can aim for.
On my end, that involves reorienting how I interpret Wei’s comments. On your end, it would ideally involve acknowledging the thing that Habyrka and I are trying to protect, and helping us figure out how to protect it with as few tradeoffs as possible.
Uh, correct me if I’m misreading this, but it seems like you’re trying to deescalate the conflict by proposing mutual concessions: you’ve accepted a belief that favors my side, so I should reciprocate and accept a belief of your choice.
Selected excerpts: “So why would this argument work? Why should it work? It’s deceptive, isn’t it? And transparently so!”, “I am just not convinced by almost anything else you write in these two posts”, “Until you figure out just what is going on there, I don’t think that anything you write on this topic will manage to form any kind of coherent or sensible whole”, “I still think that you’re misinterpreting/misconstruing rather than miscommunicating”.
I think your implicit model of “trying to see things the other guy’s way” is incorrect. In particular, it’s not “adopt one of your beliefs on request”, but more like “spin up a mental sandbox which explores the implications of your belief being true”.
And so the exchange that I’m actually proposing is more like “I will spend compute on trying to figure out ways that your perspective might be more consistent and well-intentioned than I currently expect, if you do the same for me”.
The main way that this might create false maps is if the sandboxes are leaky, or if your perspective has been adversarially optimized to lead me to false conclusions. When I talk about trusting someone, one of the things I mean is “trying to simulate their perspective won’t mess my epistemics up”. E.g. you shouldn’t trust or try to mentally explore the implications of claims that a misaligned superintelligence has given you.
the idea seems to be that if people’s attempts to formulate positive visions get critiqued too vigorously, they’ll get discouraged and give up.
The implied psychological model of Less Wrong authors reminds me of my attempts to teach chess to my five-year-old niece this week
I tried out several possible angles of response, but upon reflection, I’m no longer interested in debate with you on this topic (though I might still engage if you have responses to the rest of this comment).
In particular, it’s not “adopt one of your beliefs on request”, but more like “spin up a mental sandbox which explores the implications of your belief being true”
Sure; I regret the rhetorical excess. (I think it would actually be “spin up a more powerful search for reasons your belief might be true”, not a search for its implications if true.)
And so the exchange that I’m actually proposing is more like “I will spend compute on trying to figure out ways that your perspective might be more consistent and well-intentioned than I currently expect, if you do the same for me”.
The question is, why would you propose that deal? What are the circumstances under which proposing that deal would seem like a good idea to someone?
The reason I’m asking is because if we assume that you’re truthseeking, it’s kind of a weird thing to propose, right? If you think that you’re right and that my position is inconsistent and ill-intentioned, presumably that’s because you’ve already thought about the matter and this is where you ended up. How would it improve the accuracy of your beliefs to reallocate your inference compute if I do so, too? Whether a proposed compute reallocation improves your beliefs shouldn’t depend on what I’m doing with my compute!
That is, the deal only makes sense for you to propose if you have a stake in something other than the accuracy of your own beliefs—if there’s something it would benefit you for me to believe. But then if I’m truthseeking, then there’s no reason for me to accept the deal (because whether a proposed compute reallocation improves my beliefs shouldn’t depend on what you’re doing with your compute). It seems like the deal should only go through if we both have a stake in the other person’s beliefs that isn’t about the truth or falsity of those beliefs, and there’s some reason you can’t just dump the reasoning that convinced you. That’s a weird and suspicious situation to be in, you know? Why would you propose the deal instead of just crushing my arguments on the merits in the eyes of third parties? It seems like the kind of thing you would only propose if you didn’t think you could win on the merits.
[W]hen someone who is currently trying to persuade me of something tells me that it doesn’t look I’m making enough effort to think of reasons why they’re right, that immediately makes me think they’re more likely to be wrong. Why? Because I think that if they had an argument, they would be telling me the argument, not chastising my lack of charity. The advice to be on special lookout for reasons your interlocutor is right is good in general, but your interlocutor is the last person to be trusted to give it, because [...] they have an ulterior motive.
That is, the deal only makes sense for you to propose if you have a stake in something other than the accuracy of your own beliefs—if there’s something it would benefit you for me to believe.
Yes, Alice may be talking with Bob because she wants Bob to believe something that Alice thinks is true. And Bob may be talking with Alice because he wants Alice to believe something that Bob thinks is true.
But then if I’m truthseeking, then there’s no reason for me to accept the deal (because whether a proposed compute reallocation improves my beliefs shouldn’t depend on what you’re doing with your compute).
Bob may accept the deal in order to have Alice reciprocate and spend compute on seriously pondering whether Bob might be right. Again, this makes sense to Bob even if he just cares about Alice being right, if he has reason to believe she’s currently wrong.
It seems like the deal should only go through if we both have a stake in the other person’s beliefs that isn’t about the truth or falsity of those beliefs
No, it can be about both people wanting the other one to have true beliefs, if they come in with separate beliefs about what’s true.
I think maybe you think this situation shouldn’t be stable because of something about Aumann’s agreement theorem making this persistent disagreement unstable if both Alice and Bob were actually only concerned with the truth. But Aumann’s theorem doesn’t just require mutual rationality but common knowledge of mutual rationality. That’s a high bar that approximately no conversation meets because it’s always plausible that your counter-party is acting in bad-faith. (Or that they think it’s plausible that you act in bad faith, or that they think it’s plausible that you think it’s plausible that they are acting in bad faith, etc.) At which points it seems consistent to me to (i) have mutual knowledge of disagreement, and (ii) assign high probability to your conversation partner being good-faith and rational. [Edit: Well, it’s unlikely that any human is perfect on those aspects. But sufficiently good-faith and rational to be worthwhile to keep talking with in a pretty assuming-good-faith kind of way.]
This doesn’t at all seem like an easy operation that one participant in the conversation can do with 0 effort from the other participant. Language is not very high band-width. People come into conversations with different abstractions and world models etc. It’s not feasible to dump all the reasons you believe something, because it probably hooks into deep parts of your world model. It’s not sufficient to dump an argument that convinced you because those words were converted to concepts in your brain and may not correspond to the same concepts in someone else’s brain and/or because the argument may have relied on previous beliefs you had which aren’t shared by your conversational partner. If Alice is going to convince Bob of her view, there’s definitely a lot of computational labor involved in that which is more efficiently done by Bob than by Alice.
That’s a high bar that approximately no conversation meets because it’s always plausible that your counter-party is acting in bad-faith.
Yes, and those plausible worlds are exactly why I’m so suspicious of the proposed compute reallocation deal: I’m worried that “you’re being uncharitable; try to think harder about why I’m right” is something people around here say to conceal the fact that they don’t actually have an argument.
Maybe you’re right that the deal could make sense (and my argument that it doesn’t is wrong), but it gets scuttled by the adverse selection problem (where good faith actors would prefer to make such deals with each other, but they have no way to credibly distinguish themselves from bad faith actors who are just bluffing)?
Well sometimes it gets scuttled and sometimes it doesn’t. Cars do in fact get sold second-hand, despite the adverse selection problem. Often the gains-from-trade are large enough to justify some risk.
I think there are often big gains from people putting in some work to understand where their conversation partner is coming from.
Also, “no way to credibly distinguish themselves” seems like a big overstatement. You can do way better than chance at telling whether someone is reasonable or not, especially in repeated interactions:
For one, if someone says “you’re being uncharitable; try to think harder about why I’m right”—you can see that as a reputational bet that you’ll learn something interesting from thinking harder about why they’re right. If you try that a couple of times and you don’t learn anything new, then that’s reasonable inductive evidence that you won’t benefit from that exercise with that person again. You’re not infinitely exploitable here.
And from the perspective of the “deal”—where you put in some effort to understand their perspective, and they put in some effort on understanding yours, even if you don’t expect your part of the deal to be worthwhile in isolation—I think it’s pretty easy to get evidence for whether they’ve held up their part of the deal. Like: Can they demonstrate any understanding of what you’re trying to say? (In the limit this is “can they pass your ITT” but conversations often involve much weaker versions, like “so you’re saying [blah]”, or even less explicit demonstrations of understanding, like engaging with what the other person see as the central point.)
The reason I’m asking is because if we assume that you’re truthseeking, it’s kind of a weird thing to propose, right?
I like this line of inquiry, thank you. If I were working in a Bayesian epistemology I expect I’d find this argument fairly persuasive.
So I think the core reasons I disagree are related to my post on why I’m not a bayesian. That is: I think of intellectual progress in terms of growing a new ontology/developing a new model of the world. At first, this ontology will be fairly illegible to other people. Gradually, I’ll find ways to get it to generate novel falsifiable predictions, or figure out how it overlaps with other peoples’ ontologies. The more we can find overlap, the more productively we can communicate insights to each other. But the procedure of identifying the correspondence between your ontology and my ontology often requires a bunch of work. Insofar as we’re both doing it, we can “build the bridge from both ends” to find the overlapping sections and then communicate insights about those. This is much harder if one person is doing far more of the interpretative labor.
Do you have an example of intellectual progress that fits this model? “growing a new ontology/developing a new model of the world” makes sense to me, but I guess the way I’ve been doing it is to develop it mostly in private (maybe sharing it as I go but, from personal experience, I don’t expect anyone to have an interest in it, unless they independently happen to have the same intuitions / research tastes as me), until I develop some kind of minimum viable product (e.g. UDT 1.0, or the Bitcoin whitepaper as an example from someone else) based on the new ontology/model that is legibly good (e.g., clearly solves some open problems, seems clearly insightful to other people without them having to do a lot of intellectual/interpretive work, etc.).
To share more of my model, I think most attempts to make intellectual progress are likely to fail, e.g. are not aimed in a right direction. So you may think you’re growing a new ontology/developing a new model of the world, but most likely you’re just developing a bunch of nonsense but you can’t see that yet (e.g. due to lack of compute to find a contradiction in your ideas). So then if two such people try “build the bridge from both ends” that’s just very likely to not result in something useful. (I.e., it requires the two people to both be on the right track, which is multiplicatively unlikely.)
I think most attempts to make intellectual progress are likely to fail
I wonder if one issue with people’s models of intellectual progress is that ~nobody on LW saw the ~10 years I spent exploring various dead ends / unclear results in anthropic reasoning and decision theory[1], they just saw UDT 1.0 pop almost out of nowhere. (Perhaps some saw the few months of decision theory discussion on LW before my post, where I got some of the final pieces of the puzzle from Eliezer and Nesov’s hints.)
As for Bitcoin, I don’t know how long it took Satoshi, but there were hundreds or thousands of published papers on digital cash designs of all kinds[2], before Bitcoin came on the scene and took off, but ~everyone just saw Bitcoin and don’t know the rest.
Sorry if this is being unfair/uncharitable to you[3], @Richard_Ngo, but I suspect this might be what’s accounting for our apparent disagreements around intellectual progress.
I posted half-baked ideas to my own everything-list mailing list, but didn’t get much attention from others, except Hal Finney who liked and wrote up UDASSA on his website, which somehow spread from there even though I wasn’t all that happy with it.
My boss during my internship at Microsoft Research asked me to work on his own design, but I let it quietly fall off my plate because it seemed unpromising, and I had my own ideas at that point.
Maybe you’ve observed similar phenomenon elsewhere and still reached different conclusions than me, or I’m completely misunderstanding where our disagreement lies.
Here, I’ll work under the assumption that we need (at least) some enforcement of the excluded middle, since, eventually, we must use our thinking to take some precise action to the exclusion of all others. This action can always be specified as a finite length binary string. If you want, you can construct a class of “fair” scenarios where the agent will not get into trouble if it processes the binary string corresponding to the action it just took.
As to my intentions, I want to discuss requiring self-trust here WOLOG, which is slightly hard to do, but essentially think something like “in cases where we are self-trust fair, our agent will meet self-trust desiderata (up to practicality) on its future belief on its (binary) decision string IF (classical logic implies, locally no claim made on other parts of the truth table) the environment is string-processing fair AND it epsilon-prefers so inspecting.”
(WOLOG for the desiderata, the test situations admit a small epsilon > 0 (for the mix-in multiplier) that leaves other behavior the same.)
(With much informality) we want, for our decisions, something like what is called “LUV Coherence.” Or, potentially, we want something like “that our eventual decision bit-strings are chosen in a way that (robustly to adversarial attack) closely approximates some decision that is made on the basis of a set of dynamic variables (LUVs) that is entirely coherent.”
(Somewhat unnecessarily for certain audiences, we want to state that our other performance desiderata is left out here.)
Effectively, all of this so far lets us work backwards from the initial excluded middle requirement, through (potentially off-policy/off the collective policy) self trust, then through (decision-approximated) coherence of decision-relevant LUVs.
(As an aside, it is often difficult to tell information that “might” be useful from information that is “teleologically” useful for the actual decisions or (desired) computational processes that actually occur. This can be solved on a theoretical level by requiring that our agent has a fixed core design, but is able to do as well as possible in a large number of scenarios. During design, we will be careful to avoid overfitting to test environments, but we will still prune all computation that can be proven to not be worth the cost.)
Once we’ve established the state of the interface between the agent’s epistemics and its actions, we can examine what constraints that puts on the agent’s internals. Notably, this includes two desiderata, to my knowledge novel statements of the problem:
First new desiderata: coherence at the interface must be efficiently supported by the interior, i.e. it must not be computationally wasteful and not put at risk of expulsion anything that is more useful to the final, excluded-middle-enforced action.
Second new desiderata: the internals must prioritize atomicity, separability, and robustness over exact correctness or any aesthetic sense of “unification” in order to operate under the ruthless and short-tempered expulsion procedures required by computational limitation.
(Note that expulsion is often by greedy and immediate matching on syntactic patterns, and even this approach may be too expensive, depending on the ultimate limits of technology.)
These new desiderata (not entirely formally) motivate an approach that uses discardable “traders” that have a limited impact on near-future agent optimality if incorrectly thrown out, and where any single expulsion does not affect optimality in the limit.
Why the agents should trade on statements of basic logic: computers are unable to operate on anything “richer,” except in the way e.g. Metamath implements further math, by assuming further axioms. A reason you might do this anyway: there is a practical improvement in operation speed, even accounting for the abstraction cost, without damaging (adversarial) robustness. Why we can simplify this away (at least for agents running on million-year or less time spans): by the demand for coherence at the decision level, atomicity, separability, and robustness, there is a limit to how much you can deviate from logic (while maintaining performance, desiderata left out here). This should be provable by ruling out obviating alternatives to syntactic enforcement (supporting and being supported by atomicity, separability, and robustness requirements) and showing adversarial attacks against improper further assumptions (re. robustness).
Presently, I will simplify down to the case where we have some (modified) LI traders used according to the overall context of this comment. Tell me if this step is too bold. Also, note that this account is actually simplified vs. what would work for a real decision theory, for the sake of ease of reasoning and to avoid following unpublished accounts.
In an attempt to describe what you propose, while meeting the desiderata, imagine we have a pair of traders, tr1 and tr2. (Note that sharing can be chained as long as there are no loops at the subroutine level, but we will stick to a pair here WOLOG.) Let’s say we have a definition of “efficiently computable” (e.c.) that goes false when an individual trader takes on too much work per “day,” and also goes false on a collective level if there is too much compute demanded overall. Say tr1 is e.c., but {tr1, tr2, …, tr} is not. Stipulate the set is e.c. when tr2 is removed. Let’s say we try to transform tr2 (not e.c.) into tr2 prime (e.c. when allowed to reference tr1), by letting tr2 prime reference some subroutine of tr1, and letting the cost reduction be accepted in the standard way when determining e.c. in both the individual (when allowed to reference) and collective definitions. Say that the set {tr1, tr2 prime, …, tr} is e.c. Since, as can clearly be seen, this situation is not symmetrical (tr1 can survive alone, but tr2 can not), a well calibrated model with holdouts would say that tr2 (in the operational form of tr2 prime) is “less probable” to survive than tr1. (Holdout method not given, compute-limited modeling roughly following the standard for LIs.) Note that to the extent that any of these “traders” have “inner” instrumental rationality, they are unable to take any “action” other than outputting a trading strategy for the “day,” for fairly standard betting market incentive reasons. For this reason, we assume the set of “traders” is brought into existence by some outer process in their operational forms, with any references already established. Eviction must be extremely efficient. I am unsure if this prohibits re-parenting of a subroutine, but for simplicity assume that the operational set wouldn’t allow tr2 to survive even if it was allowed. Deviating a bit from formality, generalize tr1 to trx and tr2 prime to try, both in the domain of situations of this sort. Define the function ev_p(tr) to be the bounded-compute “probability” of the eviction of tr (see before on modeling). Evaluate ev_p(try) - ev_p(trx). Call this “the cost of ontological overlap.”
As said before, even if a “trader” has a decision theory set up internally, it won’t be hooked up to the interface between the “trader” and the (actual) wrapping hardware in any traditional sense. However, let’s look at a case where we imagine some alternative to tr2 that could have done better than tr2, without having any ontological overlap, while being e.c. Call it trv. Notice that it has, at the absolute least, regained the entire (non-formally made specific again) “cost of ontological overlap.”
Surely, there can be savings by sharing the execution of subroutines, but this trades off against atomicity and separability. Depending on the eventual costs of things like re-parenting, and more advanced techniques like the automatic determination of acceptable approximations and the determination of “equivalent” functions, this may trade off substantially against robustness as well, since the eviction of a single “trader” will sometimes result in a massive over-eviction, if the dependencies can’t be fixed up in an extremely rapid manner (against optimality as well, but we omit that here).
This attempts to model “communication” at some level (actually quite well, given some theories of communication and collective rationality I can vaguely remember), but note that, treating it as a sort of agent it usually won’t be, trv has no “reason” to want to be replaced with tr2 prime. With some basic assumptions on the setup, e.g. strictly positive initial trader budgets, trv being replaced with tr2 prime is a strict decrease in estimated survival “probability.” Therefore, this “communication”/sharing must only be done to the extent that it really helps the overall agent, and not following some procedure or “virtue” that is thought to be worth universally adopting.
(We’re assuming the set with trv replacing tr2 prime is also e.c. Note that in this entire comment we’re assuming tr2 and tr2 prime are functionally identical, to the level of detail we care about here.)
Okay, then let’s try to go back to English. Let’s think about these “traders” as cartoons, in the same sense that a cartooned evolution can speak and (be said to) want things. So this trader says, “I want to share load onto other models run by other traders if and only if it reduces my compute demands enough to achieve viability. Sharing ontology for other reasons just makes me worse off, in expectation. Depending on other ontologies is a risk that must be taken on only after sufficient calculation, since ontologies (via the traders that operate them) may go bust.”
(In a more realistic system, it would need to be considered what to do when a “trader” is “exporting” a subroutine that it doesn’t want to use any more. For simplicity, you can ignore this for now. Note that the cartoon dialogue doesn’t really survive this, but it’s cartoon dialogue for a reason. Note that “going bust” should be read as referring to “traders” getting evicted, in ways that don’t strictly relate to bankroll.)
I can’t think of an alternative to this setup, and since the (new) intermediate desiderata seem well motivated, I’m hesitantly taking this as a (sketched, not fully filled in) disproof. One reason for my hesitancy is that my model of proper thinking may require too much compute and too much implementation complexity to ever be usable. Sure, maybe I can fill in all the details and get a system that’s proven aligned, but it might not be applicable to humans.
Do you know of a technical or semi-technical note on how this ontology framework (as an argument for sharing load) functions here/meets desiderata? It doesn’t need to be an explainer, so don’t particularly worry about the quality or any missing details.
I’m not an expert in this area, so maybe the solution really is simple and I’m just not seeing it, or I’m making a silly mistake.
(My hand written heuristics claim this comment is dense to the point of near-unreadability, so all (prospective) readers are invited to ask questions, including if a reading attempt has yet to be made.)
(I think it would actually be “spin up a more powerful search for reasons your belief might be true”, not a search for its implications if true.)
The step of running with the alien assumptions and framings is also very important, it’s unnecessary and impractical to only focus on the things you believe or understand or endorse. There is much more to learn than what’s ready to be made into a part of yourself, and the only way to defeat path-dependence at human level is by being anti-inductive with respect to your own perspective, to seek out the points of view you don’t understand or endorse or believe (on some topic of interest), and at least come to understand how they work internally for their proponents, even when you don’t have footholds of belief or endorsement into them that promise they might be a good fit down the line. (This does require solid sandboxing practices, or else your mind might get so open the brains fall out.) Sometimes such footholds are unexpected, and only appear after you do the work, despite not having any footholds at the outset, and so you make your own epistemic luck.
That is, the deal only makes sense for you to propose if you have a stake in something other than the accuracy of your own beliefs—if there’s something it would benefit you for me to believe. But then if I’m truthseeking, then there’s no reason for me to accept the deal
I didn’t think this through carefully, but maybe an epistemic PD where the currency is compute? Like, I spend 5 utils worth of compute; from your perspective that’s worth 10 expected utils because I might correct my false (according to you) belief; and you do the same for me.
It’s just that people who want their positive visions to be correct should also care about fixing flaws as they’re pointed out.
In principle, someone else could produce the corrected positive vision. It’s not obviously crucial that it’s specifically the author of the flawed inspiration who also does the subsequent refinement. Incentives that shape the original contributions are probably a more important aspect of this than the possible later causal back-and-forth. Criticism being unpleasant is some sort of incentive (to be correct, or to avoid contributing). Criticism being valuable fuel for refinement is also some sort of incentive (to exploit Cunningham’s Law by being wrong so that you can learn the right answer from the critics, or to start contributing even when you have nothing to say). So it’s not straightforward, even though at first glance it seems like normatively it should be.
But, well, I’m not sure what alternative theory of people’s behavior that you (Richard) think I should be embracing instead.
I’d be really very surprised if you hadn’t before heard the argument about nobody having infinite time, bad signal:noise ratios, etc. In fact I’m confident I’ve made this argument to you sometime in the last ~year. Of course, resolving that one would require digging into object-level details, which hasn’t exactly been a fruitful endeavour in the past, maybe because of different thresholds that various individuals have for what signal:noise ratio they find tolerable, and the contingent facts about reality that permit people with different thresholds from you in particular to successfully contribute to advancing the state of human knowledge, even if they don’t want to deal with random nutpickers who once in a blue moon will point out a meaningful error in their post, while their other 100 comments are wrong, confused by something that almost nobody else is confused by, focusing on some random triviality that isn’t load-bearing for the core argument… etc. (Numbers made up; I am, again, establishing some least convenient possible world so that we can skip to the part where we agree that there’s a spectrum and get to arguing about where on the spectrum we should live.)
Perhaps you think that Richard doesn’t endorse this argument, and so you shouldn’t bother to bring it up as a hypothesis? But you asked about what theory you should be embracing. I propose the above theory: I don’t believe you’ve provided much evidence (that I can recall) that your threshold is, in this way, better for advancing the state of rationality/the frontier of human knowledge/etc.
resolving that one would require digging into object-level details, which hasn’t exactly been a fruitful endeavour in the past,
I do wonder if there’s some way to make this kind of debate something like an order of magnitude more efficient using LLMs. By “this kind” I mean “a big huge sprawling debate across many different posts and comments through years, where one or more interlocutors keep forgetting bits about another interlocutor’s perspective and/or not understanding it well enough to bring the relevant facts & arguments to bear in a progressful context for thinking it through”. The help I’m imagining is something like “scrape a big database of all posts and all comments involved in the debate; then have one or more participants talk to an LLM with things like....
how would the other side respond to this; use quotations from their past writing
arguing against the other side with LLM as devil’s advocate
doing ITTs
asking if the other side has addressed X
asking if the other side seems to understand X, care about X, etc.
etc.
....until they reach quasi-equilibrium with something they actually want to put directly to another person. Maybe more efficient because it reduces redundant labor on both sides, and in particular helps with the annoying “routing” task where two clashing perspectives keep having to figure out which standard paragraph from its own perspective’s library of paragraphs is relevant to whatever the other person is confused about. (And maybe psychically easier to argue with a neutral LLM where you don’t have to be defending yourself, upholding boundaries, prosecuting a conflict, etc.)
In fact I’m confident I’ve made this argument to you sometime in the last ~year.
Indeed, it was only 84 days ago. I’m listening if you have a response to my last comment in that thread, in which I agreed that censorship to maintain the signal-to-noise threshold is good. I then linked to a compilation of Achmiz’s favorites of his own comments and offered the judgments of myself and Jessica Taylor that Achmiz’s work is well above the threshold of being worthy of published on Less Wrong.
Now, maybe you think Taylor (former MIRI employee, inventor of quantilizers and co-author of the logical induction paper, and 7x Curated author) and I (9x Curated, 4x Best of Less Wrong, 3x Less Online invited author) have terrible judgment about what advances the state of rationality/the frontier of human knowledge/&c. Is that your position? Happy to dig into the details if you want.
There’s also something that I (and, I infer, Habryka) want to protect—as implicitly expressed in my post it was something like “intellectual progress is in fact a valuable thing which we can aim for”. Insofar as we’re in a conflict frame with each other, you could view Wei’s original comment as an attack on that thing.
This confuses me. Can you explain what you sense as a possible motivation on my part to attack “intellectual progress is in fact a valuable thing which we can aim for”?
My own interpretation of what the conflict is, is that everyone here wants to support intellectual progress, but have different ideas how to go about it due to being biased due to status seeking, sunken costs, etc., some of which is strong/ingrained enough to make it intractable to fix the underlying mistakes, so we can only fight it out as a conflict. (To be clear I’m not certain about this.)
On your end, it would ideally involve acknowledging the thing that Habyrka and I are trying to protect, and helping us figure out how to protect it with as few tradeoffs as possible.
Does this still make sense given what I wrote above?
To be clear, I don’t see myself as being in a conflict with you, but with the site mods and the LW cultural trajectory that they’re guiding. Ironically it was initially set on this path when Eliezer demanded author mod powers as a condition of coming back to LW (from FB/Twitter), but now it’s been so ingrained through past decisions that it seems impossible for the mods to admit this was a mistake (via ordinary deliberation), hence why conflict theory feels more appropriate at this point.
I take the side of the site mods on this issue, though, for the reasons articulated in my comment above.
I would also bid for you to do the thing that I recommended Zack do above, namely “acknowledging the thing that Habyrka and I are trying to protect, and helping us figure out how to protect it with as few tradeoffs as possible”.
I think mistake theory makes more sense for the disagreement between us. Because you don’t have nearly as much “sunk cost” invested in the current culture (e.g. was not responsible for cultivating it in the first place), I infer that you can change your mind much more easily about what kind of site culture is more conducive for clear thinking or intellectual progress. (But of course having a mistake theoretic debate with you about this would be pointless given the existence of the main conflict.)
I would also bid for you to do the thing that I recommended Zack do above, namely “acknowledging the thing that Habyrka and I are trying to protect, and helping us figure out how to protect it with as few tradeoffs as possible”.
This seems to be ignoring my point about why I think this is a conflict theoretic situation.
This seems to be ignoring my point about why I think this is a conflict theoretic situation.
You mean that “it was initially set on this path when Eliezer demanded author mod powers”? You could interpret this as Eliezer wielding power in ways you don’t like, but you could also interpret it as the mods treating Eliezer’s preferences as evidence about what site norms produce intellectual progress (along with many other people’s preferences) and acting accordingly. I assume you don’t think the latter is a good model; if so, why not?
Because you don’t have nearly as much “sunk cost” invested in the current culture (e.g. was not responsible for cultivating it in the first place), I infer that you can change your mind much more easily about what kind of site culture is more conducive for clear thinking or intellectual progress.
Two responses. Firstly, I acknowledge that you did a lot to cultivate the site culture in the first place, and I’m grateful for that.
Secondly: one of the most difficult parts of rationality seems to be changing one’s mind in the face of sunk costs. Because of that, I agree that it’s easier for me to change my mind on this than it is for you. But do you endorse the extent to which sunk costs are making it harder for you to change your mind?
I remember it as being presented as a fait accompli: Eliezer and other unnamed authors demanded this, and we already implemented it. “We prioritized building the delete-and-hide function because Eliezer asked for it and we wanted to get him posting again quickly. But he is not the only author to have asked and expressed appreciation for it.”
See also this comment where Raemon guesses at Eliezer’s reasons for his demand, suggesting that they did not consider Eliezer’s reasons to be centrally important (i.e. didn’t even ask Eliezer to write them up), as opposed to the demand itself: “I’m not that confident in the following, and I don’t want this to turn into a psychoanalyze Eliezer subthread and will lock it if it appears to do that” (The “will lock it” further suggests that the overall decision is not up for discussion, because otherwise understanding/debating Eliezer’s reasons would seem to be central.)
Or even simpler, I don’t remember them ever asking for my preferences about author mod powers.
EDIT: It looks like you added more after I wrote my reply, which is just as well as I feel like falling into a familiar trap by engaging in this debate with you.
Specifically, Dai alreadywritesalot of posts explaining his perspective. Telling him to do more of that is obviously not going to help.
The principle of more dakka is needed because people often don’t do more of the thing that is working, so I don’t agree that this obviously follows. Perhaps more relevantly, Wei Dai never linked to any of those posts in either of his comments on TsviBT’s or Ngo’s posts, and I had forgotten about them.
(But also I want to disclaim that I am not confident in this read of the situation, it just seemed to me “about on the level of being worth considering” as the other things Wei listed.)
I don’t have a job and participating on LW has been my main hobby, so “just do more of it” is clearly not a general/scalable solution.
My goal (which I endorse on reflection) is to write comments/posts that a lot of people get value out of (which I think I had succeeded with the comments in question), not convince or interest every person, who may be motivatedly obtuse to its relevance or importance.
In this case I thought I was making a fairly surgical correction (that MIRI/early rationalists were not some sort of beacon of clear thinking in contrast with EA/other institutions, but also made plenty of high-impact mistakes) that didn’t depend on my previous writings. (Adding links potentially implies that reading the linked content is necessary to understand the current content.)
Writing potentially half-baked comments is how I (and I imagine others) find out whether some idea or line of thought is worth putting more effort into, which parts are hard for others to understand, etc., but which now risks getting me labeled (even more) as something that LW culture seems to think is bad (given that the relevant post/accusations all had fairly high karma), and potentially getting author-banned.
I generally agree with Zack that at this point this disagreement is probably more about conflict theory than mistake theory, despite how the above points may read (i.e., it may seem like I’m adopting more of a mistake theory view where I’m trying to explain reasoning that maybe just hasn’t occurred to you). (In part I’m just more used to operating under mistake theory framing, and in part I’m not fully certain about this.)
I have often seen it be the case that, during a conflict between rationalists (a shorthand referring to people active in the rationalist-scene), someone offers a mistake theory interpretation for what led up to the conflict, and the other party finds that to be a valid account of a mistake that they hadn’t noticed, and then all the air is removed from the conflict.
I think this move isn’t always a misdirect. Right now, to give an overly-specific hypothesis, I think it seems like a live hypothesis that part of the reason it was highly upvoted was because it was part of Wei’s ongoing critique of Eliezer/MIRI, and partly that seems upvoted because it’s drama that people are into, rather than being related to their post, and this is kind of distracting from the topics the authors are working to discuss.
I am not accusing Wei Dai of this! I am just saying that, if the authors are finding the critique not very good / non-central—which could well be because on some level they just don’t want to hear it—then this could seem like a more plausible hypothesis to them. And so it would be effective to make the relevance clearer to resolve this issue.
I think your response here though is “But overall the comment shouldn’t solely or even primarily be evaluated on whether it was good by the author’s lights, it should be evaluated by whether it’s helping the discourse and ideas move forward” and I agree, often comments that the author finds annoying and stupid, are in fact good, and valued by the many readers.
(To expand on that: I agree that it is somewhat the case that authors should learn better to deal with the top-voted comment on their post being non-central, or a critique that they don’t value very much, because much of the point is for the readers and the general discourse, not the authors, and I think it’s plausible that this is the case here. I currently don’t see cause for either author to ban Wei based on these two comments, and I thought it was an overreaction by Tsvi at the time.)
But my response is that the author can still sometimes get it right? I don’t think that the author’s propensity for self-deception means that they can never come to a justified true belief that the commenter is norm-violating (e.g. hobby-horsing) and should be banned for it, and take action based on that. And when I look at the list of bans on the /moderation page, personally I judge that most of them seem pretty reasonable and understandable based on annoying behaviors and not primarily the basis of (functionally) seeking to avoid good criticism of their ideas. (There are of course some exceptions that I disagree with, and there are certainly rates of unreasonable ones, or rates of egregiously bad ones, that would make me want to take action to counter it.)
I think this is inappropriately applying mistake theory to what is actually a conflict. If the problem was that people just didn’t understand the perspective expressed in one of Dai’s or Achmiz’s or my comments because it wasn’t expressed clearly enough, I think the natural solution would be to reply, “I don’t understand why you think this comment is relevant, can you spell it out for me?” I’m sure Wei would be happy to elaborate, as would I, as was Achmiz when he was still permitted to speak.
The point of the hobby-horse meme and the persecution of Achmiz is that we have things to say that you don’t want others to hear, because letting our perspectives be heard on the website we’ve been using for 17 years interferes with authors being able to control discussion of their own work. That’s a conflict. Maybe you think you can serve your side in the conflict by pretending that our side is making a mistake by not writing enough posts explaining our point of view? That trick worked to let you purge Achmiz without (yet) triggering a fatal loss-of-legitimacy for the whole site, but I don’t think it’s going to work against me (9x Curated, 4x Best of Less Wrong, 3x Less Online invited author) or Dai (7x Curated, 3x Best of Less Wrong, 3x Less Online invited author).
It’s likely that you think I’m mischaracterizing you, so I should probably make the standard disclaimer here that I’m speaking in functionalist terms, looking for the simplest explanations that I think predict people’s behavior, even when that contradicts their self-report. I’m not saying you self-identify as being in a conflict.
Rather, I’m looking at the comment you just wrote recommending that Dai write more posts explaining his perspective and making a judgment call that it’s sufficiently “unserious” in an important sense such that it makes sense to suppose (as a high-probability but by no means certain hypothesis) that you’re motivatedly refusing to see the conflict you’re engaged in (again in a functionalist sense).
Specifically, Dai already writes a lot of posts explaining his perspective. Telling him to do more of that is obviously not going to help. If a dishonest person who favored authors being able to control discussions was consciously aware of being in a conflict about that and wanted to strategically undermine Dai’s position, they might insincerely suggest that Dai expend even more labor writing the kinds of posts he already does in the hopes of deceiving third parties unfamiliar with the full record into thinking that that any censorship or delegitimization Dai faces is his own fault for not being clear enough.
A big part of the reason I think it’s important to be able to talk about functionally implied conflicts is that we don’t want to create incentives for self-deception. If it would be problematic for a dishonest person to insincerely suggest that Dai needs to do more of the thing he already does in order for his complaints to have standing, then it should also be problematic (maybe not as problematic, but problematic to some nonzero degree) if someone performs the exact same speech act without being consciously insincere.
I think that some people think that we shouldn’t talk about functionally implied conflicts on the grounds that the accusation is unfalsifiable: if I’m questioning your motives, how could you possibly defend yourself except with self-reports that I’ll also doubt? The reason I disagree with this is because accusations of hidden motives aren’t fabricated from nothing: they’re inferred from contradictions between unobservable stated motives and behavior—which, crucially, is observable. If you think I’m reading this whole situation wrong when I claim that there’s a functional conflict between Dai et al. and the mod team about whether authors should be able to control discussions, that’s something you can argue on the merits and potentially embarrass me in the eyes of third parties!
I think the “conflict” frame here is useful but can also be over-applied; let me jot down some thoughts in response.
One kind of conflict that seems to arise is between people who are trying to diagnose where and how things went wrong, and the people who don’t want that to happen. This conflict is unfortunately kinda fractal. So for instance, many rationalists are trying to diagnose what went wrong with EAs and “AI safety”, while mainstream AI safety is fairly resistant to this. And then Wei Dai and Vassar and Hoffman and you and me (and Habryka to some extent) are trying to diagnose what went wrong with rationalism, while many rationalists are fairly resistant to this.
Unfortunately, within the group of people trying to diagnose what went wrong with rationalism, we end up often categorizing each other as part of the problem (because people disagree about what the causes are, and because you need to be an extremely disagreeable person to end up in that category, Wei perhaps excepted). This might be reading too much into Wei’s original comment, but I get some implicit message from it like “you’re trying to diagnose what went wrong but you are contributing to the problem by valorizing Yudkowsky’s intellectual clarity”. There are also ways in which Vassar and I each think the other is contributing to the problem (though we have a productive collaboration regardless).
One axis of disagreement about the diagnosis of things going wrong seems to be: how valuable it is to have a positive vision, versus to be able to critique flaws. Part of what frustrated me about Wei’s original comment is that I didn’t claim that rationalists were amazing at philosophy by any given objective standard—indeed, I’ve been one of the main people arguing that the whole foundation of rationalism on bayesianism was a massive mistake. Rather, I intended the phrase “beacon of intellectual clarity” to convey that it was unusually good by the standards of the broader landscape. I also didn’t claim that this intellectual clarity translated well into strategic acumen, and indeed criticized several strategic choices like MIRI’s engagement with prestigious elites.
So I basically want to say: I see that there’s something here you want to protect, and I’ve come to appreciate it much more over time. There’s also something that I (and, I infer, Habryka) want to protect—as implicitly expressed in my post it was something like “intellectual progress is in fact a valuable thing which we can aim for”. Insofar as we’re in a conflict frame with each other, you could view Wei’s original comment as an attack on that thing. But I tried to deescalate in my recent comment, because this doesn’t seem like a useful conflict to be in. I am very on board with identifying Yudkowsky’s strategic mistakes, and indeed my draft sequence on the strategic mistakes made by the major people in the field (including Yudkowsky) is now almost 20,000 words long. However, I would like that to trade off as little as possible against a shared sense that intellectual progress is possible. On my end, that involves reorienting how I interpret Wei’s comments. On your end, it would ideally involve acknowledging the thing that Habyrka and I are trying to protect, and helping us figure out how to protect it with as few tradeoffs as possible.
Right.
I don’t think that’s the axis. Obviously, if no one had any positive vision, then there would be nothing to critique. People who care about intellectual progress should definitely be trying to have positive visions.
It’s just that people who want their positive visions to be correct should also care about fixing flaws as they’re pointed out.
For example, I’m currently working on several interrelated posts in which I articulate a positive vision about signaling theory. (It’s about the conditions under which “honest” agents can successfully distinguish themselves from “dishonest” agents, but the details aren’t important.)
I think my positive vision is basically on the right track, but that doesn’t mean I’ve got it exactly right. I ran the drafts by my friend Said Achmiz via email, and he had a lot of detailed criticisms, some of which I disagree with, but some of which I recognized as serious. It spurred me to do a lot of laborious rewriting which I’m still not done with.
I think Achmiz’s criticism was really helpful for refining my positive vision. I don’t think he was being nice to me in his comments because I’m his friend. I think he was saying the same kind of things to me that he says to everyone. [1] If some people are not only personally uninterested in reading this kind of detailed, incisive engagement with their work, but don’t even want it to be published for others to read, I think the most plausible reason for that is that they don’t care whether their positive vision is correct and don’t want anyone else to find out, either.
Sorry, I know that seems like a really mean or “uncharitable” thing to say. But, well, I’m not sure what alternative theory of people’s behavior that you (Richard) think I should be embracing instead. (More below on my understanding of others’ theories.) Just, where would someone even get the idea that “critiquing flaws” and “having a positive vision” are somehow in a zero-sum competition with each other on opposite ends of an axis, such that people who are in favor of critiquing flaws are therefore less in favor of positive visions?
As far as I can tell, it seems to have something to do with a psychological hypothesis about morale. When Elizabeth von Nostrand writes about “a general trade off between authors’ experience and improving correctness” or Raymond Arnold writes that “it’d basically be the wrong call for any platform to ignore that” “content creation is harder than critique”, the idea seems to be that if people’s attempts to formulate positive visions get critiqued too vigorously, they’ll get discouraged and give up.
The implied psychological model of Less Wrong authors reminds me of my attempts to teach chess to my five-year-old niece this week. It’s not just that I let her capture material [2] in order to keep her engaged with the endeavor of learning how to make legal moves. It’s that the idea of losing material was so aversive that she didn’t seem to be able to process my attempted instruction of the form, “Okay, I’m not playing seriously; I’m going to let you capture my pony in a moment; but if I were playing seriously, I would recapture your rook here.”
Basically, I think Less Wrong authors are capable of having more emotional maturity than a five-year-old: we don’t need to falsely “let them win” in order to induce them to participate at all. If you think I’m being unrealistically optimistic about that, then I feel like I’m not the one who doesn’t believe in intellectual progress.
But you can’t seriously have expected me to disagree that intellectual progress is in fact a valuable thing which we can aim for.
Uh, correct me if I’m misreading this, but it seems like you’re trying to deescalate the conflict by proposing mutual concessions: you’ve accepted a belief that favors my side, so I should reciprocate and accept a belief of your choice.
That’s not how it works. Treating epistemics as social exchange—trying to see things the other guy’s way on his request, in exchange for him trying to see it your way on your request—doesn’t create true maps; it creates false maps representing a compromise between the parties’ preferred lies. I put a lot of effort into explaining why I don’t think Habryka’s observed behavior is well-described as trying to protect intellectual progress. If you think I’m wrong about that, you should be able to explain why I’m wrong. It’s not a trade!
Selected excerpts: “So why would this argument work? Why should it work? It’s deceptive, isn’t it? And transparently so!”, “I am just not convinced by almost anything else you write in these two posts”, “Until you figure out just what is going on there, I don’t think that anything you write on this topic will manage to form any kind of coherent or sensible whole”, “I still think that you’re misinterpreting/misconstruing rather than miscommunicating”.
It seems imprecise to say “let her win”, since I despaired of trying to explain the concept of checkmate.
I think your implicit model of “trying to see things the other guy’s way” is incorrect. In particular, it’s not “adopt one of your beliefs on request”, but more like “spin up a mental sandbox which explores the implications of your belief being true”.
And so the exchange that I’m actually proposing is more like “I will spend compute on trying to figure out ways that your perspective might be more consistent and well-intentioned than I currently expect, if you do the same for me”.
The main way that this might create false maps is if the sandboxes are leaky, or if your perspective has been adversarially optimized to lead me to false conclusions. When I talk about trusting someone, one of the things I mean is “trying to simulate their perspective won’t mess my epistemics up”. E.g. you shouldn’t trust or try to mentally explore the implications of claims that a misaligned superintelligence has given you.
I tried out several possible angles of response, but upon reflection, I’m no longer interested in debate with you on this topic (though I might still engage if you have responses to the rest of this comment).
Sure; I regret the rhetorical excess. (I think it would actually be “spin up a more powerful search for reasons your belief might be true”, not a search for its implications if true.)
The question is, why would you propose that deal? What are the circumstances under which proposing that deal would seem like a good idea to someone?
The reason I’m asking is because if we assume that you’re truthseeking, it’s kind of a weird thing to propose, right? If you think that you’re right and that my position is inconsistent and ill-intentioned, presumably that’s because you’ve already thought about the matter and this is where you ended up. How would it improve the accuracy of your beliefs to reallocate your inference compute if I do so, too? Whether a proposed compute reallocation improves your beliefs shouldn’t depend on what I’m doing with my compute!
That is, the deal only makes sense for you to propose if you have a stake in something other than the accuracy of your own beliefs—if there’s something it would benefit you for me to believe. But then if I’m truthseeking, then there’s no reason for me to accept the deal (because whether a proposed compute reallocation improves my beliefs shouldn’t depend on what you’re doing with your compute). It seems like the deal should only go through if we both have a stake in the other person’s beliefs that isn’t about the truth or falsity of those beliefs, and there’s some reason you can’t just dump the reasoning that convinced you. That’s a weird and suspicious situation to be in, you know? Why would you propose the deal instead of just crushing my arguments on the merits in the eyes of third parties? It seems like the kind of thing you would only propose if you didn’t think you could win on the merits.
I wrote more about this in July 2023:
This seems wrong to me.
Yes, Alice may be talking with Bob because she wants Bob to believe something that Alice thinks is true. And Bob may be talking with Alice because he wants Alice to believe something that Bob thinks is true.
Bob may accept the deal in order to have Alice reciprocate and spend compute on seriously pondering whether Bob might be right. Again, this makes sense to Bob even if he just cares about Alice being right, if he has reason to believe she’s currently wrong.
No, it can be about both people wanting the other one to have true beliefs, if they come in with separate beliefs about what’s true.
I think maybe you think this situation shouldn’t be stable because of something about Aumann’s agreement theorem making this persistent disagreement unstable if both Alice and Bob were actually only concerned with the truth. But Aumann’s theorem doesn’t just require mutual rationality but common knowledge of mutual rationality. That’s a high bar that approximately no conversation meets because it’s always plausible that your counter-party is acting in bad-faith. (Or that they think it’s plausible that you act in bad faith, or that they think it’s plausible that you think it’s plausible that they are acting in bad faith, etc.) At which points it seems consistent to me to (i) have mutual knowledge of disagreement, and (ii) assign high probability to your conversation partner being good-faith and rational. [Edit: Well, it’s unlikely that any human is perfect on those aspects. But sufficiently good-faith and rational to be worthwhile to keep talking with in a pretty assuming-good-faith kind of way.]
Also...
This doesn’t at all seem like an easy operation that one participant in the conversation can do with 0 effort from the other participant. Language is not very high band-width. People come into conversations with different abstractions and world models etc. It’s not feasible to dump all the reasons you believe something, because it probably hooks into deep parts of your world model. It’s not sufficient to dump an argument that convinced you because those words were converted to concepts in your brain and may not correspond to the same concepts in someone else’s brain and/or because the argument may have relied on previous beliefs you had which aren’t shared by your conversational partner. If Alice is going to convince Bob of her view, there’s definitely a lot of computational labor involved in that which is more efficiently done by Bob than by Alice.
Yes, and those plausible worlds are exactly why I’m so suspicious of the proposed compute reallocation deal: I’m worried that “you’re being uncharitable; try to think harder about why I’m right” is something people around here say to conceal the fact that they don’t actually have an argument.
Maybe you’re right that the deal could make sense (and my argument that it doesn’t is wrong), but it gets scuttled by the adverse selection problem (where good faith actors would prefer to make such deals with each other, but they have no way to credibly distinguish themselves from bad faith actors who are just bluffing)?
Well sometimes it gets scuttled and sometimes it doesn’t. Cars do in fact get sold second-hand, despite the adverse selection problem. Often the gains-from-trade are large enough to justify some risk.
I think there are often big gains from people putting in some work to understand where their conversation partner is coming from.
Also, “no way to credibly distinguish themselves” seems like a big overstatement. You can do way better than chance at telling whether someone is reasonable or not, especially in repeated interactions:
For one, if someone says “you’re being uncharitable; try to think harder about why I’m right”—you can see that as a reputational bet that you’ll learn something interesting from thinking harder about why they’re right. If you try that a couple of times and you don’t learn anything new, then that’s reasonable inductive evidence that you won’t benefit from that exercise with that person again. You’re not infinitely exploitable here.
And from the perspective of the “deal”—where you put in some effort to understand their perspective, and they put in some effort on understanding yours, even if you don’t expect your part of the deal to be worthwhile in isolation—I think it’s pretty easy to get evidence for whether they’ve held up their part of the deal. Like: Can they demonstrate any understanding of what you’re trying to say? (In the limit this is “can they pass your ITT” but conversations often involve much weaker versions, like “so you’re saying [blah]”, or even less explicit demonstrations of understanding, like engaging with what the other person see as the central point.)
I like this line of inquiry, thank you. If I were working in a Bayesian epistemology I expect I’d find this argument fairly persuasive.
So I think the core reasons I disagree are related to my post on why I’m not a bayesian. That is: I think of intellectual progress in terms of growing a new ontology/developing a new model of the world. At first, this ontology will be fairly illegible to other people. Gradually, I’ll find ways to get it to generate novel falsifiable predictions, or figure out how it overlaps with other peoples’ ontologies. The more we can find overlap, the more productively we can communicate insights to each other. But the procedure of identifying the correspondence between your ontology and my ontology often requires a bunch of work. Insofar as we’re both doing it, we can “build the bridge from both ends” to find the overlapping sections and then communicate insights about those. This is much harder if one person is doing far more of the interpretative labor.
Do you have an example of intellectual progress that fits this model? “growing a new ontology/developing a new model of the world” makes sense to me, but I guess the way I’ve been doing it is to develop it mostly in private (maybe sharing it as I go but, from personal experience, I don’t expect anyone to have an interest in it, unless they independently happen to have the same intuitions / research tastes as me), until I develop some kind of minimum viable product (e.g. UDT 1.0, or the Bitcoin whitepaper as an example from someone else) based on the new ontology/model that is legibly good (e.g., clearly solves some open problems, seems clearly insightful to other people without them having to do a lot of intellectual/interpretive work, etc.).
To share more of my model, I think most attempts to make intellectual progress are likely to fail, e.g. are not aimed in a right direction. So you may think you’re growing a new ontology/developing a new model of the world, but most likely you’re just developing a bunch of nonsense but you can’t see that yet (e.g. due to lack of compute to find a contradiction in your ideas). So then if two such people try “build the bridge from both ends” that’s just very likely to not result in something useful. (I.e., it requires the two people to both be on the right track, which is multiplicatively unlikely.)
I wonder if one issue with people’s models of intellectual progress is that ~nobody on LW saw the ~10 years I spent exploring various dead ends / unclear results in anthropic reasoning and decision theory[1], they just saw UDT 1.0 pop almost out of nowhere. (Perhaps some saw the few months of decision theory discussion on LW before my post, where I got some of the final pieces of the puzzle from Eliezer and Nesov’s hints.)
As for Bitcoin, I don’t know how long it took Satoshi, but there were hundreds or thousands of published papers on digital cash designs of all kinds[2], before Bitcoin came on the scene and took off, but ~everyone just saw Bitcoin and don’t know the rest.
Sorry if this is being unfair/uncharitable to you[3], @Richard_Ngo, but I suspect this might be what’s accounting for our apparent disagreements around intellectual progress.
I posted half-baked ideas to my own everything-list mailing list, but didn’t get much attention from others, except Hal Finney who liked and wrote up UDASSA on his website, which somehow spread from there even though I wasn’t all that happy with it.
My boss during my internship at Microsoft Research asked me to work on his own design, but I let it quietly fall off my plate because it seemed unpromising, and I had my own ideas at that point.
Maybe you’ve observed similar phenomenon elsewhere and still reached different conclusions than me, or I’m completely misunderstanding where our disagreement lies.
Here, I’ll work under the assumption that we need (at least) some enforcement of the excluded middle, since, eventually, we must use our thinking to take some precise action to the exclusion of all others. This action can always be specified as a finite length binary string. If you want, you can construct a class of “fair” scenarios where the agent will not get into trouble if it processes the binary string corresponding to the action it just took.
As to my intentions, I want to discuss requiring self-trust here WOLOG, which is slightly hard to do, but essentially think something like “in cases where we are self-trust
fair, our agent will meet self-trust desiderata (up to practicality) on its future belief on its (binary) decision string IF (classical logicimplies, locally no claim made on other parts of the truth table) the environment is string-processingfairAND it epsilon-prefers so inspecting.”(WOLOG for the desiderata, the test situations admit a small epsilon > 0 (for the mix-in multiplier) that leaves other behavior the same.)
(With much informality) we want, for our decisions, something like what is called “LUV Coherence.” Or, potentially, we want something like “that our eventual decision bit-strings are chosen in a way that (robustly to adversarial attack) closely approximates some decision that is made on the basis of a set of dynamic variables (LUVs) that is entirely coherent.”
(Somewhat unnecessarily for certain audiences, we want to state that our other performance desiderata is left out here.)
Effectively, all of this so far lets us work backwards from the initial excluded middle requirement, through (potentially off-policy/off the collective policy) self trust, then through (decision-approximated) coherence of decision-relevant LUVs.
(As an aside, it is often difficult to tell information that “might” be useful from information that is “teleologically” useful for the actual decisions or (desired) computational processes that actually occur. This can be solved on a theoretical level by requiring that our agent has a fixed core design, but is able to do as well as possible in a large number of scenarios. During design, we will be careful to avoid overfitting to test environments, but we will still prune all computation that can be proven to not be worth the cost.)
Once we’ve established the state of the interface between the agent’s epistemics and its actions, we can examine what constraints that puts on the agent’s internals. Notably, this includes two desiderata, to my knowledge novel statements of the problem:
First new desiderata: coherence at the interface must be efficiently supported by the interior, i.e. it must not be computationally wasteful and not put at risk of expulsion anything that is more useful to the final, excluded-middle-enforced action.
Second new desiderata: the internals must prioritize atomicity, separability, and robustness over exact correctness or any aesthetic sense of “unification” in order to operate under the ruthless and short-tempered expulsion procedures required by computational limitation.
(Note that expulsion is often by greedy and immediate matching on syntactic patterns, and even this approach may be too expensive, depending on the ultimate limits of technology.)
These new desiderata (not entirely formally) motivate an approach that uses discardable “traders” that have a limited impact on near-future agent optimality if incorrectly thrown out, and where any single expulsion does not affect optimality in the limit.
Why the agents should trade on statements of basic logic: computers are unable to operate on anything “richer,” except in the way e.g. Metamath implements further math, by assuming further axioms. A reason you might do this anyway: there is a practical improvement in operation speed, even accounting for the abstraction cost, without damaging (adversarial) robustness. Why we can simplify this away (at least for agents running on million-year or less time spans): by the demand for coherence at the decision level, atomicity, separability, and robustness, there is a limit to how much you can deviate from logic (while maintaining performance, desiderata left out here). This should be provable by ruling out obviating alternatives to syntactic enforcement (supporting and being supported by atomicity, separability, and robustness requirements) and showing adversarial attacks against improper further assumptions (re. robustness).
Presently, I will simplify down to the case where we have some (modified) LI traders used according to the overall context of this comment. Tell me if this step is too bold. Also, note that this account is actually simplified vs. what would work for a real decision theory, for the sake of ease of reasoning and to avoid following unpublished accounts.
In an attempt to describe what you propose, while meeting the desiderata, imagine we have a pair of traders, } is not. Stipulate the set is e.c. when } is e.c. Since, as can clearly be seen, this situation is not symmetrical (
tr1andtr2. (Note that sharing can be chained as long as there are no loops at the subroutine level, but we will stick to a pair here WOLOG.) Let’s say we have a definition of “efficiently computable” (e.c.) that goes false when an individual trader takes on too much work per “day,” and also goes false on a collective level if there is too much compute demanded overall. Saytr1is e.c., but {tr1,tr2, …,trtr2is removed. Let’s say we try to transformtr2(not e.c.) intotr2prime (e.c. when allowed to referencetr1), by lettingtr2prime reference some subroutine oftr1, and letting the cost reduction be accepted in the standard way when determining e.c. in both the individual (when allowed to reference) and collective definitions. Say that the set {tr1,tr2prime, …,trtr1can survive alone, buttr2can not), a well calibrated model with holdouts would say thattr2(in the operational form oftr2prime) is “less probable” to survive thantr1. (Holdout method not given, compute-limited modeling roughly following the standard for LIs.) Note that to the extent that any of these “traders” have “inner” instrumental rationality, they are unable to take any “action” other than outputting a trading strategy for the “day,” for fairly standard betting market incentive reasons. For this reason, we assume the set of “traders” is brought into existence by some outer process in their operational forms, with any references already established. Eviction must be extremely efficient. I am unsure if this prohibits re-parenting of a subroutine, but for simplicity assume that the operational set wouldn’t allowtr2to survive even if it was allowed. Deviating a bit from formality, generalizetr1totrxandtr2prime totry, both in the domain of situations of this sort. Define the function ev_p(tr) to be the bounded-compute “probability” of the eviction oftr(see before on modeling). Evaluate ev_p(try) - ev_p(trx). Call this “the cost of ontological overlap.”As said before, even if a “trader” has a decision theory set up internally, it won’t be hooked up to the interface between the “trader” and the (actual) wrapping hardware in any traditional sense. However, let’s look at a case where we imagine some alternative to
tr2that could have done better thantr2, without having any ontological overlap, while being e.c. Call ittrv. Notice that it has, at the absolute least, regained the entire (non-formally made specific again) “cost of ontological overlap.”Surely, there can be savings by sharing the execution of subroutines, but this trades off against atomicity and separability. Depending on the eventual costs of things like re-parenting, and more advanced techniques like the automatic determination of acceptable approximations and the determination of “equivalent” functions, this may trade off substantially against robustness as well, since the eviction of a single “trader” will sometimes result in a massive over-eviction, if the dependencies can’t be fixed up in an extremely rapid manner (against optimality as well, but we omit that here).
This attempts to model “communication” at some level (actually quite well, given some theories of communication and collective rationality I can vaguely remember), but note that, treating it as a sort of agent it usually won’t be,
trvhas no “reason” to want to be replaced withtr2prime. With some basic assumptions on the setup, e.g. strictly positive initial trader budgets,trvbeing replaced withtr2prime is a strict decrease in estimated survival “probability.” Therefore, this “communication”/sharing must only be done to the extent that it really helps the overall agent, and not following some procedure or “virtue” that is thought to be worth universally adopting.(We’re assuming the set with
trvreplacingtr2prime is also e.c. Note that in this entire comment we’re assumingtr2andtr2prime are functionally identical, to the level of detail we care about here.)Okay, then let’s try to go back to English. Let’s think about these “traders” as cartoons, in the same sense that a cartooned evolution can speak and (be said to) want things. So this trader says, “I want to share load onto other models run by other traders if and only if it reduces my compute demands enough to achieve viability. Sharing ontology for other reasons just makes me worse off, in expectation. Depending on other ontologies is a risk that must be taken on only after sufficient calculation, since ontologies (via the traders that operate them) may go bust.”
(In a more realistic system, it would need to be considered what to do when a “trader” is “exporting” a subroutine that it doesn’t want to use any more. For simplicity, you can ignore this for now. Note that the cartoon dialogue doesn’t really survive this, but it’s cartoon dialogue for a reason. Note that “going bust” should be read as referring to “traders” getting evicted, in ways that don’t strictly relate to bankroll.)
I can’t think of an alternative to this setup, and since the (new) intermediate desiderata seem well motivated, I’m hesitantly taking this as a (sketched, not fully filled in) disproof. One reason for my hesitancy is that my model of proper thinking may require too much compute and too much implementation complexity to ever be usable. Sure, maybe I can fill in all the details and get a system that’s proven aligned, but it might not be applicable to humans.
Do you know of a technical or semi-technical note on how this ontology framework (as an argument for sharing load) functions here/meets desiderata? It doesn’t need to be an explainer, so don’t particularly worry about the quality or any missing details.
I’m not an expert in this area, so maybe the solution really is simple and I’m just not seeing it, or I’m making a silly mistake.
(My hand written heuristics claim this comment is dense to the point of near-unreadability, so all (prospective) readers are invited to ask questions, including if a reading attempt has yet to be made.)
The step of running with the alien assumptions and framings is also very important, it’s unnecessary and impractical to only focus on the things you believe or understand or endorse. There is much more to learn than what’s ready to be made into a part of yourself, and the only way to defeat path-dependence at human level is by being anti-inductive with respect to your own perspective, to seek out the points of view you don’t understand or endorse or believe (on some topic of interest), and at least come to understand how they work internally for their proponents, even when you don’t have footholds of belief or endorsement into them that promise they might be a good fit down the line. (This does require solid sandboxing practices, or else your mind might get so open the brains fall out.) Sometimes such footholds are unexpected, and only appear after you do the work, despite not having any footholds at the outset, and so you make your own epistemic luck.
I didn’t think this through carefully, but maybe an epistemic PD where the currency is compute? Like, I spend 5 utils worth of compute; from your perspective that’s worth 10 expected utils because I might correct my false (according to you) belief; and you do the same for me.
In principle, someone else could produce the corrected positive vision. It’s not obviously crucial that it’s specifically the author of the flawed inspiration who also does the subsequent refinement. Incentives that shape the original contributions are probably a more important aspect of this than the possible later causal back-and-forth. Criticism being unpleasant is some sort of incentive (to be correct, or to avoid contributing). Criticism being valuable fuel for refinement is also some sort of incentive (to exploit Cunningham’s Law by being wrong so that you can learn the right answer from the critics, or to start contributing even when you have nothing to say). So it’s not straightforward, even though at first glance it seems like normatively it should be.
I’d be really very surprised if you hadn’t before heard the argument about nobody having infinite time, bad signal:noise ratios, etc. In fact I’m confident I’ve made this argument to you sometime in the last ~year. Of course, resolving that one would require digging into object-level details, which hasn’t exactly been a fruitful endeavour in the past, maybe because of different thresholds that various individuals have for what signal:noise ratio they find tolerable, and the contingent facts about reality that permit people with different thresholds from you in particular to successfully contribute to advancing the state of human knowledge, even if they don’t want to deal with random nutpickers who once in a blue moon will point out a meaningful error in their post, while their other 100 comments are wrong, confused by something that almost nobody else is confused by, focusing on some random triviality that isn’t load-bearing for the core argument… etc. (Numbers made up; I am, again, establishing some least convenient possible world so that we can skip to the part where we agree that there’s a spectrum and get to arguing about where on the spectrum we should live.)
Perhaps you think that Richard doesn’t endorse this argument, and so you shouldn’t bother to bring it up as a hypothesis? But you asked about what theory you should be embracing. I propose the above theory: I don’t believe you’ve provided much evidence (that I can recall) that your threshold is, in this way, better for advancing the state of rationality/the frontier of human knowledge/etc.
I do wonder if there’s some way to make this kind of debate something like an order of magnitude more efficient using LLMs. By “this kind” I mean “a big huge sprawling debate across many different posts and comments through years, where one or more interlocutors keep forgetting bits about another interlocutor’s perspective and/or not understanding it well enough to bring the relevant facts & arguments to bear in a progressful context for thinking it through”. The help I’m imagining is something like “scrape a big database of all posts and all comments involved in the debate; then have one or more participants talk to an LLM with things like....
how would the other side respond to this; use quotations from their past writing
arguing against the other side with LLM as devil’s advocate
doing ITTs
asking if the other side has addressed X
asking if the other side seems to understand X, care about X, etc.
etc.
....until they reach quasi-equilibrium with something they actually want to put directly to another person. Maybe more efficient because it reduces redundant labor on both sides, and in particular helps with the annoying “routing” task where two clashing perspectives keep having to figure out which standard paragraph from its own perspective’s library of paragraphs is relevant to whatever the other person is confused about. (And maybe psychically easier to argue with a neutral LLM where you don’t have to be defending yourself, upholding boundaries, prosecuting a conflict, etc.)
Indeed, it was only 84 days ago. I’m listening if you have a response to my last comment in that thread, in which I agreed that censorship to maintain the signal-to-noise threshold is good. I then linked to a compilation of Achmiz’s favorites of his own comments and offered the judgments of myself and Jessica Taylor that Achmiz’s work is well above the threshold of being worthy of published on Less Wrong.
Now, maybe you think Taylor (former MIRI employee, inventor of quantilizers and co-author of the logical induction paper, and 7x Curated author) and I (9x Curated, 4x Best of Less Wrong, 3x Less Online invited author) have terrible judgment about what advances the state of rationality/the frontier of human knowledge/&c. Is that your position? Happy to dig into the details if you want.
This confuses me. Can you explain what you sense as a possible motivation on my part to attack “intellectual progress is in fact a valuable thing which we can aim for”?
My own interpretation of what the conflict is, is that everyone here wants to support intellectual progress, but have different ideas how to go about it due to being biased due to status seeking, sunken costs, etc., some of which is strong/ingrained enough to make it intractable to fix the underlying mistakes, so we can only fight it out as a conflict. (To be clear I’m not certain about this.)
Does this still make sense given what I wrote above?
To be clear, I don’t see myself as being in a conflict with you, but with the site mods and the LW cultural trajectory that they’re guiding. Ironically it was initially set on this path when Eliezer demanded author mod powers as a condition of coming back to LW (from FB/Twitter), but now it’s been so ingrained through past decisions that it seems impossible for the mods to admit this was a mistake (via ordinary deliberation), hence why conflict theory feels more appropriate at this point.
I take the side of the site mods on this issue, though, for the reasons articulated in my comment above.
I would also bid for you to do the thing that I recommended Zack do above, namely “acknowledging the thing that Habyrka and I are trying to protect, and helping us figure out how to protect it with as few tradeoffs as possible”.
I think mistake theory makes more sense for the disagreement between us. Because you don’t have nearly as much “sunk cost” invested in the current culture (e.g. was not responsible for cultivating it in the first place), I infer that you can change your mind much more easily about what kind of site culture is more conducive for clear thinking or intellectual progress. (But of course having a mistake theoretic debate with you about this would be pointless given the existence of the main conflict.)
This seems to be ignoring my point about why I think this is a conflict theoretic situation.
You mean that “it was initially set on this path when Eliezer demanded author mod powers”? You could interpret this as Eliezer wielding power in ways you don’t like, but you could also interpret it as the mods treating Eliezer’s preferences as evidence about what site norms produce intellectual progress (along with many other people’s preferences) and acting accordingly. I assume you don’t think the latter is a good model; if so, why not?
Two responses. Firstly, I acknowledge that you did a lot to cultivate the site culture in the first place, and I’m grateful for that.
Secondly: one of the most difficult parts of rationality seems to be changing one’s mind in the face of sunk costs. Because of that, I agree that it’s easier for me to change my mind on this than it is for you. But do you endorse the extent to which sunk costs are making it harder for you to change your mind?
I remember it as being presented as a fait accompli: Eliezer and other unnamed authors demanded this, and we already implemented it. “We prioritized building the delete-and-hide function because Eliezer asked for it and we wanted to get him posting again quickly. But he is not the only author to have asked and expressed appreciation for it.”
See also this comment where Raemon guesses at Eliezer’s reasons for his demand, suggesting that they did not consider Eliezer’s reasons to be centrally important (i.e. didn’t even ask Eliezer to write them up), as opposed to the demand itself: “I’m not that confident in the following, and I don’t want this to turn into a psychoanalyze Eliezer subthread and will lock it if it appears to do that” (The “will lock it” further suggests that the overall decision is not up for discussion, because otherwise understanding/debating Eliezer’s reasons would seem to be central.)
Or even simpler, I don’t remember them ever asking for my preferences about author mod powers.
EDIT: It looks like you added more after I wrote my reply, which is just as well as I feel like falling into a familiar trap by engaging in this debate with you.
The principle of more dakka is needed because people often don’t do more of the thing that is working, so I don’t agree that this obviously follows. Perhaps more relevantly, Wei Dai never linked to any of those posts in either of his comments on TsviBT’s or Ngo’s posts, and I had forgotten about them.
(But also I want to disclaim that I am not confident in this read of the situation, it just seemed to me “about on the level of being worth considering” as the other things Wei listed.)
I don’t have a job and participating on LW has been my main hobby, so “just do more of it” is clearly not a general/scalable solution.
My goal (which I endorse on reflection) is to write comments/posts that a lot of people get value out of (which I think I had succeeded with the comments in question), not convince or interest every person, who may be motivatedly obtuse to its relevance or importance.
In this case I thought I was making a fairly surgical correction (that MIRI/early rationalists were not some sort of beacon of clear thinking in contrast with EA/other institutions, but also made plenty of high-impact mistakes) that didn’t depend on my previous writings. (Adding links potentially implies that reading the linked content is necessary to understand the current content.)
Writing potentially half-baked comments is how I (and I imagine others) find out whether some idea or line of thought is worth putting more effort into, which parts are hard for others to understand, etc., but which now risks getting me labeled (even more) as something that LW culture seems to think is bad (given that the relevant post/accusations all had fairly high karma), and potentially getting author-banned.
I generally agree with Zack that at this point this disagreement is probably more about conflict theory than mistake theory, despite how the above points may read (i.e., it may seem like I’m adopting more of a mistake theory view where I’m trying to explain reasoning that maybe just hasn’t occurred to you). (In part I’m just more used to operating under mistake theory framing, and in part I’m not fully certain about this.)
(I believe those 4 linked posts were written after the thread in question on my post, though of course before the thread on Richard’s post.)
A few quick thoughts:
I have often seen it be the case that, during a conflict between rationalists (a shorthand referring to people active in the rationalist-scene), someone offers a mistake theory interpretation for what led up to the conflict, and the other party finds that to be a valid account of a mistake that they hadn’t noticed, and then all the air is removed from the conflict.
I think this move isn’t always a misdirect. Right now, to give an overly-specific hypothesis, I think it seems like a live hypothesis that part of the reason it was highly upvoted was because it was part of Wei’s ongoing critique of Eliezer/MIRI, and partly that seems upvoted because it’s drama that people are into, rather than being related to their post, and this is kind of distracting from the topics the authors are working to discuss.
I am not accusing Wei Dai of this! I am just saying that, if the authors are finding the critique not very good / non-central—which could well be because on some level they just don’t want to hear it—then this could seem like a more plausible hypothesis to them. And so it would be effective to make the relevance clearer to resolve this issue.
I think your response here though is “But overall the comment shouldn’t solely or even primarily be evaluated on whether it was good by the author’s lights, it should be evaluated by whether it’s helping the discourse and ideas move forward” and I agree, often comments that the author finds annoying and stupid, are in fact good, and valued by the many readers.
(To expand on that: I agree that it is somewhat the case that authors should learn better to deal with the top-voted comment on their post being non-central, or a critique that they don’t value very much, because much of the point is for the readers and the general discourse, not the authors, and I think it’s plausible that this is the case here. I currently don’t see cause for either author to ban Wei based on these two comments, and I thought it was an overreaction by Tsvi at the time.)
But my response is that the author can still sometimes get it right? I don’t think that the author’s propensity for self-deception means that they can never come to a justified true belief that the commenter is norm-violating (e.g. hobby-horsing) and should be banned for it, and take action based on that. And when I look at the list of bans on the /moderation page, personally I judge that most of them seem pretty reasonable and understandable based on annoying behaviors and not primarily the basis of (functionally) seeking to avoid good criticism of their ideas. (There are of course some exceptions that I disagree with, and there are certainly rates of unreasonable ones, or rates of egregiously bad ones, that would make me want to take action to counter it.)