I think your implicit model of “trying to see things the other guy’s way” is incorrect. In particular, it’s not “adopt one of your beliefs on request”, but more like “spin up a mental sandbox which explores the implications of your belief being true”.
And so the exchange that I’m actually proposing is more like “I will spend compute on trying to figure out ways that your perspective might be more consistent and well-intentioned than I currently expect, if you do the same for me”.
The main way that this might create false maps is if the sandboxes are leaky, or if your perspective has been adversarially optimized to lead me to false conclusions. When I talk about trusting someone, one of the things I mean is “trying to simulate their perspective won’t mess my epistemics up”. E.g. you shouldn’t trust or try to mentally explore the implications of claims that a misaligned superintelligence has given you.
the idea seems to be that if people’s attempts to formulate positive visions get critiqued too vigorously, they’ll get discouraged and give up.
The implied psychological model of Less Wrong authors reminds me of my attempts to teach chess to my five-year-old niece this week
I tried out several possible angles of response, but upon reflection, I’m no longer interested in debate with you on this topic (though I might still engage if you have responses to the rest of this comment).
In particular, it’s not “adopt one of your beliefs on request”, but more like “spin up a mental sandbox which explores the implications of your belief being true”
Sure; I regret the rhetorical excess. (I think it would actually be “spin up a more powerful search for reasons your belief might be true”, not a search for its implications if true.)
And so the exchange that I’m actually proposing is more like “I will spend compute on trying to figure out ways that your perspective might be more consistent and well-intentioned than I currently expect, if you do the same for me”.
The question is, why would you propose that deal? What are the circumstances under which proposing that deal would seem like a good idea to someone?
The reason I’m asking is because if we assume that you’re truthseeking, it’s kind of a weird thing to propose, right? If you think that you’re right and that my position is inconsistent and ill-intentioned, presumably that’s because you’ve already thought about the matter and this is where you ended up. How would it improve the accuracy of your beliefs to reallocate your inference compute if I do so, too? Whether a proposed compute reallocation improves your beliefs shouldn’t depend on what I’m doing with my compute!
That is, the deal only makes sense for you to propose if you have a stake in something other than the accuracy of your own beliefs—if there’s something it would benefit you for me to believe. But then if I’m truthseeking, then there’s no reason for me to accept the deal (because whether a proposed compute reallocation improves my beliefs shouldn’t depend on what you’re doing with your compute). It seems like the deal should only go through if we both have a stake in the other person’s beliefs that isn’t about the truth or falsity of those beliefs, and there’s some reason you can’t just dump the reasoning that convinced you. That’s a weird and suspicious situation to be in, you know? Why would you propose the deal instead of just crushing my arguments on the merits in the eyes of third parties? It seems like the kind of thing you would only propose if you didn’t think you could win on the merits.
[W]hen someone who is currently trying to persuade me of something tells me that it doesn’t look I’m making enough effort to think of reasons why they’re right, that immediately makes me think they’re more likely to be wrong. Why? Because I think that if they had an argument, they would be telling me the argument, not chastising my lack of charity. The advice to be on special lookout for reasons your interlocutor is right is good in general, but your interlocutor is the last person to be trusted to give it, because [...] they have an ulterior motive.
That is, the deal only makes sense for you to propose if you have a stake in something other than the accuracy of your own beliefs—if there’s something it would benefit you for me to believe.
Yes, Alice may be talking with Bob because she wants Bob to believe something that Alice thinks is true. And Bob may be talking with Alice because he wants Alice to believe something that Bob thinks is true.
But then if I’m truthseeking, then there’s no reason for me to accept the deal (because whether a proposed compute reallocation improves my beliefs shouldn’t depend on what you’re doing with your compute).
Bob may accept the deal in order to have Alice reciprocate and spend compute on seriously pondering whether Bob might be right. Again, this makes sense to Bob even if he just cares about Alice being right, if he has reason to believe she’s currently wrong.
It seems like the deal should only go through if we both have a stake in the other person’s beliefs that isn’t about the truth or falsity of those beliefs
No, it can be about both people wanting the other one to have true beliefs, if they come in with separate beliefs about what’s true.
I think maybe you think this situation shouldn’t be stable because of something about Aumann’s agreement theorem making this persistent disagreement unstable if both Alice and Bob were actually only concerned with the truth. But Aumann’s theorem doesn’t just require mutual rationality but common knowledge of mutual rationality. That’s a high bar that approximately no conversation meets because it’s always plausible that your counter-party is acting in bad-faith. (Or that they think it’s plausible that you act in bad faith, or that they think it’s plausible that you think it’s plausible that they are acting in bad faith, etc.) At which points it seems consistent to me to (i) have mutual knowledge of disagreement, and (ii) assign high probability to your conversation partner being good-faith and rational. [Edit: Well, it’s unlikely that any human is perfect on those aspects. But sufficiently good-faith and rational to be worthwhile to keep talking with in a pretty assuming-good-faith kind of way.]
This doesn’t at all seem like an easy operation that one participant in the conversation can do with 0 effort from the other participant. Language is not very high band-width. People come into conversations with different abstractions and world models etc. It’s not feasible to dump all the reasons you believe something, because it probably hooks into deep parts of your world model. It’s not sufficient to dump an argument that convinced you because those words were converted to concepts in your brain and may not correspond to the same concepts in someone else’s brain and/or because the argument may have relied on previous beliefs you had which aren’t shared by your conversational partner. If Alice is going to convince Bob of her view, there’s definitely a lot of computational labor involved in that which is more efficiently done by Bob than by Alice.
That’s a high bar that approximately no conversation meets because it’s always plausible that your counter-party is acting in bad-faith.
Yes, and those plausible worlds are exactly why I’m so suspicious of the proposed compute reallocation deal: I’m worried that “you’re being uncharitable; try to think harder about why I’m right” is something people around here say to conceal the fact that they don’t actually have an argument.
Maybe you’re right that the deal could make sense (and my argument that it doesn’t is wrong), but it gets scuttled by the adverse selection problem (where good faith actors would prefer to make such deals with each other, but they have no way to credibly distinguish themselves from bad faith actors who are just bluffing)?
Well sometimes it gets scuttled and sometimes it doesn’t. Cars do in fact get sold second-hand, despite the adverse selection problem. Often the gains-from-trade are large enough to justify some risk.
I think there are often big gains from people putting in some work to understand where their conversation partner is coming from.
Also, “no way to credibly distinguish themselves” seems like a big overstatement. You can do way better than chance at telling whether someone is reasonable or not, especially in repeated interactions:
For one, if someone says “you’re being uncharitable; try to think harder about why I’m right”—you can see that as a reputational bet that you’ll learn something interesting from thinking harder about why they’re right. If you try that a couple of times and you don’t learn anything new, then that’s reasonable inductive evidence that you won’t benefit from that exercise with that person again. You’re not infinitely exploitable here.
And from the perspective of the “deal”—where you put in some effort to understand their perspective, and they put in some effort on understanding yours, even if you don’t expect your part of the deal to be worthwhile in isolation—I think it’s pretty easy to get evidence for whether they’ve held up their part of the deal. Like: Can they demonstrate any understanding of what you’re trying to say? (In the limit this is “can they pass your ITT” but conversations often involve much weaker versions, like “so you’re saying [blah]”, or even less explicit demonstrations of understanding, like engaging with what the other person see as the central point.)
The reason I’m asking is because if we assume that you’re truthseeking, it’s kind of a weird thing to propose, right?
I like this line of inquiry, thank you. If I were working in a Bayesian epistemology I expect I’d find this argument fairly persuasive.
So I think the core reasons I disagree are related to my post on why I’m not a bayesian. That is: I think of intellectual progress in terms of growing a new ontology/developing a new model of the world. At first, this ontology will be fairly illegible to other people. Gradually, I’ll find ways to get it to generate novel falsifiable predictions, or figure out how it overlaps with other peoples’ ontologies. The more we can find overlap, the more productively we can communicate insights to each other. But the procedure of identifying the correspondence between your ontology and my ontology often requires a bunch of work. Insofar as we’re both doing it, we can “build the bridge from both ends” to find the overlapping sections and then communicate insights about those. This is much harder if one person is doing far more of the interpretative labor.
Do you have an example of intellectual progress that fits this model? “growing a new ontology/developing a new model of the world” makes sense to me, but I guess the way I’ve been doing it is to develop it mostly in private (maybe sharing it as I go but, from personal experience, I don’t expect anyone to have an interest in it, unless they independently happen to have the same intuitions / research tastes as me), until I develop some kind of minimum viable product (e.g. UDT 1.0, or the Bitcoin whitepaper as an example from someone else) based on the new ontology/model that is legibly good (e.g., clearly solves some open problems, seems clearly insightful to other people without them having to do a lot of intellectual/interpretive work, etc.).
To share more of my model, I think most attempts to make intellectual progress are likely to fail, e.g. are not aimed in a right direction. So you may think you’re growing a new ontology/developing a new model of the world, but most likely you’re just developing a bunch of nonsense but you can’t see that yet (e.g. due to lack of compute to find a contradiction in your ideas). So then if two such people try “build the bridge from both ends” that’s just very likely to not result in something useful. (I.e., it requires the two people to both be on the right track, which is multiplicatively unlikely.)
I think most attempts to make intellectual progress are likely to fail
I wonder if one issue with people’s models of intellectual progress is that ~nobody on LW saw the ~10 years I spent exploring various dead ends / unclear results in anthropic reasoning and decision theory[1], they just saw UDT 1.0 pop almost out of nowhere. (Perhaps some saw the few months of decision theory discussion on LW before my post, where I got some of the final pieces of the puzzle from Eliezer and Nesov’s hints.)
As for Bitcoin, I don’t know how long it took Satoshi, but there were hundreds or thousands of published papers on digital cash designs of all kinds[2], before Bitcoin came on the scene and took off, but ~everyone just saw Bitcoin and don’t know the rest.
Sorry if this is being unfair/uncharitable to you[3], @Richard_Ngo, but I suspect this might be what’s accounting for our apparent disagreements around intellectual progress.
I posted half-baked ideas to my own everything-list mailing list, but didn’t get much attention from others, except Hal Finney who liked and wrote up UDASSA on his website, which somehow spread from there even though I wasn’t all that happy with it.
My boss during my internship at Microsoft Research asked me to work on his own design, but I let it quietly fall off my plate because it seemed unpromising, and I had my own ideas at that point.
Maybe you’ve observed similar phenomenon elsewhere and still reached different conclusions than me, or I’m completely misunderstanding where our disagreement lies.
Here, I’ll work under the assumption that we need (at least) some enforcement of the excluded middle, since, eventually, we must use our thinking to take some precise action to the exclusion of all others. This action can always be specified as a finite length binary string. If you want, you can construct a class of “fair” scenarios where the agent will not get into trouble if it processes the binary string corresponding to the action it just took.
As to my intentions, I want to discuss requiring self-trust here WOLOG, which is slightly hard to do, but essentially think something like “in cases where we are self-trust fair, our agent will meet self-trust desiderata (up to practicality) on its future belief on its (binary) decision string IF (classical logic implies, locally no claim made on other parts of the truth table) the environment is string-processing fair AND it epsilon-prefers so inspecting.”
(WOLOG for the desiderata, the test situations admit a small epsilon > 0 (for the mix-in multiplier) that leaves other behavior the same.)
(With much informality) we want, for our decisions, something like what is called “LUV Coherence.” Or, potentially, we want something like “that our eventual decision bit-strings are chosen in a way that (robustly to adversarial attack) closely approximates some decision that is made on the basis of a set of dynamic variables (LUVs) that is entirely coherent.”
(Somewhat unnecessarily for certain audiences, we want to state that our other performance desiderata is left out here.)
Effectively, all of this so far lets us work backwards from the initial excluded middle requirement, through (potentially off-policy/off the collective policy) self trust, then through (decision-approximated) coherence of decision-relevant LUVs.
(As an aside, it is often difficult to tell information that “might” be useful from information that is “teleologically” useful for the actual decisions or (desired) computational processes that actually occur. This can be solved on a theoretical level by requiring that our agent has a fixed core design, but is able to do as well as possible in a large number of scenarios. During design, we will be careful to avoid overfitting to test environments, but we will still prune all computation that can be proven to not be worth the cost.)
Once we’ve established the state of the interface between the agent’s epistemics and its actions, we can examine what constraints that puts on the agent’s internals. Notably, this includes two desiderata, to my knowledge novel statements of the problem:
First new desiderata: coherence at the interface must be efficiently supported by the interior, i.e. it must not be computationally wasteful and not put at risk of expulsion anything that is more useful to the final, excluded-middle-enforced action.
Second new desiderata: the internals must prioritize atomicity, separability, and robustness over exact correctness or any aesthetic sense of “unification” in order to operate under the ruthless and short-tempered expulsion procedures required by computational limitation.
(Note that expulsion is often by greedy and immediate matching on syntactic patterns, and even this approach may be too expensive, depending on the ultimate limits of technology.)
These new desiderata (not entirely formally) motivate an approach that uses discardable “traders” that have a limited impact on near-future agent optimality if incorrectly thrown out, and where any single expulsion does not affect optimality in the limit.
Why the agents should trade on statements of basic logic: computers are unable to operate on anything “richer,” except in the way e.g. Metamath implements further math, by assuming further axioms. A reason you might do this anyway: there is a practical improvement in operation speed, even accounting for the abstraction cost, without damaging (adversarial) robustness. Why we can simplify this away (at least for agents running on million-year or less time spans): by the demand for coherence at the decision level, atomicity, separability, and robustness, there is a limit to how much you can deviate from logic (while maintaining performance, desiderata left out here). This should be provable by ruling out obviating alternatives to syntactic enforcement (supporting and being supported by atomicity, separability, and robustness requirements) and showing adversarial attacks against improper further assumptions (re. robustness).
Presently, I will simplify down to the case where we have some (modified) LI traders used according to the overall context of this comment. Tell me if this step is too bold. Also, note that this account is actually simplified vs. what would work for a real decision theory, for the sake of ease of reasoning and to avoid following unpublished accounts.
In an attempt to describe what you propose, while meeting the desiderata, imagine we have a pair of traders, tr1 and tr2. (Note that sharing can be chained as long as there are no loops at the subroutine level, but we will stick to a pair here WOLOG.) Let’s say we have a definition of “efficiently computable” (e.c.) that goes false when an individual trader takes on too much work per “day,” and also goes false on a collective level if there is too much compute demanded overall. Say tr1 is e.c., but {tr1, tr2, …, tr} is not. Stipulate the set is e.c. when tr2 is removed. Let’s say we try to transform tr2 (not e.c.) into tr2 prime (e.c. when allowed to reference tr1), by letting tr2 prime reference some subroutine of tr1, and letting the cost reduction be accepted in the standard way when determining e.c. in both the individual (when allowed to reference) and collective definitions. Say that the set {tr1, tr2 prime, …, tr} is e.c. Since, as can clearly be seen, this situation is not symmetrical (tr1 can survive alone, but tr2 can not), a well calibrated model with holdouts would say that tr2 (in the operational form of tr2 prime) is “less probable” to survive than tr1. (Holdout method not given, compute-limited modeling roughly following the standard for LIs.) Note that to the extent that any of these “traders” have “inner” instrumental rationality, they are unable to take any “action” other than outputting a trading strategy for the “day,” for fairly standard betting market incentive reasons. For this reason, we assume the set of “traders” is brought into existence by some outer process in their operational forms, with any references already established. Eviction must be extremely efficient. I am unsure if this prohibits re-parenting of a subroutine, but for simplicity assume that the operational set wouldn’t allow tr2 to survive even if it was allowed. Deviating a bit from formality, generalize tr1 to trx and tr2 prime to try, both in the domain of situations of this sort. Define the function ev_p(tr) to be the bounded-compute “probability” of the eviction of tr (see before on modeling). Evaluate ev_p(try) - ev_p(trx). Call this “the cost of ontological overlap.”
As said before, even if a “trader” has a decision theory set up internally, it won’t be hooked up to the interface between the “trader” and the (actual) wrapping hardware in any traditional sense. However, let’s look at a case where we imagine some alternative to tr2 that could have done better than tr2, without having any ontological overlap, while being e.c. Call it trv. Notice that it has, at the absolute least, regained the entire (non-formally made specific again) “cost of ontological overlap.”
Surely, there can be savings by sharing the execution of subroutines, but this trades off against atomicity and separability. Depending on the eventual costs of things like re-parenting, and more advanced techniques like the automatic determination of acceptable approximations and the determination of “equivalent” functions, this may trade off substantially against robustness as well, since the eviction of a single “trader” will sometimes result in a massive over-eviction, if the dependencies can’t be fixed up in an extremely rapid manner (against optimality as well, but we omit that here).
This attempts to model “communication” at some level (actually quite well, given some theories of communication and collective rationality I can vaguely remember), but note that, treating it as a sort of agent it usually won’t be, trv has no “reason” to want to be replaced with tr2 prime. With some basic assumptions on the setup, e.g. strictly positive initial trader budgets, trv being replaced with tr2 prime is a strict decrease in estimated survival “probability.” Therefore, this “communication”/sharing must only be done to the extent that it really helps the overall agent, and not following some procedure or “virtue” that is thought to be worth universally adopting.
(We’re assuming the set with trv replacing tr2 prime is also e.c. Note that in this entire comment we’re assuming tr2 and tr2 prime are functionally identical, to the level of detail we care about here.)
Okay, then let’s try to go back to English. Let’s think about these “traders” as cartoons, in the same sense that a cartooned evolution can speak and (be said to) want things. So this trader says, “I want to share load onto other models run by other traders if and only if it reduces my compute demands enough to achieve viability. Sharing ontology for other reasons just makes me worse off, in expectation. Depending on other ontologies is a risk that must be taken on only after sufficient calculation, since ontologies (via the traders that operate them) may go bust.”
(In a more realistic system, it would need to be considered what to do when a “trader” is “exporting” a subroutine that it doesn’t want to use any more. For simplicity, you can ignore this for now. Note that the cartoon dialogue doesn’t really survive this, but it’s cartoon dialogue for a reason. Note that “going bust” should be read as referring to “traders” getting evicted, in ways that don’t strictly relate to bankroll.)
I can’t think of an alternative to this setup, and since the (new) intermediate desiderata seem well motivated, I’m hesitantly taking this as a (sketched, not fully filled in) disproof. One reason for my hesitancy is that my model of proper thinking may require too much compute and too much implementation complexity to ever be usable. Sure, maybe I can fill in all the details and get a system that’s proven aligned, but it might not be applicable to humans.
Do you know of a technical or semi-technical note on how this ontology framework (as an argument for sharing load) functions here/meets desiderata? It doesn’t need to be an explainer, so don’t particularly worry about the quality or any missing details.
I’m not an expert in this area, so maybe the solution really is simple and I’m just not seeing it, or I’m making a silly mistake.
(My hand written heuristics claim this comment is dense to the point of near-unreadability, so all (prospective) readers are invited to ask questions, including if a reading attempt has yet to be made.)
(I think it would actually be “spin up a more powerful search for reasons your belief might be true”, not a search for its implications if true.)
The step of running with the alien assumptions and framings is also very important, it’s unnecessary and impractical to only focus on the things you believe or understand or endorse. There is much more to learn than what’s ready to be made into a part of yourself, and the only way to defeat path-dependence at human level is by being anti-inductive with respect to your own perspective, to seek out the points of view you don’t understand or endorse or believe (on some topic of interest), and at least come to understand how they work internally for their proponents, even when you don’t have footholds of belief or endorsement into them that promise they might be a good fit down the line. (This does require solid sandboxing practices, or else your mind might get so open the brains fall out.) Sometimes such footholds are unexpected, and only appear after you do the work, despite not having any footholds at the outset, and so you make your own epistemic luck.
That is, the deal only makes sense for you to propose if you have a stake in something other than the accuracy of your own beliefs—if there’s something it would benefit you for me to believe. But then if I’m truthseeking, then there’s no reason for me to accept the deal
I didn’t think this through carefully, but maybe an epistemic PD where the currency is compute? Like, I spend 5 utils worth of compute; from your perspective that’s worth 10 expected utils because I might correct my false (according to you) belief; and you do the same for me.
I think your implicit model of “trying to see things the other guy’s way” is incorrect. In particular, it’s not “adopt one of your beliefs on request”, but more like “spin up a mental sandbox which explores the implications of your belief being true”.
And so the exchange that I’m actually proposing is more like “I will spend compute on trying to figure out ways that your perspective might be more consistent and well-intentioned than I currently expect, if you do the same for me”.
The main way that this might create false maps is if the sandboxes are leaky, or if your perspective has been adversarially optimized to lead me to false conclusions. When I talk about trusting someone, one of the things I mean is “trying to simulate their perspective won’t mess my epistemics up”. E.g. you shouldn’t trust or try to mentally explore the implications of claims that a misaligned superintelligence has given you.
I tried out several possible angles of response, but upon reflection, I’m no longer interested in debate with you on this topic (though I might still engage if you have responses to the rest of this comment).
Sure; I regret the rhetorical excess. (I think it would actually be “spin up a more powerful search for reasons your belief might be true”, not a search for its implications if true.)
The question is, why would you propose that deal? What are the circumstances under which proposing that deal would seem like a good idea to someone?
The reason I’m asking is because if we assume that you’re truthseeking, it’s kind of a weird thing to propose, right? If you think that you’re right and that my position is inconsistent and ill-intentioned, presumably that’s because you’ve already thought about the matter and this is where you ended up. How would it improve the accuracy of your beliefs to reallocate your inference compute if I do so, too? Whether a proposed compute reallocation improves your beliefs shouldn’t depend on what I’m doing with my compute!
That is, the deal only makes sense for you to propose if you have a stake in something other than the accuracy of your own beliefs—if there’s something it would benefit you for me to believe. But then if I’m truthseeking, then there’s no reason for me to accept the deal (because whether a proposed compute reallocation improves my beliefs shouldn’t depend on what you’re doing with your compute). It seems like the deal should only go through if we both have a stake in the other person’s beliefs that isn’t about the truth or falsity of those beliefs, and there’s some reason you can’t just dump the reasoning that convinced you. That’s a weird and suspicious situation to be in, you know? Why would you propose the deal instead of just crushing my arguments on the merits in the eyes of third parties? It seems like the kind of thing you would only propose if you didn’t think you could win on the merits.
I wrote more about this in July 2023:
This seems wrong to me.
Yes, Alice may be talking with Bob because she wants Bob to believe something that Alice thinks is true. And Bob may be talking with Alice because he wants Alice to believe something that Bob thinks is true.
Bob may accept the deal in order to have Alice reciprocate and spend compute on seriously pondering whether Bob might be right. Again, this makes sense to Bob even if he just cares about Alice being right, if he has reason to believe she’s currently wrong.
No, it can be about both people wanting the other one to have true beliefs, if they come in with separate beliefs about what’s true.
I think maybe you think this situation shouldn’t be stable because of something about Aumann’s agreement theorem making this persistent disagreement unstable if both Alice and Bob were actually only concerned with the truth. But Aumann’s theorem doesn’t just require mutual rationality but common knowledge of mutual rationality. That’s a high bar that approximately no conversation meets because it’s always plausible that your counter-party is acting in bad-faith. (Or that they think it’s plausible that you act in bad faith, or that they think it’s plausible that you think it’s plausible that they are acting in bad faith, etc.) At which points it seems consistent to me to (i) have mutual knowledge of disagreement, and (ii) assign high probability to your conversation partner being good-faith and rational. [Edit: Well, it’s unlikely that any human is perfect on those aspects. But sufficiently good-faith and rational to be worthwhile to keep talking with in a pretty assuming-good-faith kind of way.]
Also...
This doesn’t at all seem like an easy operation that one participant in the conversation can do with 0 effort from the other participant. Language is not very high band-width. People come into conversations with different abstractions and world models etc. It’s not feasible to dump all the reasons you believe something, because it probably hooks into deep parts of your world model. It’s not sufficient to dump an argument that convinced you because those words were converted to concepts in your brain and may not correspond to the same concepts in someone else’s brain and/or because the argument may have relied on previous beliefs you had which aren’t shared by your conversational partner. If Alice is going to convince Bob of her view, there’s definitely a lot of computational labor involved in that which is more efficiently done by Bob than by Alice.
Yes, and those plausible worlds are exactly why I’m so suspicious of the proposed compute reallocation deal: I’m worried that “you’re being uncharitable; try to think harder about why I’m right” is something people around here say to conceal the fact that they don’t actually have an argument.
Maybe you’re right that the deal could make sense (and my argument that it doesn’t is wrong), but it gets scuttled by the adverse selection problem (where good faith actors would prefer to make such deals with each other, but they have no way to credibly distinguish themselves from bad faith actors who are just bluffing)?
Well sometimes it gets scuttled and sometimes it doesn’t. Cars do in fact get sold second-hand, despite the adverse selection problem. Often the gains-from-trade are large enough to justify some risk.
I think there are often big gains from people putting in some work to understand where their conversation partner is coming from.
Also, “no way to credibly distinguish themselves” seems like a big overstatement. You can do way better than chance at telling whether someone is reasonable or not, especially in repeated interactions:
For one, if someone says “you’re being uncharitable; try to think harder about why I’m right”—you can see that as a reputational bet that you’ll learn something interesting from thinking harder about why they’re right. If you try that a couple of times and you don’t learn anything new, then that’s reasonable inductive evidence that you won’t benefit from that exercise with that person again. You’re not infinitely exploitable here.
And from the perspective of the “deal”—where you put in some effort to understand their perspective, and they put in some effort on understanding yours, even if you don’t expect your part of the deal to be worthwhile in isolation—I think it’s pretty easy to get evidence for whether they’ve held up their part of the deal. Like: Can they demonstrate any understanding of what you’re trying to say? (In the limit this is “can they pass your ITT” but conversations often involve much weaker versions, like “so you’re saying [blah]”, or even less explicit demonstrations of understanding, like engaging with what the other person see as the central point.)
I like this line of inquiry, thank you. If I were working in a Bayesian epistemology I expect I’d find this argument fairly persuasive.
So I think the core reasons I disagree are related to my post on why I’m not a bayesian. That is: I think of intellectual progress in terms of growing a new ontology/developing a new model of the world. At first, this ontology will be fairly illegible to other people. Gradually, I’ll find ways to get it to generate novel falsifiable predictions, or figure out how it overlaps with other peoples’ ontologies. The more we can find overlap, the more productively we can communicate insights to each other. But the procedure of identifying the correspondence between your ontology and my ontology often requires a bunch of work. Insofar as we’re both doing it, we can “build the bridge from both ends” to find the overlapping sections and then communicate insights about those. This is much harder if one person is doing far more of the interpretative labor.
Do you have an example of intellectual progress that fits this model? “growing a new ontology/developing a new model of the world” makes sense to me, but I guess the way I’ve been doing it is to develop it mostly in private (maybe sharing it as I go but, from personal experience, I don’t expect anyone to have an interest in it, unless they independently happen to have the same intuitions / research tastes as me), until I develop some kind of minimum viable product (e.g. UDT 1.0, or the Bitcoin whitepaper as an example from someone else) based on the new ontology/model that is legibly good (e.g., clearly solves some open problems, seems clearly insightful to other people without them having to do a lot of intellectual/interpretive work, etc.).
To share more of my model, I think most attempts to make intellectual progress are likely to fail, e.g. are not aimed in a right direction. So you may think you’re growing a new ontology/developing a new model of the world, but most likely you’re just developing a bunch of nonsense but you can’t see that yet (e.g. due to lack of compute to find a contradiction in your ideas). So then if two such people try “build the bridge from both ends” that’s just very likely to not result in something useful. (I.e., it requires the two people to both be on the right track, which is multiplicatively unlikely.)
I wonder if one issue with people’s models of intellectual progress is that ~nobody on LW saw the ~10 years I spent exploring various dead ends / unclear results in anthropic reasoning and decision theory[1], they just saw UDT 1.0 pop almost out of nowhere. (Perhaps some saw the few months of decision theory discussion on LW before my post, where I got some of the final pieces of the puzzle from Eliezer and Nesov’s hints.)
As for Bitcoin, I don’t know how long it took Satoshi, but there were hundreds or thousands of published papers on digital cash designs of all kinds[2], before Bitcoin came on the scene and took off, but ~everyone just saw Bitcoin and don’t know the rest.
Sorry if this is being unfair/uncharitable to you[3], @Richard_Ngo, but I suspect this might be what’s accounting for our apparent disagreements around intellectual progress.
I posted half-baked ideas to my own everything-list mailing list, but didn’t get much attention from others, except Hal Finney who liked and wrote up UDASSA on his website, which somehow spread from there even though I wasn’t all that happy with it.
My boss during my internship at Microsoft Research asked me to work on his own design, but I let it quietly fall off my plate because it seemed unpromising, and I had my own ideas at that point.
Maybe you’ve observed similar phenomenon elsewhere and still reached different conclusions than me, or I’m completely misunderstanding where our disagreement lies.
Here, I’ll work under the assumption that we need (at least) some enforcement of the excluded middle, since, eventually, we must use our thinking to take some precise action to the exclusion of all others. This action can always be specified as a finite length binary string. If you want, you can construct a class of “fair” scenarios where the agent will not get into trouble if it processes the binary string corresponding to the action it just took.
As to my intentions, I want to discuss requiring self-trust here WOLOG, which is slightly hard to do, but essentially think something like “in cases where we are self-trust
fair, our agent will meet self-trust desiderata (up to practicality) on its future belief on its (binary) decision string IF (classical logicimplies, locally no claim made on other parts of the truth table) the environment is string-processingfairAND it epsilon-prefers so inspecting.”(WOLOG for the desiderata, the test situations admit a small epsilon > 0 (for the mix-in multiplier) that leaves other behavior the same.)
(With much informality) we want, for our decisions, something like what is called “LUV Coherence.” Or, potentially, we want something like “that our eventual decision bit-strings are chosen in a way that (robustly to adversarial attack) closely approximates some decision that is made on the basis of a set of dynamic variables (LUVs) that is entirely coherent.”
(Somewhat unnecessarily for certain audiences, we want to state that our other performance desiderata is left out here.)
Effectively, all of this so far lets us work backwards from the initial excluded middle requirement, through (potentially off-policy/off the collective policy) self trust, then through (decision-approximated) coherence of decision-relevant LUVs.
(As an aside, it is often difficult to tell information that “might” be useful from information that is “teleologically” useful for the actual decisions or (desired) computational processes that actually occur. This can be solved on a theoretical level by requiring that our agent has a fixed core design, but is able to do as well as possible in a large number of scenarios. During design, we will be careful to avoid overfitting to test environments, but we will still prune all computation that can be proven to not be worth the cost.)
Once we’ve established the state of the interface between the agent’s epistemics and its actions, we can examine what constraints that puts on the agent’s internals. Notably, this includes two desiderata, to my knowledge novel statements of the problem:
First new desiderata: coherence at the interface must be efficiently supported by the interior, i.e. it must not be computationally wasteful and not put at risk of expulsion anything that is more useful to the final, excluded-middle-enforced action.
Second new desiderata: the internals must prioritize atomicity, separability, and robustness over exact correctness or any aesthetic sense of “unification” in order to operate under the ruthless and short-tempered expulsion procedures required by computational limitation.
(Note that expulsion is often by greedy and immediate matching on syntactic patterns, and even this approach may be too expensive, depending on the ultimate limits of technology.)
These new desiderata (not entirely formally) motivate an approach that uses discardable “traders” that have a limited impact on near-future agent optimality if incorrectly thrown out, and where any single expulsion does not affect optimality in the limit.
Why the agents should trade on statements of basic logic: computers are unable to operate on anything “richer,” except in the way e.g. Metamath implements further math, by assuming further axioms. A reason you might do this anyway: there is a practical improvement in operation speed, even accounting for the abstraction cost, without damaging (adversarial) robustness. Why we can simplify this away (at least for agents running on million-year or less time spans): by the demand for coherence at the decision level, atomicity, separability, and robustness, there is a limit to how much you can deviate from logic (while maintaining performance, desiderata left out here). This should be provable by ruling out obviating alternatives to syntactic enforcement (supporting and being supported by atomicity, separability, and robustness requirements) and showing adversarial attacks against improper further assumptions (re. robustness).
Presently, I will simplify down to the case where we have some (modified) LI traders used according to the overall context of this comment. Tell me if this step is too bold. Also, note that this account is actually simplified vs. what would work for a real decision theory, for the sake of ease of reasoning and to avoid following unpublished accounts.
In an attempt to describe what you propose, while meeting the desiderata, imagine we have a pair of traders, } is not. Stipulate the set is e.c. when } is e.c. Since, as can clearly be seen, this situation is not symmetrical (
tr1andtr2. (Note that sharing can be chained as long as there are no loops at the subroutine level, but we will stick to a pair here WOLOG.) Let’s say we have a definition of “efficiently computable” (e.c.) that goes false when an individual trader takes on too much work per “day,” and also goes false on a collective level if there is too much compute demanded overall. Saytr1is e.c., but {tr1,tr2, …,trtr2is removed. Let’s say we try to transformtr2(not e.c.) intotr2prime (e.c. when allowed to referencetr1), by lettingtr2prime reference some subroutine oftr1, and letting the cost reduction be accepted in the standard way when determining e.c. in both the individual (when allowed to reference) and collective definitions. Say that the set {tr1,tr2prime, …,trtr1can survive alone, buttr2can not), a well calibrated model with holdouts would say thattr2(in the operational form oftr2prime) is “less probable” to survive thantr1. (Holdout method not given, compute-limited modeling roughly following the standard for LIs.) Note that to the extent that any of these “traders” have “inner” instrumental rationality, they are unable to take any “action” other than outputting a trading strategy for the “day,” for fairly standard betting market incentive reasons. For this reason, we assume the set of “traders” is brought into existence by some outer process in their operational forms, with any references already established. Eviction must be extremely efficient. I am unsure if this prohibits re-parenting of a subroutine, but for simplicity assume that the operational set wouldn’t allowtr2to survive even if it was allowed. Deviating a bit from formality, generalizetr1totrxandtr2prime totry, both in the domain of situations of this sort. Define the function ev_p(tr) to be the bounded-compute “probability” of the eviction oftr(see before on modeling). Evaluate ev_p(try) - ev_p(trx). Call this “the cost of ontological overlap.”As said before, even if a “trader” has a decision theory set up internally, it won’t be hooked up to the interface between the “trader” and the (actual) wrapping hardware in any traditional sense. However, let’s look at a case where we imagine some alternative to
tr2that could have done better thantr2, without having any ontological overlap, while being e.c. Call ittrv. Notice that it has, at the absolute least, regained the entire (non-formally made specific again) “cost of ontological overlap.”Surely, there can be savings by sharing the execution of subroutines, but this trades off against atomicity and separability. Depending on the eventual costs of things like re-parenting, and more advanced techniques like the automatic determination of acceptable approximations and the determination of “equivalent” functions, this may trade off substantially against robustness as well, since the eviction of a single “trader” will sometimes result in a massive over-eviction, if the dependencies can’t be fixed up in an extremely rapid manner (against optimality as well, but we omit that here).
This attempts to model “communication” at some level (actually quite well, given some theories of communication and collective rationality I can vaguely remember), but note that, treating it as a sort of agent it usually won’t be,
trvhas no “reason” to want to be replaced withtr2prime. With some basic assumptions on the setup, e.g. strictly positive initial trader budgets,trvbeing replaced withtr2prime is a strict decrease in estimated survival “probability.” Therefore, this “communication”/sharing must only be done to the extent that it really helps the overall agent, and not following some procedure or “virtue” that is thought to be worth universally adopting.(We’re assuming the set with
trvreplacingtr2prime is also e.c. Note that in this entire comment we’re assumingtr2andtr2prime are functionally identical, to the level of detail we care about here.)Okay, then let’s try to go back to English. Let’s think about these “traders” as cartoons, in the same sense that a cartooned evolution can speak and (be said to) want things. So this trader says, “I want to share load onto other models run by other traders if and only if it reduces my compute demands enough to achieve viability. Sharing ontology for other reasons just makes me worse off, in expectation. Depending on other ontologies is a risk that must be taken on only after sufficient calculation, since ontologies (via the traders that operate them) may go bust.”
(In a more realistic system, it would need to be considered what to do when a “trader” is “exporting” a subroutine that it doesn’t want to use any more. For simplicity, you can ignore this for now. Note that the cartoon dialogue doesn’t really survive this, but it’s cartoon dialogue for a reason. Note that “going bust” should be read as referring to “traders” getting evicted, in ways that don’t strictly relate to bankroll.)
I can’t think of an alternative to this setup, and since the (new) intermediate desiderata seem well motivated, I’m hesitantly taking this as a (sketched, not fully filled in) disproof. One reason for my hesitancy is that my model of proper thinking may require too much compute and too much implementation complexity to ever be usable. Sure, maybe I can fill in all the details and get a system that’s proven aligned, but it might not be applicable to humans.
Do you know of a technical or semi-technical note on how this ontology framework (as an argument for sharing load) functions here/meets desiderata? It doesn’t need to be an explainer, so don’t particularly worry about the quality or any missing details.
I’m not an expert in this area, so maybe the solution really is simple and I’m just not seeing it, or I’m making a silly mistake.
(My hand written heuristics claim this comment is dense to the point of near-unreadability, so all (prospective) readers are invited to ask questions, including if a reading attempt has yet to be made.)
The step of running with the alien assumptions and framings is also very important, it’s unnecessary and impractical to only focus on the things you believe or understand or endorse. There is much more to learn than what’s ready to be made into a part of yourself, and the only way to defeat path-dependence at human level is by being anti-inductive with respect to your own perspective, to seek out the points of view you don’t understand or endorse or believe (on some topic of interest), and at least come to understand how they work internally for their proponents, even when you don’t have footholds of belief or endorsement into them that promise they might be a good fit down the line. (This does require solid sandboxing practices, or else your mind might get so open the brains fall out.) Sometimes such footholds are unexpected, and only appear after you do the work, despite not having any footholds at the outset, and so you make your own epistemic luck.
I didn’t think this through carefully, but maybe an epistemic PD where the currency is compute? Like, I spend 5 utils worth of compute; from your perspective that’s worth 10 expected utils because I might correct my false (according to you) belief; and you do the same for me.