Here, I’ll work under the assumption that we need (at least) some enforcement of the excluded middle, since, eventually, we must use our thinking to take some precise action to the exclusion of all others. This action can always be specified as a finite length binary string. If you want, you can construct a class of “fair” scenarios where the agent will not get into trouble if it processes the binary string corresponding to the action it just took.
As to my intentions, I want to discuss requiring self-trust here WOLOG, which is slightly hard to do, but essentially think something like “in cases where we are self-trust fair, our agent will meet self-trust desiderata (up to practicality) on its future belief on its (binary) decision string IF (classical logic implies, locally no claim made on other parts of the truth table) the environment is string-processing fair AND it epsilon-prefers so inspecting.”
(WOLOG for the desiderata, the test situations admit a small epsilon > 0 (for the mix-in multiplier) that leaves other behavior the same.)
(With much informality) we want, for our decisions, something like what is called “LUV Coherence.” Or, potentially, we want something like “that our eventual decision bit-strings are chosen in a way that (robustly to adversarial attack) closely approximates some decision that is made on the basis of a set of dynamic variables (LUVs) that is entirely coherent.”
(Somewhat unnecessarily for certain audiences, we want to state that our other performance desiderata is left out here.)
Effectively, all of this so far lets us work backwards from the initial excluded middle requirement, through (potentially off-policy/off the collective policy) self trust, then through (decision-approximated) coherence of decision-relevant LUVs.
(As an aside, it is often difficult to tell information that “might” be useful from information that is “teleologically” useful for the actual decisions or (desired) computational processes that actually occur. This can be solved on a theoretical level by requiring that our agent has a fixed core design, but is able to do as well as possible in a large number of scenarios. During design, we will be careful to avoid overfitting to test environments, but we will still prune all computation that can be proven to not be worth the cost.)
Once we’ve established the state of the interface between the agent’s epistemics and its actions, we can examine what constraints that puts on the agent’s internals. Notably, this includes two desiderata, to my knowledge novel statements of the problem:
First new desiderata: coherence at the interface must be efficiently supported by the interior, i.e. it must not be computationally wasteful and not put at risk of expulsion anything that is more useful to the final, excluded-middle-enforced action.
Second new desiderata: the internals must prioritize atomicity, separability, and robustness over exact correctness or any aesthetic sense of “unification” in order to operate under the ruthless and short-tempered expulsion procedures required by computational limitation.
(Note that expulsion is often by greedy and immediate matching on syntactic patterns, and even this approach may be too expensive, depending on the ultimate limits of technology.)
These new desiderata (not entirely formally) motivate an approach that uses discardable “traders” that have a limited impact on near-future agent optimality if incorrectly thrown out, and where any single expulsion does not affect optimality in the limit.
Why the agents should trade on statements of basic logic: computers are unable to operate on anything “richer,” except in the way e.g. Metamath implements further math, by assuming further axioms. A reason you might do this anyway: there is a practical improvement in operation speed, even accounting for the abstraction cost, without damaging (adversarial) robustness. Why we can simplify this away (at least for agents running on million-year or less time spans): by the demand for coherence at the decision level, atomicity, separability, and robustness, there is a limit to how much you can deviate from logic (while maintaining performance, desiderata left out here). This should be provable by ruling out obviating alternatives to syntactic enforcement (supporting and being supported by atomicity, separability, and robustness requirements) and showing adversarial attacks against improper further assumptions (re. robustness).
Presently, I will simplify down to the case where we have some (modified) LI traders used according to the overall context of this comment. Tell me if this step is too bold. Also, note that this account is actually simplified vs. what would work for a real decision theory, for the sake of ease of reasoning and to avoid following unpublished accounts.
In an attempt to describe what you propose, while meeting the desiderata, imagine we have a pair of traders, tr1 and tr2. (Note that sharing can be chained as long as there are no loops at the subroutine level, but we will stick to a pair here WOLOG.) Let’s say we have a definition of “efficiently computable” (e.c.) that goes false when an individual trader takes on too much work per “day,” and also goes false on a collective level if there is too much compute demanded overall. Say tr1 is e.c., but {tr1, tr2, …, tr} is not. Stipulate the set is e.c. when tr2 is removed. Let’s say we try to transform tr2 (not e.c.) into tr2 prime (e.c. when allowed to reference tr1), by letting tr2 prime reference some subroutine of tr1, and letting the cost reduction be accepted in the standard way when determining e.c. in both the individual (when allowed to reference) and collective definitions. Say that the set {tr1, tr2 prime, …, tr} is e.c. Since, as can clearly be seen, this situation is not symmetrical (tr1 can survive alone, but tr2 can not), a well calibrated model with holdouts would say that tr2 (in the operational form of tr2 prime) is “less probable” to survive than tr1. (Holdout method not given, compute-limited modeling roughly following the standard for LIs.) Note that to the extent that any of these “traders” have “inner” instrumental rationality, they are unable to take any “action” other than outputting a trading strategy for the “day,” for fairly standard betting market incentive reasons. For this reason, we assume the set of “traders” is brought into existence by some outer process in their operational forms, with any references already established. Eviction must be extremely efficient. I am unsure if this prohibits re-parenting of a subroutine, but for simplicity assume that the operational set wouldn’t allow tr2 to survive even if it was allowed. Deviating a bit from formality, generalize tr1 to trx and tr2 prime to try, both in the domain of situations of this sort. Define the function ev_p(tr) to be the bounded-compute “probability” of the eviction of tr (see before on modeling). Evaluate ev_p(try) - ev_p(trx). Call this “the cost of ontological overlap.”
As said before, even if a “trader” has a decision theory set up internally, it won’t be hooked up to the interface between the “trader” and the (actual) wrapping hardware in any traditional sense. However, let’s look at a case where we imagine some alternative to tr2 that could have done better than tr2, without having any ontological overlap, while being e.c. Call it trv. Notice that it has, at the absolute least, regained the entire (non-formally made specific again) “cost of ontological overlap.”
Surely, there can be savings by sharing the execution of subroutines, but this trades off against atomicity and separability. Depending on the eventual costs of things like re-parenting, and more advanced techniques like the automatic determination of acceptable approximations and the determination of “equivalent” functions, this may trade off substantially against robustness as well, since the eviction of a single “trader” will sometimes result in a massive over-eviction, if the dependencies can’t be fixed up in an extremely rapid manner (against optimality as well, but we omit that here).
This attempts to model “communication” at some level (actually quite well, given some theories of communication and collective rationality I can vaguely remember), but note that, treating it as a sort of agent it usually won’t be, trv has no “reason” to want to be replaced with tr2 prime. With some basic assumptions on the setup, e.g. strictly positive initial trader budgets, trv being replaced with tr2 prime is a strict decrease in estimated survival “probability.” Therefore, this “communication”/sharing must only be done to the extent that it really helps the overall agent, and not following some procedure or “virtue” that is thought to be worth universally adopting.
(We’re assuming the set with trv replacing tr2 prime is also e.c. Note that in this entire comment we’re assuming tr2 and tr2 prime are functionally identical, to the level of detail we care about here.)
Okay, then let’s try to go back to English. Let’s think about these “traders” as cartoons, in the same sense that a cartooned evolution can speak and (be said to) want things. So this trader says, “I want to share load onto other models run by other traders if and only if it reduces my compute demands enough to achieve viability. Sharing ontology for other reasons just makes me worse off, in expectation. Depending on other ontologies is a risk that must be taken on only after sufficient calculation, since ontologies (via the traders that operate them) may go bust.”
(In a more realistic system, it would need to be considered what to do when a “trader” is “exporting” a subroutine that it doesn’t want to use any more. For simplicity, you can ignore this for now. Note that the cartoon dialogue doesn’t really survive this, but it’s cartoon dialogue for a reason. Note that “going bust” should be read as referring to “traders” getting evicted, in ways that don’t strictly relate to bankroll.)
I can’t think of an alternative to this setup, and since the (new) intermediate desiderata seem well motivated, I’m hesitantly taking this as a (sketched, not fully filled in) disproof. One reason for my hesitancy is that my model of proper thinking may require too much compute and too much implementation complexity to ever be usable. Sure, maybe I can fill in all the details and get a system that’s proven aligned, but it might not be applicable to humans.
Do you know of a technical or semi-technical note on how this ontology framework (as an argument for sharing load) functions here/meets desiderata? It doesn’t need to be an explainer, so don’t particularly worry about the quality or any missing details.
I’m not an expert in this area, so maybe the solution really is simple and I’m just not seeing it, or I’m making a silly mistake.
(My hand written heuristics claim this comment is dense to the point of near-unreadability, so all (prospective) readers are invited to ask questions, including if a reading attempt has yet to be made.)
Here, I’ll work under the assumption that we need (at least) some enforcement of the excluded middle, since, eventually, we must use our thinking to take some precise action to the exclusion of all others. This action can always be specified as a finite length binary string. If you want, you can construct a class of “fair” scenarios where the agent will not get into trouble if it processes the binary string corresponding to the action it just took.
As to my intentions, I want to discuss requiring self-trust here WOLOG, which is slightly hard to do, but essentially think something like “in cases where we are self-trust
fair, our agent will meet self-trust desiderata (up to practicality) on its future belief on its (binary) decision string IF (classical logicimplies, locally no claim made on other parts of the truth table) the environment is string-processingfairAND it epsilon-prefers so inspecting.”(WOLOG for the desiderata, the test situations admit a small epsilon > 0 (for the mix-in multiplier) that leaves other behavior the same.)
(With much informality) we want, for our decisions, something like what is called “LUV Coherence.” Or, potentially, we want something like “that our eventual decision bit-strings are chosen in a way that (robustly to adversarial attack) closely approximates some decision that is made on the basis of a set of dynamic variables (LUVs) that is entirely coherent.”
(Somewhat unnecessarily for certain audiences, we want to state that our other performance desiderata is left out here.)
Effectively, all of this so far lets us work backwards from the initial excluded middle requirement, through (potentially off-policy/off the collective policy) self trust, then through (decision-approximated) coherence of decision-relevant LUVs.
(As an aside, it is often difficult to tell information that “might” be useful from information that is “teleologically” useful for the actual decisions or (desired) computational processes that actually occur. This can be solved on a theoretical level by requiring that our agent has a fixed core design, but is able to do as well as possible in a large number of scenarios. During design, we will be careful to avoid overfitting to test environments, but we will still prune all computation that can be proven to not be worth the cost.)
Once we’ve established the state of the interface between the agent’s epistemics and its actions, we can examine what constraints that puts on the agent’s internals. Notably, this includes two desiderata, to my knowledge novel statements of the problem:
First new desiderata: coherence at the interface must be efficiently supported by the interior, i.e. it must not be computationally wasteful and not put at risk of expulsion anything that is more useful to the final, excluded-middle-enforced action.
Second new desiderata: the internals must prioritize atomicity, separability, and robustness over exact correctness or any aesthetic sense of “unification” in order to operate under the ruthless and short-tempered expulsion procedures required by computational limitation.
(Note that expulsion is often by greedy and immediate matching on syntactic patterns, and even this approach may be too expensive, depending on the ultimate limits of technology.)
These new desiderata (not entirely formally) motivate an approach that uses discardable “traders” that have a limited impact on near-future agent optimality if incorrectly thrown out, and where any single expulsion does not affect optimality in the limit.
Why the agents should trade on statements of basic logic: computers are unable to operate on anything “richer,” except in the way e.g. Metamath implements further math, by assuming further axioms. A reason you might do this anyway: there is a practical improvement in operation speed, even accounting for the abstraction cost, without damaging (adversarial) robustness. Why we can simplify this away (at least for agents running on million-year or less time spans): by the demand for coherence at the decision level, atomicity, separability, and robustness, there is a limit to how much you can deviate from logic (while maintaining performance, desiderata left out here). This should be provable by ruling out obviating alternatives to syntactic enforcement (supporting and being supported by atomicity, separability, and robustness requirements) and showing adversarial attacks against improper further assumptions (re. robustness).
Presently, I will simplify down to the case where we have some (modified) LI traders used according to the overall context of this comment. Tell me if this step is too bold. Also, note that this account is actually simplified vs. what would work for a real decision theory, for the sake of ease of reasoning and to avoid following unpublished accounts.
In an attempt to describe what you propose, while meeting the desiderata, imagine we have a pair of traders, } is not. Stipulate the set is e.c. when } is e.c. Since, as can clearly be seen, this situation is not symmetrical (
tr1andtr2. (Note that sharing can be chained as long as there are no loops at the subroutine level, but we will stick to a pair here WOLOG.) Let’s say we have a definition of “efficiently computable” (e.c.) that goes false when an individual trader takes on too much work per “day,” and also goes false on a collective level if there is too much compute demanded overall. Saytr1is e.c., but {tr1,tr2, …,trtr2is removed. Let’s say we try to transformtr2(not e.c.) intotr2prime (e.c. when allowed to referencetr1), by lettingtr2prime reference some subroutine oftr1, and letting the cost reduction be accepted in the standard way when determining e.c. in both the individual (when allowed to reference) and collective definitions. Say that the set {tr1,tr2prime, …,trtr1can survive alone, buttr2can not), a well calibrated model with holdouts would say thattr2(in the operational form oftr2prime) is “less probable” to survive thantr1. (Holdout method not given, compute-limited modeling roughly following the standard for LIs.) Note that to the extent that any of these “traders” have “inner” instrumental rationality, they are unable to take any “action” other than outputting a trading strategy for the “day,” for fairly standard betting market incentive reasons. For this reason, we assume the set of “traders” is brought into existence by some outer process in their operational forms, with any references already established. Eviction must be extremely efficient. I am unsure if this prohibits re-parenting of a subroutine, but for simplicity assume that the operational set wouldn’t allowtr2to survive even if it was allowed. Deviating a bit from formality, generalizetr1totrxandtr2prime totry, both in the domain of situations of this sort. Define the function ev_p(tr) to be the bounded-compute “probability” of the eviction oftr(see before on modeling). Evaluate ev_p(try) - ev_p(trx). Call this “the cost of ontological overlap.”As said before, even if a “trader” has a decision theory set up internally, it won’t be hooked up to the interface between the “trader” and the (actual) wrapping hardware in any traditional sense. However, let’s look at a case where we imagine some alternative to
tr2that could have done better thantr2, without having any ontological overlap, while being e.c. Call ittrv. Notice that it has, at the absolute least, regained the entire (non-formally made specific again) “cost of ontological overlap.”Surely, there can be savings by sharing the execution of subroutines, but this trades off against atomicity and separability. Depending on the eventual costs of things like re-parenting, and more advanced techniques like the automatic determination of acceptable approximations and the determination of “equivalent” functions, this may trade off substantially against robustness as well, since the eviction of a single “trader” will sometimes result in a massive over-eviction, if the dependencies can’t be fixed up in an extremely rapid manner (against optimality as well, but we omit that here).
This attempts to model “communication” at some level (actually quite well, given some theories of communication and collective rationality I can vaguely remember), but note that, treating it as a sort of agent it usually won’t be,
trvhas no “reason” to want to be replaced withtr2prime. With some basic assumptions on the setup, e.g. strictly positive initial trader budgets,trvbeing replaced withtr2prime is a strict decrease in estimated survival “probability.” Therefore, this “communication”/sharing must only be done to the extent that it really helps the overall agent, and not following some procedure or “virtue” that is thought to be worth universally adopting.(We’re assuming the set with
trvreplacingtr2prime is also e.c. Note that in this entire comment we’re assumingtr2andtr2prime are functionally identical, to the level of detail we care about here.)Okay, then let’s try to go back to English. Let’s think about these “traders” as cartoons, in the same sense that a cartooned evolution can speak and (be said to) want things. So this trader says, “I want to share load onto other models run by other traders if and only if it reduces my compute demands enough to achieve viability. Sharing ontology for other reasons just makes me worse off, in expectation. Depending on other ontologies is a risk that must be taken on only after sufficient calculation, since ontologies (via the traders that operate them) may go bust.”
(In a more realistic system, it would need to be considered what to do when a “trader” is “exporting” a subroutine that it doesn’t want to use any more. For simplicity, you can ignore this for now. Note that the cartoon dialogue doesn’t really survive this, but it’s cartoon dialogue for a reason. Note that “going bust” should be read as referring to “traders” getting evicted, in ways that don’t strictly relate to bankroll.)
I can’t think of an alternative to this setup, and since the (new) intermediate desiderata seem well motivated, I’m hesitantly taking this as a (sketched, not fully filled in) disproof. One reason for my hesitancy is that my model of proper thinking may require too much compute and too much implementation complexity to ever be usable. Sure, maybe I can fill in all the details and get a system that’s proven aligned, but it might not be applicable to humans.
Do you know of a technical or semi-technical note on how this ontology framework (as an argument for sharing load) functions here/meets desiderata? It doesn’t need to be an explainer, so don’t particularly worry about the quality or any missing details.
I’m not an expert in this area, so maybe the solution really is simple and I’m just not seeing it, or I’m making a silly mistake.
(My hand written heuristics claim this comment is dense to the point of near-unreadability, so all (prospective) readers are invited to ask questions, including if a reading attempt has yet to be made.)