A very large fraction of white-collar work is based around the need for monitoring for misbehaviour of various sorts, even aspects that are not directly connected to that. A very large proportion is building and following processes that generate incentive structures to reduce the fraction of people misbehaving, even if the time spent following those processes is not in itself directly “monitoring” anyone in particular.
JBlack
Was this post a good decision with no reasoning chain, or a bad decision with no reasoning chain?
If it’s in a private posting space, probably not unless the quality of posts also drops. If it’s in a forum or other shared posting space, quite likely yes even if the quality doesn’t drop.
What I mean is that for EDT agents in the population that are in this scenario, there will be zero correlation, and the agents should be able to deduce that and therefore smoke.
Whether there is a positive correlation among people who aren’t in this type of scenario or aren’t EDT agents should be irrelevant to any EDT agents.
More precisely: The factor P(outcome | action, known information) in the EDT formula is not the same thing as P(outcome | action, some prior or other) and yet the correlation stated in the scenario setup is of the latter form. The calculation that supposedly leads an EDT agent to refrain from smoking is incorrect due to this. The agent would have to ignore known information to arrive at the supposed conclusion.
I was not very surprised, because dumber models do the dumber equivalents all the time.
The most recent smarter models can talk the talk of longer-term strategic planning, but frequently what they say has huge blind spots, and even when they do state reasonable plans there are often major failures to adhere to them.
(CDT smokes extra hard, EDT stops smoking. Next statistic collection correlation vanishes, EDT starts to smoke too.)
Why did the correlation exist in the first place?
EDT agents should smoke because they all know that they are EDT agents, and conditional upon being an EDT agent in this scenario, smoking and cancer are independent. P(cancer & smoking | EDT & scenario) = P(cancer | EDT & scenario) P(smoking | EDT & scenario) because P(smoking | EDT & scenario) = 0 or 1 and in either case the equality holds.
Selection effects can’t change anything for EDT agents here.
I’m definitely on the pessimistic side, but yes I can see that spending more on safety may shorten the timeline for solving death which is probably a good thing if it doesn’t also increase s-risks.
$0 trillion on capabilities and $1 trillion on safety would be equally fine.
This impossibility appears to hold in non-anthropic situations as well.
For example, in the “Tails False Memory” scenario I commented to earlier posts: a coin is flipped and on Tails, Beauty is given a false memory of being asked “what is your credence for heads” and having answered. Then unconditionally Beauty is asked “are you sure”. Beauty knows these rules.
In both possible worlds at every point in time, there is only one agent and therefore at most one agent with any given history. So this is a non-anthropic situation at all times by the given definition.
Let the relevant histories be h1 = “I remember being asked the first question only and not having answered yet”, and h2 = “I remember being asked both questions, having answered the first”. Let the worlds be wT = “the coin comes up Tails” and wH = “the coin comes up heads”.
Beauty knows that (wT, h1) is impossible, and so Q(wT | h1) = 0 and therefore Q(wH | h1) = 1. The subsequent state (wH, h2) is a non-anthropic situation and so Beauty must use the Simple Bayes rule to update, so Q(wH | h2) = 1 also. But Q(wH | h2) + Q(wT | h2) = 1, so Q(wT | h2) = 0 which requires that Q(wT, h2) = 0.
But (wT, h2) is a possible state and must not be assigned zero credence, concluding the impossibility proof.
A moderately intelligent conscious egoist cleans the kitchen because dirty kitchens lead to food poisoning which very much sucks.
I consider the cd = ud+ck equivalence to be much more fundamental than any martingale property, so I guess that’s the difference.
If we have an epistemic model in which we need to consider whether or not a duplicate was created and destroyed without ever making any observations, then as I see it that model should be discarded as it contains internal flaws. It shouldn’t even make any difference if we bring in an inert lump of random matter instead of a non-conscious duplicate, either. From an epistemic point of view, all non-observers are equivalently irrelevant.
I’m not sure whether you’re agreeing with my first point or not, regarding the equivalence of conditional duplication with unconditional duplication + conditional killing.
Since you said that the martingale condition is fully satisfied for both unconditional duplication and for conditional killing, do you believe that it is necessarily (regardless of SIA belief) satisfied for the sequence of both?
If not, how does the first equivalence break?
Yes, I have posted an example of suddenly changing rational updates with nothing unusual happening in these comments, and previously elsewhere.
It’s more fragile than that. E.g. suppose a coin is flipped and on tails, you’ll be given a false memory of being asked your credence for heads, and answering “100%”. On heads you’ll simply be asked the question with no memory alterations. Then in both cases you’ll be asked “are you sure?”
If you’re in the situation of remembering having been asked the question but not answering, then you must be in the heads case, and should answer “100%”. But then when you’re asked “are you sure”, you are no longer sure because you now have exactly the same sort of memories that would have had in the tails case. Your memory hasn’t been altered in this case and you have no new information, so why are you now uncertain about something you were previously rationally certain of?
Any conditional duplication is (for the purposes of the scenario’s credence of heads when asked) identical to unconditionally duplicating and then killing the duplicate in room 2 on heads before awakening.
What’s more, it’s the same as just not asking the duplicate in room 2 on heads, without any killing involved.
What’s more, it doesn’t matter when that duplication happens prior to being asked, before the coin flip or after, so longer as neither knows which one is going to be asked on heads.
So the duplication isn’t the problem.
If you ask about probabilities of observer-events not relying on personal identifications, none of this is any problem whatsoever. There are epistemic symmetries between some observer-events, but not all.
In particular, the martingale condition relies on the fact that ordinarily there is a certain type of time-translation symmetry over personal identity. In this case there obviously isn’t, so that applying the consequences of such a symmetry is simply incorrect.
Given the decomposition “one believes that X iff one alieves that X and one judges that X”, would you say that the person terrified to cross the glass walkway does not believe that it is safe, but merely judges that it is safe?
I don’t understand what is meant by the second one. Do you mean that the oracle could be extremely expensive up-front, but determines the result for any given query very cheaply?
I think your comment could be correct in principle, but is applicable only in worlds that aren’t this one.
The first point doesn’t apply since we don’t have any labs where models are already pretty safe / aligned.
The second doesn’t either, because the race dynamics appear to be already at maximum and can’t go any harder due to engineering and economic limits. We also don’t have any evidence that existing labs would—or could—push any harder on capabilities if their models were better aligned.
Inability to collapse quick takes without reloading the whole page is definitely a UI hindrance, and not only on mobile. If there is a way to do it, then it is not at all obvious.
I don’t mind them being one line in desktop, since you can quickly check another 10 lines or so by hovering the mouse over them.