Eudaimonia, a Greek word literally translating to the state or condition of ‘good spirit’, and which is commonly translated as ‘happiness’ or ‘welfare’.
Canaletto
This adds some weird assumption that all of what I value can be expressed as my experience per unit of time? Like, what if I have preferences that concern the real world, besides preferences that concern my feelings.
Hmm. Wouldn’t you have to work with its approximations or approximations of its variants, as irl systems have to take finite time to decide on any action? Irl systems such as “real LLM-based systems”.
I think it’s a bit precarious. If you made AIXI variant that actually works, do you think it would go well?
Do you have a plan for how to deal with discoveries you could make that enable capability gain for such extremely by-construction sociopathic systems.
Well, N=1, but I leave comments regularly on old posts.
Also, if you look at https://www.greaterwrong.com/?sort=active there are some posts from 2022 − 2025, so it happens regularly, maybe rarely per post tho.
In my default mindstate, thinking of ‘evil’ is like a weird alien writhing wriggling thing that other humans engage in for some reason
It’s pretty funny, I think most common evils are done by people with that mindset. Evil is what other people do. I’m fully justified in everything I do, and even if not, they are completely understandable and excusable things under unfortunate circumstances.
Well, everyone would prefer the model that does not whistleblow for themselves, but maybe not for other people to have that model.
Like, imagine, someone using Claude to screw you up, and it whistleblows and the thing gets averted. That would be pretty nice.
Omega: Perfect predictor, works by looking at your state.
Diluted Omega: if random()<P: Omega(opponent) else Coin(opponent)
Hyperstition Omega: Predicts by assuming its opponent will play best response to its current track record, as if it’s a Diluted Omega with matching P.It’s kinda iffy to define Omega by its track record, it might have been just lucky Hyperstition Omega, e.g. in Newcomb’s it got few first predictions correct, then players on seeing its track record started one boxing, and it kept providing full boxes, and it reinforced its track record.
You can coordinate on exploiting Hyperstition Omega for more money than you could have got from actual Omega. You need to keep its track record just barely above the point where it switches to giving out empty box, if everyone coordinates on particular one/two boxing mixed strategy—but this is basically multiplayer Prisoners’ dilemma
Well, it’s a bit tricky, if you think the predictor is bad, and two box, then you become predictable, and the predictor becomes better. There is some logical time game of tag, but with very good predictors this is irrelevant, yeah.
You can also one up bad predictors, by looking like you are going to one box and then two boxing, but that’s distinct from randomizing. There is no good formalism for this, c.f. Schelling points.
So if we transform this into a Newcombe-like scenario, I think the rational strategy would be to two-box, because being able to predict my System 1 isn’t able to predict me as an agent.
Predictor: He will think I will not be able to predict him, so he will two box.
You: It will not be able to predict me, so I two box.
Harmony
FDT only applies to embedded agents
I mean, unembedded agents are not necessarily unpredictable? You don’t need to fully instantiate an agent to predict it, e.g. 100 parameter N-grams model can predict human choice in Rock Paper Scissors, after observing some number of moves. So, even if agent and the universe are separate and converse through a narrow channel, you need to track logical/correlational dependencies of your choice with stuff in universe.
Functional Decision Theory (FDT) is a decision theory for a rational agent X who holds the rational belief “the probability that agent Y correctly predicts my actions is very high.”
You don’t need high prediction accuracy for that, any edge would work. Also, it’s more than that. E.g. you can describe EDT like that, and it does not pay in Counterfactual mugging for example.
Also, I think isomorphism should also be complete, not just partially true, you can’t map yourself to only one model, you need to map yourself to all models out there. If you have two clones, you need to map yourself to both, you should not consider (true but partial) isomorphism that maps only to one clone.
e.g.
There are two independent predictors who predict what would you pick in Rock Paper Scissors, each puts money on some symbol and debt notes on symbol it beats, of the same amount of money. When you pick a symbol you get the money from symbol it beats, but have to pay the debt from your chosen symbol.
(basically additive Rock Paper Scissors against multiple opponents at the same time)
Additionally, first predictor P1 is (4/10, 3⁄10, 3⁄10), and second predictor P2 is (6/10, 2⁄10, 2⁄10), on win-draw-loss. They place $10 bet each. What amount of money you would have to be paid to play this game?
Isomorphism that maps you only to the model of P1 gets you more expected value than isomorphism that maps you to P2′s model or both models.
Phi(P1) > Phi(P2) > Phi(P1, P2)
Misleadingly.
Does p-FDT pay in Counterfactual mugging? E.g. Omega tosses the coin, shows it to you, explains the setup, asks you to pay $10 to get $100 in other hypothetical branch.
I think many humans would reason like “those who in general pay in such (honest, reliable, non adversarially picked) muggings are in general richer” and pay. Can p-FDT do that?
I don’t think there exists exploitable map here? Like, it seems like your version of this decision theory is non updateless, it thinks “what should I do in this situation, given that there also are hidden ways my choice can influence the world”. But it did the update already?
The “find the isomorphism phi” part works well.
Does it? I think you defined it weighted by probability it’s true, but I’m confident that N-gram model is not isomorphic to me in the same way calculators with and without—on the end result are isomorphic.
EDIT Unrelatedly, I also think if “Omega predicts what percent of people from your country would have paid in this situation, then asks you to pay it $10, and then gives you (fraction of people who paid) * $1000 of value in consumables”—then you should pay it, it’s basically voting problem.
I also agree with the caveat you spotted, but it was not what I intended to point at haha.
I mean, this post is talking about simple tiling processes, but what about complex ones? E.g. ASI grabbing the matter in its bubble, and accelerating with time up to some 0.999*c or whatever, if it manages to print first stage receivers on distant matter with light or something tricky like that.
You can explicitly say what happens if you are not in selection, I guess. Like “Omega takes a look at you from stealth. If it thinks you would not say banana if given $100 and explanation of the setup, it approaches you and does it. Otherwise it leaves without turning off stealth or explaining anything”—then everyone is in selection again.
I think XOR problem is of this type. Smokers lesion deals with this incorrectly / underdefined. Transparent boxes Newcomb’s can be fixed in the same way. Although I think the one fixed with counterfactual predictor is better as a problem, although is a different one.
That’s not my point. I claim that some problems are stated badly, stated incorrectly, illposed. Like “Would you glomp when it’s romcsing?” or “It’s raining in Brazil. How many times will you jump?”. It’s not about what happens in problem, it’s about what happens among people who consider it.
Please explain what makes a Banana problem well posed.
It seems to me your isomorphism idea is too restrictive. Humans at least, can be predicted by very simple algorithms e.g. N-grams, they can have an edge on you in Rock Paper Scissors, so you would be foolish to play against it. I’m not isomorphic to a 100 parameter model, and yet I should take into account its predictions of my choices.
Does your formulation able to accommodate that?
Some problem statements in decision theory can be non universal. Like, they are possible as situations that can happen, but they are invalid as problems to ask what would you do in them.
Suppose Omega gives a single box with $100 to people who don’t say banana after seeing the explanation of the setup and the box. You encounter such situation. Would you say “banana”?
Like, it’s funky premise? It’s not applicable to all people? It would not give me such opportunity, so the premise is invalid.
It pre selects what kind of guys get to participate. So, some guys don’t. So, if you ask what those guys will do in that situation? It has invalid premise of them being in selection.
Newcomb’s problem is fine on that front, it accommodates every kind of strategy, namely, “see opaque and transparent boxes” → two box, and “see opaque and transparent boxes” → one box.
Transparent Newcomb’s problem needs adjustment for this selection effect danger, because in that case you can go against the prediction, e.g. one box if it’s empty and two box if they both are full. What Omega offers to those guys, huh?
One solution is “Omega leaves both full IFF you one box after seeing both full”, then some contrarians do receive only $1000 and leave with $0 after one boxing. Another “Omega leaves both full IFF no matter what you see you take one”—but this one is closer to counterfactual mugging.
Smoker’s lesion might be one such problematic problem. Agents will act on correlation they perceive, but in doing so, destroy the correlation. You can fix this by introducing staged information gain, like “People decided to collect the statistic each 10 years. This was the first time they did that. There is such correlation. Before next statistic collection, what do you do? ” or selection effects, like “even after acting on this correlation, correlation remains”—and this looks like it makes weird postulates, that might be non universal.
EDIT
XOR blackmail uses this selection effect deliberately. There is a selection effect to what kind of agents the letters arrives. If you received a letter, then you have termites XOR you will pay small fee. But it’s evidence about the state of the world not your decision, if you are the kind of person who would not pay. There is no way to act on this information and get out of bounds of what is postulated.
I sense you mean something else than this wiki quote above when you use that word, given downvote and disagree vote? But, man, that’s exactly what made me think “rules out you having preferences that aren’t about your feelings” or at least establishes conversion to them.