AIS student, self-proclaimed aspiring rationalist, very fond of game theory.
”The only good description is a self-referential description, just like this one.”
momom2
If the agent can’t buy both boxes, probably you mean that she predicts between buy A, buy B or buy none?
Note, thanks to Claude Opus 5:
Since the payoff is linear in the number of boxes predicted not matching the agent’s choice, it doesn’t matter how the predictor predicts, only her probability for each box of predicting it to match the agent’s choice.
So I guess objection retracted; there’s only one relevant meaning for your setup, the one in which the predictor is right about any given box 75% of the time (though perhaps different framings may have some philosophical importance or lead to anthropic arguments).
Since the tickle defence is so critical to rescuing EDT in the smoking lesion problem, I don’t think it’s quite enough to just say that it exists.
Or well, the way you described it is consistent, but it does feel you’re waving your hands in a “trust me bro” way, considering your stated intent is to make EDT look good compared to CDT. It’d be much more enjoyable reading if you explained the tickle defence in enough detail to convince me.
Lastly, I’m very passionate about decision theory and I think I understand FDT pretty well. I’ll be very happy to chat more in DMs or on Discord (#momom2) if you want.
It seems to me that we can formalize FDT in the adversarial offer as such:
- FDT considers that they choose the outcome of a program Decision(adversarial_offer) which outputs either None, A, B or AB.
- They consider that the predictor runs a program Prediction(A) which outputs (A \in Decision(adversarial_offer)) with probability 0.75 and not(A \in Decision(adversarial_offer)) with probability 0.25, then the same for B, and fills the boxes accordingly.
- Then, if FDT chooses None, they get 0.
- If FDT chooses A, they get 0.25 * 3 − 1 = −0.25
- Likewise if FDT chooses B.
- If FDT chooses AB, they get 0.25 * 3 * 2 − 2 = −0.5
So they pick None.
Suppose now that they consider that the predictor runs a program Prediction() which outputs Decision(adversarial_offer) with probability 0.75, and otherwise picks at random among None, A, B and AB, then fills the boxes accordingly.
- Then if FDT chooses A, they get (0.75 * 0) + (0.25 * 0.25 * 2 * 3) − 1 = −0.625
- Likewise for B.
- If FDT chooses AB, they get (0.75 * 0) + (0.25 * 0.25 * 2 * 3) + (0.25 * 0.25 * 6) − 2 = −1.25
So they pick None.
Suppose now that they consider that the predictor runs a program Prediction() which outputs Decision(adversarial_offer) with probability 0.75 and otherwise (AB—Decision(adversarial_offer)) then fills the boxes accordingly.
- Then if FDT chooses A, they get (0.75 * 0) + (0.25 * 3) − 1 = −0.25
- Likewise for B.
- If FDT chooses AB, they get (0.75 * 0) + (0.25 * 6) − 2 = −0.5
So they pick None.
So under all three readings I can make of the AO problem, FDT picks None.
Your adversarial offer situation is poorly explained.
- Does the predictor fill each box independently or does she predict the overall (double-box) choice of the agent? You describe her behavior as if she performs the former, but then implicitly assume it’s impossible for both boxes to be empty.
- How do we know the predictor’s reliability, and what does the agent know about it? Are they 100% certain that the predictor is 75% likely to predict their choice correctly (which seems contradictory with CDT’s premise that their choice is causally downstream of the prediction)? In the case where the predictor is wrong, does she know she’s wrong? Does she pick at random or does she pick the opposite of the agent for each box?
I find very suspicious the fact that you claimed you could derive that CDT would prefer buying at least one box, but not in a constructive way; I’m not quite sure why, because of the above imprecision in the specification, so I’d like more clarity.
But… Where does the intuitive reaction of raging against the unfairness of the world come from?
Is it the case that this intuition was evolved in an environment where you could always flip the table if you didn’t like the dice’s outcome and you were angry enough? (If yes, okay, but if not, then there’s probably some reason it’s there which we should understand before tearing down the intuition.)
Also, treating the outcome as information about where I am is the wrong framing. It’s basically EDT; you should take it as information about where the kind of person I am leads.
Bravo.
I know it’s not usual for this place, but in this case I think it’s worth giving non-constructive praise just to signal your virtue.
Read Ra.
Instead of systematization and explicit legibility, Ra chooses an impression of abstract generality which, upon inspection, turns out to be zillions of ad hoc special cases.
I see Ra in my own vibe-coded software: I ask AI to systematically seek generalizations, build from sound principles and strong foundations. But I rarely check which foundations, if any, it builds from—occasionally I catch it in egregious fakery and ad-hoccing and correct it, but often the directive of “find sound principles to build from” is a vibe more than a constraint.
There is nothing Ra-like, for instance, about noticing that software is a fully general force multiplier and trying to invest in or make better software. Ra comes in when you start admiring force multipliers for no specific goal, just because they’re shiny.
I’m guilty of this, which XKCD calls premature optimization, and it’s a big drain on my resources and obstacle to finishing software projects.
It seems to be clear that in practice, we cannot push towards pausing at a specific point in the future. There is no Schelling point nor a clear red line common for everyone.
Anthropic just posted a new post explaining what the economic future will look like. (Technical report)
They explicitly refuse to examine takeoff dynamics and implicitly present things as “business as usual” even in their extreme scenario.
They have a probabilistic model, but don’t price in existential risk at all.
This is very disappointing. I read it as normie-washing.
I don’t feel like I misunderstood your phrasing:
- One of the main arguments in favor of the Grabby Aliens model is that it explains why we don’t see aliens.
- Under your model, we should see aliens unless there’s some reason they’re not visible from far away.
- Therefore, one argument against your model is that it does not solve the Fermi question.Thus goes the argument for Grabby Aliens:
- Assuming interstellar civilization moved very slowly compared to the speed light, there should be a bunch of visible aliens.
- There aren’t a bunch of visible aliens in the sky (Fermi question)
- Therefore if there are interstellar civilizations, they probably move quite fast relative to the speed of light.You can say, “well, let’s discuss a bunch of possible reasons aliens are not visible”, which is basically the literature on the Fermi paradox, but it doesn’t solve it in the elegant way the Grabby Aliens model does.
So I guess your take on this is to compare the probability of expanding at c being possible based on this article’s arguments, and the probability of Grabby Aliens model based on its own arguments?
This is a surprisingly optimistic result! It is already well known that models cheat on evals, but this shows that in some way they care about honesty and response quality when they can get away with it.
Yeah, I read that article, and it convincingly argues that ARA won’t kill literally everyone. I think the main insight is that ARAs have much more selection pressure on persisting as worms than on getting more intelligence and setting up the autonomous production economy that a rogue AI would need to subsist without human-made infrastructure.
But it does not make much of a good job arguing that ARAs won’t destroy the Internet, and are not doing so right now. Its main argument there is that humans won’t accept it because it’d destroy a lot of value, and, okay, that makes sense, but the missing step here is that even if we wanted to, I don’t see how we could prevent the destruction of the Internet.We can rebuild; maybe airgapped local communities will become the norm after a large fraction of public spaces are consumed by ARAs. I’m interested in preventing even that from happening, since I believe it would already be a big loss.
Also, who is looking out for ARAs and how? From the discussion, it seems to be all “I considered it for 5 mn”, not any actual professional effort.
Tell, is anyone investigating to try to find adaptative worms and rogue agents in the wild?
Considering the strong theoretical arguments behind them, and the lack of empirical evidence against them, it seems quite likely that in fact at this very moment rogue agents are hacking into compute resources to reproduce. I just don’t see a reason why it wouldn’t be happening.
Is anyone looking for evidence of such happening?
If we don’t see other civilizations because [...] they expand at only a small fraction of the speed of light
That seems backward to me. If a civilization expands at a small fraction of the speed of light, they should both reach a large fraction of space and reach us much later than their light reaches us, making them maximally visible.
On the contrary, one of the main advantages of the Grabby Aliens model is that it explains why we don’t see alien civilizations: by expanding at an appreciable fraction of the speed of light, the time between the moment they become a galactic expanding civilization and the moment they reach us is relatively small, giving low prior probability of us noticing them the moment we turn our eyes to the sky (thus explaining away the Fermi question).
How does one subscribe to your journal?
A followup question I’d be interested in the answer of is how accurate are these opinions? Do you have a dataset with gender-tagged messages to check?
(It would especially help me better understand how much to trust these probes.)
Thanks! I don’t like podcasts and long conversations in audio form, but I really enjoyed reading this.
Vibe embodiment.
Thanks for the tips; they sound useful, but also, the general vibe of what you describe is so viscerally repulsive to me that I can’t bring myself to appreciate the usefulness of your post.
I’ve already introspected. I know what I’ll find, and I won’t like it.