You seem confused; I can’t make out what claims you think you’re addressing or contradicting. I did not say the Swarm thought they were being graded by an LLM. Did anyone else think I said that?
Eliezer Yudkowsky
Thank you for recounting item 4. I think this is an important point to propagate.
I have tried bolding the contestants; I wonder if that helped by way of contrast?
Then we should look around ourselves and see an older universe.
Also this falls under the category of “distinguish the possible from the probable”. No species anywhere cares about leaving one little probe behind to stop Holocausts?
Thanks for this high-priority correction, I’ll edit.
Now is when you’ve got leverage. In 3-24 months threats to leave will have zero leverage on them because they no longer need additional employees.
I surely do credit with wisdom those EAs who went off and quietly bought bednets and saved a few kids, rather than posting to the Internet their arguments for why MIRI couldn’t possibly be a better bet than bednets. The difference between actual humility and the posturing that is Modesty.
The ones endlessly going “But what if there’s a mistake in your argument?” are the ones I’m rolling my eyes at, here. Well, golly fuck, what if there’s a mistake in your fucking argument?
Now the quiet modest ones who bought bednets and didn’t fuck around in difficult-to-analyze grownup business are going to have their work undone when the children in Africa die anyway, because other people borrowed their name and used that reputational credit to, among other things, shit all over the attempts to have Earth do something other than die.
A hope for debate might be that, if a person has a good-enough hidden stat to construct an adequate alignment plan in the traditional way (reading, thinking, corresponding with other humans, etc.) given enough time
They don’t. They were not converging toward correct answers; their updates from debate were not moving in a correct direction. So far as I can tell, EA given unlimited time to think, but no further evidence, would’ve not figured out that AGI wasn’t scheduled for 2050.
Looks like you got your key insight down to one sentence, rather than needing to start out with a long description of a complicated mechanism that would choose eleven mice for some unknown reason that you could only explain later!
If the EAs couldn’t do it on AI timelines, neither can whoever Dario picks to be a debate judge.
An AI smart enough to solve alignment is smart enough to fool me about whether it has solved alignment, so I’m not particularly telling you to solve this problem by slotting me in as a judge. TBC, I expect to be much, much better than whoever the AI companies try to use; but ASI alignment is not graded on a curve.
It should be taken as, “In PauseAI Global’s shoes, I’d be frantic to sever ties with PauseAI US and would probably change my name in order to get that done” not a paragraph-wise endorsement of all their given reasoning.
No, it is not “petty infighting”; so far as I currently know, PauseAI are okay people, and they have legit reasons for needing to sever ties with “PauseAI US”.
For a CDT agent to visualize different policy-worlds depending on whether it picks box B1 or box B2 requires that it violate “fixity of the past”, because it is visualizing a past fact varying with its own choices. This is basically the same as one-boxing in Newcomb’s Problem.
It may of course be CDT doctrine itself that is confused, rather than your interpretation of CDT. But the obvious interpretation of CDT’s doctrine of mixed strategies, as borne out by a large volume of discussion about Death in Damascus which bites the bullet of Death as a good predictor (or alternatively, of other agents being allowed to treat with you differently if you make the choice to consult a random-number generator), is that you may be subjectively uncertain of your choice of boxes and think you are choosing 50-50, but of course you are bad at randomization so a mixture of 2nd-5th order Markov models of you is better than 50-50 at predicting what you will actually do, even though you yourself are at 50-50 about what you will pick before you “randomize” and actually choose something.
I treat ratifiable mixed strategy CDT as standard CDT. My objection to CDT is “the standard sophisticated variant ratifies bad bets” not “no known variant of CDT ever makes a choice”.
I still hold out some hope that Carl has not gone full Evil, but then I also held that hope about Peter Thiel back when he started endorsing Trump, so.
Carl was a rationalist before there was an EA/rationalist dichotomy. After the EAs showed up and started being Modest, it started being clear after a few years that Carl was Modesty-seduceable and hence leaning increasingly OpenPhil over MIRI.
That’s an update for me if true. My recollection is that he was an OpenPhil intern-or-something.
Then I’m sad about it. I don’t know what else you want me to say.
Yeah, IDK what Carl is thinking besides something something Modesty.
Inference from “Suppose Leopold is an employee at OpenPhil and that other employees at OpenPhil are exposed to his RL signal.”
If intelligent life is spectacularly rare, we should look around ourselves and see an older universe, rather than the sort of young universe that no other intelligent life has had time to reach and colonize.