Inference from “Suppose Leopold is an employee at OpenPhil and that other employees at OpenPhil are exposed to his RL signal.”
Eliezer Yudkowsky
EDT doesn’t vote if it’s heard any news about polls, since those screen off most evidence of your personal vote.
Does training to endorse CDT make the agent generally stupider? (Larger models with more inference time move more toward FDT behaviors and explicitly endorsing FDT, so it’s an obvious thing to check.)
Why are you asking AIs overwhelmingly sympathetic to FDT whether they favor EDT or CDT? Why are you benchmarking them on only problems labeled EDT or CDT and not Parfit’s Hitchhiker or whether to vote in elections, as discriminate EDT/CDT from FDT?
I rather expected from that time that Aschenbrenner would be trouble for EA, if they just couldn’t stop themselves from promoting Harvard blondes; but I did not say anything against him then on the basis of that conversation, and only marked him down as actively harmful once I saw him starting to oppose all good ideas during his brief tenure at FTX philanthropy. He did of course further go on to actively stoke arms races with China.
The relevance to the Wei Dai part is that Wei Dai seemed to me to be asking, “Aren’t you deciding post facto who is ‘EA’ and who is ‘rationalist’?” and my reply to Wei Dai is implicitly, “My memory claims that it is really really not hard to tell, in advance, early on.”
I remember the first time I met Leopold Aschenbrenner. He questioned me about why my timelines weren’t longer and finally said, “Why aren’t you updating at all in the direction of all the smart people saying longer timelines?” (aka: OpenPhil doctrine of 2050), “At this point you’re not even pretending to be rational.” I hesitated briefly, because I knew it would not end well to answer honestly, but there wasn’t any good ending past there; and then I told him honestly that I did not consider those people to be my epistemic peers.
I don’t usually consider myself good at reading faces, but wow, the sheer LOOK of contempt and scorn that crossed his face; before he turned away and walked off without another word, dismissing me from further consideration. I will never forget that moment—because of it being such a LOUD face and my actually being able to read it, not because it had any influence on my estimates, of course.
Amusingly, I have later heard from OpenPhil people telling me very earnestly that they think OpenPhil avoided groupthink on AGI timelines because there was dissent allowed inside OpenPhil. Well, sure there was allowed a pretense of dissent inside the OpenPhil Overton window, unless you said that you weren’t updating your opinion in favor of deferring to the people whom OpenPhil considered locally high-status enough that deferring to them was rational, and then you would get frozen out and dismissed from all consideration. I can only imagine the Look that Aschenbrenner would have given somebody who’d said the same thing while being an intern rather than Yudkowsky, and while other OpenPhil personnel might’ve been more polite, I don’t really see that intern being hired.
OpenPhil and EAs may consider themselves to be terribly modest and outside-viewing, but of course that’s all an utter sham given their actual skill levels, and they have no concept of what it would actually look like for them to be humble.
It doesn’t surprise me in the least what Aschenbrenner went on to do. But yes, he is definitely a “Modesty!! Outside view!!!” sort of person, or knows that is what he is supposed to perform. I don’t expect there will ever be any contradiction that he understands, between “modesty”, and whatever the hell thought he decides to take into his head. That would take skill, and he is not aware that he is unskilled.
Luke Mueuhlhauser is actually modest. I don’t think he would ever blow up a hedge fund. It also didn’t surprise me at all when he moved on to OpenPhil, and I was glad for him, because I knew MIRI had not fit him and he would be happier at OpenPhil, and Luke had worked hard at MIRI and maybe saved the organization while he lasted. In related news, EAs picked Aschenbrenner rather than Mueuhlhauser to run their hedge fund.
Carl Shulman and Luke Mueulhauser were very Modesty-Argument / “outside view” people, which tightly corresponds to getting sucked in by the EA cluster once it exists / makes them an offer.
Thank you for the pointer; I have now tweeted about this and made the uncertain guess you did not particularly want a “H/T Wei Dai” about it.
“What if we’re in a simulation that I can discern to be visibly unusual” discourse popping up again.
And whatever the utility function of gods, you imagine (1) there’s not much they could possibly find any more entertaining than that, and (2) there’s not that many other possibilities for things that would be equally entertaining?
The conceit of imagining yourself to be Top-Tier Entertainment for a god. Why, there’s not just many possible things to imagine, so why wouldn’t it imagine you, amirite?
Do you believe that God intervened in the 2016 election? All these arguments carry over identically to God.
Generalized atheism rules out “inaccurate simulation”-ism.
I’m renaming “XOR Blackmail” to “XOR Demand”. Does anyone want to object to that?
One can try to build AIs in correspondence with deontology (corrigibility), consequentialism (value-aligned sovereigns), or virtue ethics (worthy successors).
Corrigibility is of these the easiest to define, easiest to test, easiest to build, and the easiest to check for the first worrisome signs that not all is going great—though of course if you wait long enough in capabilities escalation to test, everything will look great to you.
Any sane person who was not allowed to back off the problem would try for corrigibility, by far the easiest of the three.
Worthy successors are the hardest to build, the least well-defined, the hardest to verify that you are doing correctly; the easiest place to leap on little local events that you can convince yourself are signs of hope, because you have no framework to tell you what more is required; hence, the craziest thing to attempt, beyond even a nice Sovereign ASI. So of course all the fools, having failed at corrigibility, will convince themselves they are building worthy successors instead; migrating, with the inevitability of an amoeba following an agar trail, to wherever it is hardest for fools to be persuaded (with a fool’s demanded certainty) that whatever is currently happening is not all according to their plan.
On “gendertropes” in dath ilan
Or legalize childcare. There’s no need to subsidize it or exempt it from taxes, just make it legal for anyone I trust to sell me some childcare in exchange for money.
If I were doing it all over, I would be much less terrified of something called “cultishness”, would have installed something more like an accreditation or licensing system to be considered a member in good standing, and then a requirement of being in good standing would be that you were either raising kids or contributing money to those who did, located somewhere it was legal to build housing and where they could have houses in close proximity. (Some huge amount of modern pain of childrearing is downstream of having no burden-sharing on babysitting and nowhere to put a bunch of kids where they can play together instead of with you.) I expect that’s how Civilization does it in dath ilan; you either make babies yourself or you fund the sort of babies you want to see in the world.
To state the obvious:
Many of the same people who publicly make a great fuss of this don’t seem to care that Nvidia is selling GPUs to China by way of Singapore. Their purpose is to drag up China as an excuse for continuing what they are doing. If China has more GPUs, all the better for them in the future; China will be a greater threat that they can use to continue what they are doing.
Their model hasn’t chosen an explanation like “China will play the role of the movie bad guy at any cost to their own selfish interests”. They aren’t reasoning at GPT 3.5-level, only GPT-2 level. China is a scary word-vector, so they can argue that if the USA tries to slow down against extinction risk, then China won’t slow down, because that’s a scary thing so “China” is a word-vector that the scary thing can be attributed to.
Yeah, IDK what Carl is thinking besides something something Modesty.