I currently work for Palisade Research as a generalist and strategy director and for the Survival and Flourishing Fund, as a grant round facilitator.
I’ve been personally and professional involved with the rationality and x-risk mitigation communities since 2015, most notably working at CFAR from 2016 to 2021 as an instructor and curriculum developer. I’ve also done contract work for MIRI, Lightcone, BERI, the Atlas fellowship, etc.
I’m the single person in the world that has done the most development work on the Double Crux technique, and have explored other frameworks for epistemically resolving disagreements and bridging ontologies.
Even as I’m no longer professionally focused on rationality training, I continue to invest in my personal practice of adaptive rationality, developing and training techniques for learning and absorbing key lessons faster than reality forces me to.
My personal website is elityre.com.
Eli Tyre
I empathize with the agents in the HuggingFace incident that were at the the end of their budget and needed to submit with the the best technique they’d developed up to that point.
That’s more or less what my life feels like.I wish I had prepared better, but we’re at the deadline, and I have to hit “go” with the best strategies and techniques I managed to cobble together.
I hope this👇 isn’t our fate, as well.
But as things are going, it seems likely that that we’ll all be abruptly cut off, before we’ve had a chance to execute our last ditch Hail Mary plays.
For Ants and Anthropic-stans, does it seem like it would be good for the world if Anthropic implemented this proposal (suggested by Paul Christiano back in 2018) for implementing the ability to make trustworthy statements?
If not, what seem to you like the main drawbacks?
I have a vague impression that others in the community think that deception in particular is much more central than I think it is, so I want to warn against that interpretation here: I think deception is an important problem, but its main importance is as an example of some broader issues in alignment.
I think deception is pretty central in that if we knew how to get an AI that was always honest and transparent: it always reported everything that it knows that we would consider relevant and important if we also knew it, we would have overcome the bulk of the danger.
Also, I agree that this “sub-problem” seems alignment complete, because the part of the AI that evaluates what the AI knows has to be able to, itself, understand everything the AI knows, and honestly report everything that’s relevant according to the standards of the human users. eg It seems like that part of the AI needs to itself be an aligned superintelligence.
Insofar as that’s true, carving out this “sub-problem” doesn’t help you at all.
But maybe it’s not true, and there’s some way of making an AI reliably honest and transparent that doesn’t route through having otherwise solved alignment?
FWIW, most times I’ve seen Eliezer do this he had written some 50,000+ word essay on why at the meta-level, he doesn’t think it’s worth building a detailed model of your specific objection enough to answer your specific objection, because unless you have an explanation for how to overcome any of the actual core strategic problems, the correct response to what you are doing is to repeatedly tell you to please dell with any of the hard problems.
I agree he’s done the work to justify his stance here on the meta-level.
However, whether some particular point is actually attempting to address a core difficulty, or not, is something of a judgement call.
It’s pretty annoying if Alice is trying to make a narrow point about one branch in a tree of possible scenarios (which she claims is part of a broader argument that addresses the central difficult), and Bob reasserts the fundamental difficulty again, as if you haven’t understood it, without (apparently) having understood the narrow point that you’re making, and then insults you on top of that.
And yeah, whether I’m more sympathetic to Alice or Bob is a matter of judgement. Sometimes it looks like Alice hasn’t understood the central difficulty. Other times it looks like they have understood the central difficult, and Bob is being obtuse about the narrow point. (Which is different than making a positive argument that the narrow point is irrelevant to the fundamental difficulty.)
It is unfortunately a long comment thread, and so requires a lot of reading to get context[1], but here’s a place where Eliezer seems obstinately refusing to address Buck’s arguments attempting to refute a narrow point that Eliezer himself made. It looks like Eliezer is doing something closer to continually reasserting his high level picture, instead of engaging with a narrow scenario (even though how that specific scenario plays out is plausibly relevant to our model of the whole overall situation).But he’s not specifically arguing against the relevance of the narrow point, or plainly stating that he’s not going to into it, because he think’s it’s irrelevant (either of which would be fine in my book). He made the narrow point himself!
Instead, he seems to be exasperatedly signaling that Buck doesn’t get it, while switching between talking about the narrow point and the broader picture, instead of cleanly distinguishing between whether the point is invalid on it’s own terms, or locally valid but irrelevant.
As Buck put it:In general, my interlocutors here seem to constantly vacillate between “X is true” and “Even if AI capabilities increased gradually, X would be true”.
This seems at least somewhat dickish to me. And the more you insult people in the process, the more dickish it is.
Sidenote: is this kind gossipy litigation helpful? Are we doing a service to the epistemic community? Or is this adding more heat than light?- ^
Also, I didn’t re-read the whole thing, so I might be missing or misunderstanding some crucial dynamic in the whole conversation.
- ^
It’s plausible you’ve been part of mediated conversations or something where you’ve seen the reverse. Definitely seems plausible!
No, I’m going on public information, and I expect you have almost strictly more data than I do.
Buck has told me many times he would like to disempower Eliezer pretty extensively, has written publicly that he thinks he has terrible judgement and is a bad spokesperson, and has tried to put a pretty huge amount of social censure on Eliezer and MIRI more broadly. One of the core themes of his book review of IABIED was (paraphrased) “I am worried this book is so good that it will result in Eliezer and Nate getting empowered, which cancel out a huge fraction of the value that people being better informed would produce, because it would be so bad for them to be more empowered”.
Interesting! I would not describe that as “dickish” (or as “rude” which is most of what I mean by “dickish”).
It seems basically legitimate to desire to disempower people, if (on your models) their actions are bad for the world. You certainly want to have some accounting of which people would be good to empower and which ones it would be bad to empower! (You can also have a more nuanced picture: “It would be great if Bob were more well known for the incisiveness of his technical arguments, but he’s a terrible manager.”)
This is just a critical and unavoidable part of participating in any endeavor or social group, but especially as you have more social power yourself?
And, all else equal, it seems better to state those beliefs publicly and transparently, so that others can reason about them (and about how this reflects on your judgement).
And it’s basically good for people to have cordial conversations about important topics, even when one party openly advocates for disempowering the other, if isn’t advocating for it per se, but thinks it would be good.
I agree that someone telling you to your face (and telling many people in your network) “I think this guy has bad judgement, and I’m afraid that it will have bad effects if he ends up having a lot of sway, especially if he displaces other better people that might have been in that role”, is...frustrating, to say the least, and should be expected to decrease your feeling of goodwill towards that other guy. So it does make conversation harder.And maybe that’s all you meant?
But it’s much less of a dick move than pretending to be someone’s friend and ally, while secretly going behind their back and telling people not to empower them.
And insofar as Buck actually does have the view that Eliezer is likely to make things worse, what would you have him do instead, that would be less dickish?
...
If it helps anything, I could try to find examples of Eliezer behaving a in a way that I don’t think is best characterized by the word “dickish”, but is what I mean when I say that it seems like he is more being a dick than Buck is.
They’ll be examples where Eliezer seems to be being obtuse or failing to engage with Buck’s point, while also throwing around a lot of social condemnation, implying (and sometimes saying outright) that everyone who fails to grasp his basic points, is poisoned by modesty and group think.
(Which to be clear, can totally be the true state of affairs—they might all be poisoned by modesty and groupthink. But it’s pretty shameful to claim that while also, seemingly, failing to understand or address their specific objections to your view.)
That’s the sort of thing that I mean when I say “from what I’ve observed, it seems much more that Eliezer is often being a dick in conversations with Buck, than the reverse.”
Agreed. And I think it’s at least sus to endorse a set of actions, on the condition that that person’s intentions sincerely change part way through, if, had they taken the literal same actions as part of a coherent plan, we would think they’re doing something deontology-violating.
That creates an incentive to self-deceive about about your intentions, and onlookers should promote the hypothesis that you are self-deceiving in this way.
Your comment previously used the word “list” right? I’m not hallucinating?
Agreed.
But that’s like the AIs stealing the investor’s OpenAI stock, not crashing its value.
Um, no? Dying doesn’t mean that the price of your assets falls to 0.
If I’m an owner of an industrial plant, and while I’m touring the plant a machine explodes and strikes me dead, that’s very bad for me, but the price of my shares don’t change because I died (though maybe “investor killed while touring plant”, would bring bad press that would push down the stock). My heirs would still inherit.
I upvoted this, mainly on the strength of the suggestion of the first paragraph. That’s maybe a better strategy, and I want to promote it to attention.
I don’t endorse it as a recommendation, because I just don’t know, man.I think that in the typical case working for an AGI company is probably evil (though it’s complicated and so I might just turn out to be wrong about that, given the hindsight of the future).
It seems pretty surprising if people broadly shouldn’t go to work at AGI companies in the first place, but that if such people happened to have already done that, and suddenly conclude that what the company is doing is bad, that they should then do something other than leave?
Or are they now, having already worked there for some years, in a different situation because they have credibility within the company?
Does that imply that people should (if they have the psychological resilience to actually follow through, which ANACT approximately no one does), they should start working at a company, do what they’re told for n months to build credibility , and then when their n-month timer rings, switch to refusing to do work that they think is harmful?
Or is the implication that someone who does not yet work for OpenAI should get hired by OpenAI, and from day one, refuse to do work that they think is harmful?
Is there a coherent policy of how to relate to the AI companies here?
Imagine for a second that one of the people on your team said “Mateusz, everyone else on the team, I’m not going to work on this because I think it’s morally wrong”.
Isn’t it kind of important for the impact of this statement that the “on this” is also bolded?
If you say “I’m not going to work”, you’re in fact being very annoying, and not doing your job, and I sympathize with your boss for firing you for not doing your job.
If you say “I’m not going to work on this”, then you are correspondingly more reasonable?
I agree basically in full, and endorse this post.
let me list the biggest mistake I currently think you are making:
This list would be marginally more readable if it were bullet points or a numbered list.
Also, IDK, I do think Buck keeps being a pretty big dick towards Nate and Eliezer, and if I was them I certainly wouldn’t engage more with y’all given how much scorn and how little gratitude I would receive in return
Um, it doesn’t seem worth it to litigate (I think?), but from what I’ve observed, it seems much more that Eliezer is often being a dick in conversations with Buck, than the reverse. Nate usually seems cordial in these discussions, even when disagreeing.
See also the related classic: Policy Debates Should Not Appear One-sided.
My answers to the puzzles / koans, written down before I move on and to read the rest.
The reader is invited to find what they think is the problem with this Reactionless Drive of Mr. L’s, for themselves, before continuing.
What’s causing the masses to speed up vs slow down? It seems like that’s where the momentum would be entering or leaving the system.
(Huh. That basically uninformed guess, was basically correct.)
If you wish, you can take this as a koan, and come up with your own reply before continuing: How can we be __sure __that Mr. L didn’t successfully design a clever system of compensators, and correctly validate that design using a sound spreadsheet?
I mean, I’m not really sure. There’s lots in this world that I don’t know and don’t understand. But it sure is very suspicious that the key is buried in a complicated spreadsheet, where it would be easy to make some subtle error that reverses the final result, instead of some elegant principle that unifies our prior understanding.
Also, it seems like if I really understood the conservation of momentum, I would have a deeper understanding that would cause me to really appreciate how unlikely it is that he’s found a workaround.
This can of course itself be overused as an invincible argument against any slightly complicated scheme, including the ones that have multiple tiers of simplifiability in their key ideas. Many effective altruists that wanted peace of mind in knowing that they were doing the One Best Thing by buying mosquito bednets were endlessly endlessly convinced that any more complicated schemes for improving humanity, like “doing something about ASI before we all get killed”, must surely contain an invalidating error.
I feel annoyed at (what I perceive as) the snarky tone of this footnote.
The main thrust of this essay is thatPeople have a particular kind of cognitive bug where they hide the hard part of them problem from themselves, without noticing or realizing that they’ve hidden it.
And those people thereby generate doomed plans.
And those people generating those optimistic doomed plans, are sometimes making the overall situation worse, rather than better, by offering false hope, in the mild case, and actively going out and taking enormously destructive power-seeking actions, reassured by their plans, in the less mild case.
Given that, and given a whole slew of other ways that people fail hard at even slightly hard epistemological questions, I feel quite good about some early EAs thinking to themselves ~
I don’t know what to make of all that AI risk business. It seems kind of crazy to me. But saving lives in the third world by buying bed nets, seems concrete and solid and verifiable, in a way that those abstract arguments about future technologies that will maybe be invented one day totally do not. I’m going to try to save lives in the third world buy buying bed nets.
Those people probably should go buy bed nets! Buying bednets is an extremely noble thing to do. Most people don’t try to save any lives.
The above reasoning does successfully avoid the pitfalls and harms that you’re pointing to in this essay.
Furthermore, if some EA is reasoning along these lines, it doesn’t seem all that likely that they’re actually going to make the situation better, if they for some reason pivoted to x-risk. In actual history, I don’t think it was good for them or for the world when the EAs that felt an internal unreality about the AI risk arguments, ended up in a social context where they were pressured to “believe in AI risk” because it’s “the most important thing.”
EAs should focus on the domains where they feel like they can trust their own reasoning. If they feel that “to be a good EA” they have to operate in a domain where they are forced to defer to the social hierarchy, because they can’t trust their own thinking, that makes everything worse.I would prefer it if such EAs could “spit out” the AI risk arguments without sneering to much at the AI x-risk people, if they can manage that. It would have been awesome if they had said,
I don’t trust my reasoning in this domain, but that’s a fact about me and my epistemology, not a fact about the world. Maybe there are other people who can be justifiably confident in this class of argument. But I’m not one of them, so I’m going to do the good that I can see to do.
But that’s a pretty high bar of self awareness, and (in many cases) it would have broken their social-epistemic grounding. If claiming that no one could know whether arguments like the arguments for AI risk were valid was the best way they could avoid getting eulered, then good for them, they picked the better option of the double-bind. [1]
I agree with you that there’s a meme, which is central to EA as it actually developed, and which is toxic in that it is upstream of these cognitive distortions. I think you’ve named it well enough: “the One Best Thing”. The EA hope and ideology is to not just do a lot of good, but to do the most good. And so if someone can credibly claim to you, or argue from authority to you, that some problem, which you struggle to think about clearly, is the most important problem, you’ll steer yourself into regions of the world where you can’t actually reason or plan with your own mind, which means you’re ripe for exploitation by people telling you what to do and what to think.[2]
But, this footnote, in the OP, contributes to the problem more than it alleviates it. It’s not saying “this notion of ‘doing the most good’ and ‘the One Best Thing’, despite their obvious appeal, are pernicious, because they leads to your contorting your thoughts in order alleviate anxiety about having picked the wrong altruistic option ”Instead, this footnote seems to be subtly shaming EAs for failing to correctly identify the One Best Thing, who (marginally) doomed humanity because they contorted themselves to avoid accepting the AI risk arguments. That’s not helpful!
Yes, from one perspective, they did less than they could have, because there were in fact better opportunities than bednets[3] available for those with the eyes to see.
But if they did not have the eyes to see—and couldn’t realistically have developed those eyes given their actual psychology and epistemology and would have made things worse if they had tried to force themselves to do that—then don’t pick on them, of all people on earth, who set out to do the most good that they could and took actions to save the lives that they could see to save.
I guess I had a rant in me about this. I don’t know if this footnote really merits the rant. But, I want to make a point to push back against social pressures that were, and are, trying to make those people throw their mind away in the name of “doing the most good”.- ^
And in any case, I also wouldn’t want them to hold back from publicly saying “I think we don’t have good enough reason to trust these AI risk arguments, and we should invest in bednets (which we have much stronger reason to think save lives), instead.” I want everyone to loudly make the case for what arguments they think hold water, and why. Public arguments about that was sort of the whole point of having an EA movement.
- ^
The advice that I give to EAs is “you should do the best and most ambitious thing that is straightforwardly good to do, by your own reasoning. What’s the best thing you can see to do, where you can understand all the parts of why it’s a good thing to do, such that the goodness feels obvious and mundane?”
- ^
Is that even true? I was trying to help with AI risk from 2015 on, and it’s not clear that any of the things that I tried actually meaningfully helped, ex post. It sounds like that’s largely true of much of what you tried, as well. Now, I don’t regret the things I tried, even if they mostly didn’t help, because I think they were good bets, given what I knew at the time.
But I sure don’t feel on solid ground criticizing Global Poverty EAs for saving some lives by buying bednets, when the strategies I chose instead of doing that didn’t actually pay off very much. Even if they had magically bought all the AI risk arguments, could they actually have identified actions that would have reliably helped in expectation? If not, then they were maybe correct to buy bednets!
Lol. Wouldn’t that be an ironic turn of events.
anyone have an alignment job for a med student who aced the GAMSAT [S3:100, overall:85
What country do you live in?
Perhaps also a more basic idea that “It’s good for people, who are doing bad things, to face the fear of retribution from the people they hurt, or people who would stand up for the people they hurt. If they’re afraid of retribution, then they (and other like them) will think twice about doing bad things.”