I currently work for Palisade Research as a generalist and strategy director and for the Survival and Flourishing Fund, as a grant round facilitator.
I’ve been personally and professional involved with the rationality and x-risk mitigation communities since 2015, most notably working at CFAR from 2016 to 2021 as an instructor and curriculum developer. I’ve also done contract work for MIRI, Lightcone, BERI, the Atlas fellowship, etc.
I’m the single person in the world that has done the most development work on the Double Crux technique, and have explored other frameworks for epistemically resolving disagreements and bridging ontologies.
Even as I’m no longer professionally focused on rationality training, I continue to invest in my personal practice of adaptive rationality, developing and training techniques for learning and absorbing key lessons faster than reality forces me to.
My personal website is elityre.com.
Eli Tyre
Zach is adhering to a coherent notion of morality here. I think I agree that what he’s doing is antisocial (according to a relatively common notion of “antisocial”).
But Zach (I believe) disputes that being antisocial in this sense is bad. Indeed my impression is that he thinks the opposite—being prosocial in this way is the active source of a lot of harm and evil in the world.
Rationalists will propose a banned products store where you can buy dangerous things and say “if you choose to buy something that kills you there, it’s your own fault.”
The text of the very post you cite is arguing against that stance.
Saying “People who buy dangerous products deserve to get hurt!” is not tough-minded. It is a way of refusing to live in an unfair universe. Real tough-mindedness is saying, “Yes, sulfuric acid is a horrible painful death, and no, that mother of five children didn’t deserve it, but we’re going to keep the shops open anyway because we did this cost-benefit calculation.”
The point is NOT that “If you made your own choices then they’re your own fault”. The point is “there’s an actual cost of people who realistically will make these mistakes, who will be harmed, and we shouldn’t pretend otherwise, or slip into justifying it.”
This is a genuinely good point.
I’m not super clear on what the ideal response to this state of affairs is. I’m both confused and probably wrong.
But does “paranoia” not evoke the difference?
Should we stop attempting because it’s hard? Or notice cope and cover and correct it?
Not necessarily, but we should maybe be less gung ho about “trying to do the most good”.
It feels like there’s maybe a missing mood to a lot of EA and x-risk related efforts. Probably more people should be more paranoid about the impacts of their actions, and more willing to do ostensibly less ambitious things that are less fake to them?
It could be that we live in a world where it’s not so, and we must pick between unclear good + tight feedback, or the most good + poor feedback.
I somewhat dispute the framing here.
If the “most important” problems are intractable because their feedback loops are so bad, then our standards for “most important” are themselves mistaken.
It’s like Richard Hamming saying that in some sense, developing anti-gravity is hugely important problem (if you solved that it would have enormous commercial and humanitarian benefit), it’s importance is a mirage, because there’s no-line of attack on it.
It could be case that there are a bunch of problems that will have huge impacts on many people through the whole future, and also, in practice, it’s a bad bet to work on them, because almost no one can do non-fake work in an unaccountable domain.
This thought was inspired by reading this Ben Hoffman’s recent post, and more proximally by this Wei Dai shortfrom, which claims that approximately no one exhibits long horizon agency.
You might think that means “almost no one except these EAs and rationalists”. But, since it’s easy to use “optimizing for the long term future” as a cope to hide from accountability, I suggest that many (most?) of the people who appear to be exerting long-horizon agency are using that as a cover and less actually steering the future.
“Ambitiously” tackling “the biggest problems”, like AI risk, is often actually easier and safer than the alternatives. You get to feel cool and important and you’re bolstered by your EA student group friends who are excited about what you’re doing, and the lack of good feedback loops means that you don’t fail hard and completely and visibly (but learn and grow more).
There’s a sense in which selfish or limited goals like “make a successful dating app” or “make 10 million dollars” or “get 10% more housing units built in my local city” are easier than “reduce the risk of human extinction”. But in practice, the sense in which they’re easier is often fake.
EAs would probably do better to mostly do narrow tractable things that feel real enough that you could __maybe__ actually concretely do them.
If when you consider items above, they feel extremely daunting—if you feel a twinge of fear or resistance to the thought of trying to solve them, if you feel like you wouldn’t even know where to start—consider that “reducing x-risk” is in fact a much thornier and harder problem, where you have much less leverage over the thing you’re trying to change, and dramatically worse ability to tell if you’re making progress.Actually trying to make your startup succeed (or whatever), for real, feels real and scary, the in a way that it does not feel scary to be aiming for an abstraction like “having impact”.
“Impact” or “reducing x-risk” is too vague and too removed from feedback for anyone, including yourself, to ever really hold you accountable. So even if it’s extremely “ambitious”, it’s ultimately a safe thing to do.
Trying to “the most good you can do” means tackling the “most important problems”, which can allow you to hide from doing anything real or scary.(This is mostly advice to myself.)
I predict that a randomly selected MATS participant from cohorts in the past two years would more likely than not fail to pass my ITT about why AI poses an existential risk.
Very interested in running the experiment, if anyone has a way to source random MATS scholars.
Also, I’m not very excited about individual rationality training, generally.
The main blocker for human rationality is not that people don’t have the techniques or methods to be rational, it’s that rationality is anti-incentivized in many relevant contexts. When the incentives are towards getting the right answer, smart people, at least, organically develop most of the techniques they need to get the right answer.
I’m more bullish on things like prediction markets and prediction communities than rationality training per se.
I’m new to rationality, but I get the impression that rationalist training hasn’t been able to achieve individual transcendence either. “A Sense That More is Possible” was written 17 years ago, yet rationalists are not yet widely known for their evident auras of awesomeness, and CFAR declared “narrative bankruptcy” in 2022.
For what it’s worth, I was personally close to a lot of what was going on with CFAR and the early rationality project, and my diagnosis is more that “almost no one tried seriously at any point” not mostly “we tried hard and had little progress.”
Also, the handful of people who tried seriously, in my estimation, most do in fact have auras of formidability and impressive track records. But also, those people were unusually intelligent or capable to begin with, and it’s not clear how much to attribute to their personal practice.
I’ll probably write more about this sometime, and do a full postmortem, if the world doesn’t end soon.
This kind of consideration seems extremely important and action-relevant (certainly relevant to how I feel about my current professional-bets), and I’m surprised it only has this much karma.
I’d advocate for top-level-posting this.
And there isn’t a simple spell you can perform to—poof! - summon a complete ghost into the machine.
So, as it turns out, this is false?
The loop that does a LLM pertaining is very simple (admitting that you do have a bunch of complexity in the data you’re training on) and effectively does summon a ghost.
You could frame this as an extension of number 3, but an additional thing that foundations can do, but s-processes with rotating casts of recommenders can’t do, is active grant-making.
Foundations can decide that theres’s a specific project that they want to to exist, and then go out in the world and try to find the people who can do it, and offer to fund those people to do that project. That’s not something you can do, if you’re hired to be a recommender for 3 months and then go back to your other job, and you can’t make commitments on behalf of the grant-making org.
This seems like the same level of as other proposed alignment strategies like “maternal instincts” from Geoffrey Hinton or the “truth-seeking AI” from Elon Musk.
I disagree, on several counts. For one thing, constitutional AI is a technical spec, not just a vague idea for how things could go better? For another, it’s a technical spec that recruits an AI’s own intelligence for shaping the AI’s behavior, which seems like the kind of thing that you need in a scalable alignment scheme.
I’m not saying it’s remotely sufficient for aligning superintelligences. It seems like it isn’t for a bunch of reasons.But if you think people are wrong to be engaging with it, do you want to give some specific arguments for what’s wrong with it rather than just stating that it’s “obviously naive”?
(A) both view it as a good thing to “free people from work” and (seemingly inconsistently) continue working themselves despite adequate personal wealth?
I don’t think this is inconsistent at all.
I want to be free from work, in the sense of having enough money that I don’t need to work to pay rent. I will at that point, in all likelihood, keep working, but with more freedom to do only and exactly the work that is most important or meaningful.
I am uncompelled.
It sounds like you’re arguing “you want to make good decisions, therefore you act as if observer moments are sampled according to importance, because otherwise you’ll...waste time thinking about how weird it is that you’re in an important moment making important decisions”??
If the problem is that you’re going to “waste time” being surprised, just don’t do that. Be surprised, and do your duty anyway.
More critically, it seems like if you’re in an important moment making important decisions then it’s especially critical that you have an accurate understanding of your situation. That you’re going to “waste time” thinking about how weird your situation is seems dramatically less relevant to whether you’ll make good decisions than, for instance, whether the fish are moral patients or not.
Why are we privileging “wasting time” as the main way that you’ll make worse decisions?
Most of his social media posts and campaign actions
To add some extra detail to the model, it’s pretty likely that he doesn’t personally control his social media posts, but has hired someone to manage it for him.
On the one hand, if so, he did hire that person to represent him, and so its pretty reasonable to hold him accountable for that.
On the other hand, a person might be higher integrity than average, and struggle to hire people who are similarly high integrity, or who have a similarly good understanding of eg AI issues.
They no longer print your boarding pass if it’s within 60 minutes of an international flight, so you can’t go much lower.
Can you not checkin and get a digital boarding pass on your phone for international flights?
This is a much better statement of the AI risk problem than I’ve often given and that it resolves some qualms that I’ve had since trying to explain why any AI would be motivated to take over the world, in a recent talk.
Generally, there’s a bit of a leap in the old LessWrong arguments that any powerful AI would exhibit instrumental convergence—that claim depends on a kind of specific conception of what AI is or has to be, which Eliezer persuasively argues for, but which isn’t obvious, and is hard to convey to a layman. It is much less of a leap to say “any AI of this type, would try to take over the world if it could” and “notice that since LLM base models were invented, every single capability advance entailed making the AIs more like the scary version.” This seems to me like a more intellectually honest framing of the problem. Maybe not all powerful AIs we could invent are like this, but the ones we’re building in practice seem to be!
I also like that this frames the problem, appropriately, not as some speculative thing that might happen, but as a thing we’ve seen over and over for decades, and that we expect to keep happening. It’s just that as it occurs in more and more powerful and empowered AIs, we get more and more harmful versions of the problem.