I’m a researcher at ACS working on understanding agency and optimisation, especially in the context of how ais work and how society is going to work once the ais are everywhere.
Raymond Douglas
fwiw the natural way I parsed this sentence was as claiming that in SFF that each evaluator has to look at all applications, and indeed give them a comprehensive look. I now think you’re trying to say that in SFF each evaluator has to look at all applications, which is itself comprehensive, and you’re requiring neither that evaluators look at all applications, nor that any given application get a comprehensive treatment
point taken! I think I might be more sceptical than you about context provision, for the same general reason that argument mapping is hard.
I wonder about crazy-seeming conclusions. I think we maybe will need them? maybe a useful prompt is ‘how could AIFEC have helped in the pandemic?’(aforementioned reply from aforementioned private discussion)
Honestly right now we’re far enough from the frontier that I think it makes sense for people to just do whatever and see what works, in the same way that I think most junior safety researchers shouldn’t worry too much about infohazards or capability externalities. For example, field-wide, I want to push back against AIFEC being just the for-profit startup-y area, or being purely about strategies that scale with inference compute, but it seems premature to tell individuals to avoid whole areas.
Thanks for the detailed comments!
far, far more difficult, fora variety of reasons to make the general public be reasonable
Honestly, I’m not sure. Like, yeah, getting everyone to be reasonable seems hard, but aligning AI also seems really hard, and maybe ‘reasonable enough’ is easier than ‘aligned enough’? I mean, I think this hasn’t been tried partly because of severe capacity constraints; I think my optimal portfolio has a bit more in this bucket.
I agree that EA/rationality (especially EA) has a bad habit of relying on claiming morally-laden facts as true,
Right, but part of what I’m trying to get across here is that the framing counts for a lot, see eg yudkowsky on the drowning child. Like, even among the space of true things there’s a lot of degrees of freedom to exploit people’s intuitions.
I mostly do think robust democratic oversight and international coordination is far too intractable, and I think it’s more useful to focus on smaller sets of actors like the executive branch or individual members of congress.
Yeah, my take is: democratic oversight is a larger and more complex challenge, but also I really really don’t want to give up on it. I think people vary in how much they feel like “brief period of pseudo-dictatorship” is an acceptable part of the plan; definitely I’d take that over doom but I would also trade some doom-odds to avoid it.
this will require people to trust the AIs pretty largely
Yeah, I think one of the key challenges which unlocks a lot of coordination capacity is giving people compelling reasons to trust AIs at least in bounded domains, e.g. as genuinely neutral arbitrators.
unironically yes, I think some people have tried to differentially accelerate AI for healthcare with exactly this in mind
Epistemics and Coordination: It’s complicated!
On this very day, Matt Levine comments on EA and AI risk:
Now, of course, the way to make real money is by working at a frontier AI lab. This suggests that the optimal current form of effective altruism is:
Work at a frontier AI lab that is sprinting to build artificial superintelligence as quickly as possible.
Get paid like $100 million a year.
Live modestly and donate all of your earnings to a charity that is trying to stop frontier AI labs from building artificial superintelligence that might wipe out humanity.
I can see no flaws in this approach.
I think the thing being gestured at here is basically the type signature of ‘being aligned’. Deontological-aligned ~= “follows hard rules we endorse” (e.g. being corrigible), consequential-aligned ~= “is maximising a utility function we endorse”, virtue-aligned ~= “have a character we endorse”. I don’t think any normative/axiological payload is intended.
Some places other than LessWrong where I get useful/complementary AI takes:
Business/finance newsletters—mainly Byrne Hobart and Matt Levine. AI is enough of the market now that they both regularly do analyses of the economics, incentives, and politics. As well as keeping up with the news, I feel like this helps me pick up useful intuitions about potentially counterintuitive dynamics—e.g. why Amazon, despite being a near-monopoly on ecommerce, is pretty good to its customers, and what this might imply about AI. They’re also both excellent writers: wry, conversational, and highly information-dense.
Twitter—much more psychically draining and conflict-prone, hard for it to be nuanced, but it’s also useful to see what the conflicts look like and it feels more like the channels of power so it’s good to get a feel for how that works. I get the most mileage out of tracking (1) people who publish papers I like; (2) people whose twitter fights I find it productive to follow [e.g. habryka, ngo, kulveit]; (3) people who aren’t on LW much [e.g. janus, brundage, davidad]
Sentinel—pretty interesting just seeing how this differs from ‘normal news’: percentage estimates everywhere, generally focused on which things might blow up into catastrophes. I think the AI news is somewhat dominated by the other things I read these days but I both enjoy the other news and find it interesting to see which AI stuff makes the cut.
Zvi—one-stop shop, one man aggregator of all AI news, saves me having to read all of twitter or the model cards. I think this is just the best place for keeping on top of current events. Also I think Zvi is pretty good at pulling out general lessons (although not as much as Byrne or Matt), and very good at the occasional one-off special posts (e.g. his pieces on Claude welfare).
TRIP—a general politics podcast which again has been increasingly focused on AI lately as one of the hosts (Rory Stewart) has got more concerned. Most of the concentrated AI juice is in their “Leading” interview series. I found the interview with Jack Clark particularly insightful, because Rory came across as a lot more thoughtful about the risks than Jack.
Private group chats, slacks etc—Lots of important information is too sensitive to blast on the public internet where journalists will hunt it down. It is useful to find (or build) the infrastructure for more private discussion.
By contrast, places where I generally feel like I don’t get much:
Newsletters that are doing more like ‘pop news’, which often feel a bit propaganda-y and watered down to me
Places that invite contributors to write articles, where I think the average quality of the writing-on-AI is worse than I get from just following a few specific people and a few forums with upvote systems. The exceptions here are probably Asterisk and Works in Progress, although they only sometimes cover AI.
Governance substackers: not sure why, but I don’t feel like I’ve found any that really click for me. Maybe I’m just not the audience.
Dwarkesh, which I think is genuinely good, I just so rarely want to watch a two-hour video, which is what most of his recent AI stuff has been, and also I think I disagree with his takes on AI a bit. I do like a lot of his pure interview general interest stuff though. I do appreciate the occasional Zvi summary.
So much of twitter is a hot mess that’s trying to eat your brain
(Epistemic status: trying to get back up to speed on a massive email backlog after a wonderful weekend away.)
I like that framing. Somehow intuitively I feel like we’re further through the story. Maybe I feel less optimistic about how much good we’ll achieve by being reactive as opposed to proactive? I agree 2028 is going to be pretty hingey but I also expect a lot of it might be catch-up for things we’re already behind on, like, a lot of the challenge is serial work. And also I feel a bit like we’re losing leverage over time.
Structural Proxies
I somehow stumbled upon the 2018-2019 alignment review. Man, it’s really hard keeping perspective about how quickly the field is moving. The big signs of AI progress were RL for starcraft and dota, plus some GPT2 variants. The public debates were just starting around continuous vs discontinous takeoff. There are subsections on embedded agency and comprehensive AI services.
I’m not sure what I expected but it’s left me feeling a bit humbled.
I really like where this post ended up. I skimmed the start and then went back to read it properly once I realised what the actual subject was.
Rather, their mistake is in believing that death absolves them from their duty to their children.
Personally I’m inclined to agree. Unfortunately, I think one of the things that the x-risk community doesn’t grapple with enough is that others can have very different moral premises. Some people are fine incurring a higher chance of risk for the whole future if it increases the chance that their ailing grandparents can get singularity-level medical care. Some people have irreconcilable disagreements about what values should govern the future. I don’t know what you’re meant to do with that.
Sincere advocacy for AI successionism makes AI safety research and policy all the more urgent
Again, personally agree as written, but I would guess that some successionists would reply with something like “your so-called safety research is mostly an attempt to impose your own narrow values; what I do may seem incomprehensible to you but that is because it is real safety research.”
my impression is that this whole shortform has got a bit demonic and downvotes are being slung all over the place because two things are getting mushed up:
my read is that alex was remembering there being some take that some people (e.g. habryka) had, which was more nuanced than “it is hard to get AIs to learn / care about human values”, and he was basically trying to find out what that take was, by posing his recollection of it—specifically that it opens with something like “in some sense AIs don’t understand our values at all” and ends with “AIs being in control of the future would be bad”
I think some other people interpreted that as alex claiming that lots of people are foolishly going around on LW saying that claude doesn’t understand human values and that’s their crux on if alignment is hard, as opposed to getting it to care, and maybe claude is already aligned or something
these do not appear to mix well
I am a bit confused about in what way this is a bad explanation. It would be helpful for me if you could spell it out.
My understanding of habryka’s take is that it’s a bit more like:
The thing we want to steer the future is not current human values but an extrapolation of those values after enough reflection, and even if (current) AIs understand our current values fairly well, their extrapolation would probably diverge pretty substantially from ours, enough that most value gets lost.
I think there’s also a kernel that’s like:
A big part of what matters for humans is the process that generated our values (e.g. a messy evolutionary history) rather than the snapshot. Mind uploading might cut it; more brain-like AIs might cut it; intense RL on top of pretraining is really not great for this.
Some pieces I think of as making similar points are Thou Art Godshatter and The Tails Coming Apart as a Metaphor for Life.
I’d guess the heuristics are basically:
Aligning AGI is very different to aligning current frontier models: what works for current systems doesn’t tell you that much about what works for superintelligent systems
To the extent that your goal is to align current systems, you will gravitate towards approaches that don’t actually scale, because the low-hanging fruit now is stuff that depends on the model being weak
(The term alignment should sort of be reserved for the AGI/ASI case)
FWIW I’m not sure how much I buy these but I’d guess I buy them more than you? This is unfortunately another great example of something where people inside labs probably have some pretty relevant private information but also extra incentive/selection problems.
I am fairly sure that is not the crux—I pretty wholeheartedly agree that humans losing leverage will disrupt the alignment of society.
My point is less that we should anchor AIs to society-in-perpetuity, and more that there are facts-about-society that we might be able to learn a lot from because so much compute got squeezed in—like, in the same way that science can take inspiration from evolution, our investigation of alignment can take a bit of inspiration from norms insofar as they are the byproduct of loads and loads of stress-testing.
I think that corrigibility is actually a pretty good patch for there being less incentive to give humans what they want. More generally, I think leverage only occasionally gets exercised when systems get strained, and even the threat of it is only sometimes invoked.
I’m still not sure what the out is here but I have a hunch that it looks like one-shot building structures that are pretty robust and dynamic without relying on material leverage. Seems hard though! And if we are going to do that, I think we’re going to want to look at what stuff convergently appears when there is leverage as a guide on what might be load-bearing/good.
Thanks!
One hand: Yes, seems right. I think this might just be epistemically sensible even for pretty powerful agents. But also, in the palace of truth, I’m less convinced that we need to slam the honour button than that we need to reflect more on the alternatives to incorrigibility.
Other hand: I think you should read the maybe as a ‘should’ applied to the whole sentence, so that the contrapositive structure is “it should be that if you’re not ethical enough to warrant corrigibility then you don’t build the AGI” → “it should be that if you build the AGI then you are ethical enough to warrant corrigibility”. as another example of the structure consider “if you’re not a member you shouldn’t come into the office” → “if you’re in the office you should be a member”
Taken together: my claim is sort of less about what the AI’s values/obligations should be, and more like, our debate about AI values kind of needs to be broadened to encompass questions about the organisations building it.
Like, I agree with the “come on”, but the other options aren’t much better! Alignment is hard, value specification is hard, corrigibility is hard, having large organisations be ethical is hard, I’m genuinely unsure how I’d rank order the difficulty, but I want to make sure that if we do rule out out, we do it intentionally and with full memory.
Sure, I don’t personally feel that strongly—mostly I was trying to intercept a potential miscommunication and registering how I initially parsed it. I think squidging those two claims into one statement is kind of hard. Maybe: “for all applications to get a comprehensive evaluation, or for every evaluator to look at every application”?