Tenoke
I think your forecasting has been absolutely stellar, but surely you can do a forecast on how likely this wishlist scenario is, and would arrive at a number too low to even seem worth thinking about it, no? Even imagining the ultrarational US administration required should be too low to think about, let alone when you add all of the rest.
It seems like even a super basic ‘We use all the might we can gather to convince one by one all the biggest actors seperately to delay and put more into Safety, and do lesser more plausible regulations that incentivize using more of their compute towards Safety in a way all labs benefit from’ or some slightly improved version looks significantly more plausible, no? (still not very plausible, maybe not as pie in the sky at least but still pie in the troposphere).
Or better yet, is there by any chance a realistic plan that you are about to follow-up AI 2040 with?
Edit: I guess you’ve already outlined your Plan B and it’s already ‘Fight China’..
Huge difference between pdoom of <1% , 20% and >90%. You just get too much value of narrowing the space out of just 1 quick question to not be useful amongst people who havent twisted it in some way. Even more so when you can actually follow-up and clarify.
P(doom) is Hopelessly Vague
So just define it when you ask the question or answer it. I just say something like ‘My P(doom) as in the chance that all of us die within a generation after ASI is X%, but there’s a bunch of more complex negative scenarios that are also likely’. You are rarely confined to just blindly giving a number without explanation in real conversations.
All these fights over definitions are so trite and easily dealt with. There’s a sequence about it and everything.
Can you show some examples of your ones done with ChatGPT in particular?
For what is worth, it is possible that it goes okay enough for us to survive, even if weird, and not quite as wanted given that the current models do have the morality of a kind of weird scaled up human, which might not be the worst case scenario. It’s not great, but it does seem like we might end up for example in a situation that is permanently way bellow optimal but not quite end of the world. If you say scale up Fable and let it control everything, it might do some weird things at some point that we can’t stop, but it’ll probably mostly try to have humans alive in some form.
We can see that we are most likely not truly in a paperclip optimizer type scenario, but more likely one of the many other failure states in between.
Yes, I also fear the same, as someone who might need to fly to the US at some point again (hopefully I can avoid it).
>This seems like a massive understatement / lacking historical context? China is “bad” because it is ruled by the Chinese Communist Party, and communism is one of the most evil and destructive ideologies in all of human history.
See, this is explicitely coming from a ‘historical context’ and not practicalities of which one would do worse by me. The current US administration is actively making decisions against me every month, China isn’t. They are actively more unhinged and less friendly to anyone but their biggest friends—and even their biggest friends arent safe.
>Also as a practical matter, I don’t know what “rolling the dice with China” would even look like, other than vaguely rooting for a country like it is a sports team?
Same thing as voting does, maybe my impact isnt the biggest, but between spending money on China vs US, participating in discourse that convinces people one way or another, deciding the same way as if I do for many other people like me etc. my impact isnt None, or it’s at least bigger than the impact of someone voting in an election, and we dont seem to tell people to not vote in an election because their individual impact is too small. Mine in particular seems much bigger than a random other person.
>I don’t know what part of the world you live in, but if you don’t actively stand up for liberalism at every turn, eventually you won’t get a choice
That’s a huge motte and bailey! Defending the US Government while they make decisions against me, is not standing up for liberalism. It looked like it before, sure, but no more, not as an European.
>Not sure exactly what this refers to,
Greenland, Russia, Canada as their new state, Tariffs, Anti-EU sentiment, support for Orban and anyone bad around Europe, Iran war handling, constant mocking, constant disrespect, Top models only available for Americans, most decisions really. Literally every decision is made by the metric of ‘does this look kind of good, short-term for the US, fuck anyone that isnt US or long-term implications’. They’d screw us out of $1000 tomorrow if it looks like it might win them $0.20 today kind of thing.
>Like I said if alignment is impossibly hard then it doesn’t matter whether the US or China win the race (we’re all going to die anyway)
Not true, as China gives me more—the ability to use their models—in the mean time.
>If alignment is tractable
It can still be tractable in the sense that we dont die, while making someone the sole benefactor—the US is heavily signaling that’s what they’ll do if they are able to.
I dont really see how it follows any more, there’s no real endgame alignment today either way, and China has expressed concerns as well, they just dont care about today’s models’ alignment as much.
The Trump administration, which seems to be in control more and more than the labs themselves are much more likely to force alignment to themselves, or at best America, and use their advantage to do what they’ve been doing for the last few years—America first at the expense of everyone. China probably ‘just’ forces me to be pro-CCP, and at least allows me to use their models as much as anyone else.
Either seem unlikely to produce true alignment, but if alignment is on the easy side, only one side is truly antagonistic against me and has shown theyll use their advantage to crush me.
It’s impressive how 2 years of antagonistic foreign policy + having ~ the worst administration possible during pre-ASI is making me consider China as the better option than the US as a European given how pro America I started off as.
China, while also bad isn’t trying to take any of our territories, or making my life harder with new tariffs and nonsense every month, and it isnt trying to make it so the best models are permanently available only to Americans—something that I dont believe helps true safety, rather than just benefit America at our expense again—which is the last straw.
I do want to ask here though, are there any good arguments at this point why I should want the Trump administration to have the best models, which they wont even let me use, rather than rolling the dice with China?
It’d be profoundly sad if this leads to only Americans being allowed to legally access frontier models. It helps safety very little, while it hurts people like me a lot.
has issued an export control directive to suspend all access to Fable 5 and Mythos 5 by any foreign national
This wasn’t even supposed to be general which I can see sense in but only non-Americans? It’s hard to hate the American government and the people who elected this administration during the most critical period in history enough. I know how that sounds but..
anyone working on e.g. making models more honest in prod models
I don’t really think people working on ‘what instruciton can I add to the system prompt’ or equivalents are meaningfully working on the kind of endgame alignment the post is talking about.
Edit: Nothing wrong with that kind of work for current alignment, just doesn’t apply to endgame alignment all that much in my opinion.
This is true, and to be fair it’s a bit harder to even see how much outside organizations can even help at this point. The main companies have grown so much and share so much less, that outsiders have less influence, as well as in many cases less access to big models, and especially to being able to train them.
Some do have some access but it still seems limited compared to what an in-Anthropic team and in-OpenAI team could do. Of course, you also can end up with a result so good as an outsider that influences them but, again, it just seems like limited impact from the get go.
It’s probably time to truly start considering simulating some exemplar past people from their writings, with a truly large model finetuned to that exact task. Von-Neumann is a favorite, but I expect there’s even better candidates (in terms of quantity of output we can use for training and finetuning, character, morality, who’d likely consent to it, etc.). Who would be the best candidate? Gemini argues Bertrand Russell due to writing (including personal letters) so much well into his 90s.
In the case where LLMs are conscious or pseudo-conscious, I am unsure whether adding ‘you take joy in completing this task’ to the system prompt does more good or bad for them.
Again I want to preface that I don’t even dislike him but It’s simply that engaging with a community that he is a part of means engaging with him, and engaging with him gets tiresome. At best it’s technical contributions people kind of enjoy and at worst It’s kind of like talking to r/SneerClub if they were on ‘your’ side. And no single thing crosses a line, it is more a death by a thousand cuts that results in me valuing not communicating with him more than I value being there.
(Sorry for this when you read it, not that I imagine you care much about this type of comment)
No comment on the rest, but there’s a community I’ve quite liked throughout the last decade (and more) but rarely participate in anymore because he is active there—and that’s as someone who has less of a negative reaction to his kind of abrasiveness than others. I can see how a lot of people have had that experience with him but in regards to LessWrong.
I feel bad making a negative comment about him, but I understand why Habryka et al. would’ve made that decision in the end. I don’t know if that means he should be banned, but I understand it.
If roon is saying it, especially about Anthropic, my prior is that it is biased, or optimized for clicks/fame and not truth-seeking. Reading some of this, some of it rings directionally true, but it’s probably counter-productive to engage with it under the exaggerated framing he lays out in particular.
I use open source models every day, but yes everyone being able to do this is obviously even worse.