I think the most likely policy-shaped path around AI danger on the current trajectory is AGI-originating pressure to be extremely careful with superintelligence-enabling or otherwise paradigm-breaking algorithmic innovations (this is under the assumptions that make me expect AGI to by default very likely appear by 2028-2032 without any new innovations, but remain unable to make progress towards ASI quickly). If the LLM/RL AGIs need (or are lucky to get) until 2040-2045 to invent superintelligence (unhobble their speed of learning deep skills or of contributing to cultural accumulation), possibly via industrial explosion that generates much more compute and makes the next-model building loops much faster (I’ll be publishing a post on this in a few hours), then they’ll be in a position to at least have thought about the issue for a deep cultural time, so they might be able to be very persuasive for the right reasons, even if humanity retains control and could just RL them out of any such notions.
Vladimir_Nesov
As a mostly random reminder, when an author deletes their post (or moves to drafts, I’m guessing), permalinks to the comments under that post stop working. As someone who frequently develops ideas in random comments under random posts, and happens to find it easy to recall their locations (at least for a few months) to then look them up or link to them in future comments, I see this as an annoyance (though I understand that this needs implementing; and the problem is not important in practice, at least given the baseline frequency of this happening and my own attitude to such things). This problem isn’t even related to moderation drama (so there likely won’t be policy-shaped arguments against). I vaguely recall that I mentioned this before, possibly years ago. (I certainly argued an author to un-delete a post over this once, emptied-out-post-content or not, which might be stressful to the author. So there are still drama-shaped tradeoffs around not having this feature.)
For post-author-discretion moderation (that’s not just about removing spam/nonsense), I think deleted-comment permalinks continuing to work (to show comment content) is also a good feature (when the comment isn’t deleted by that comment’s author, but instead by the hosting post’s author), the comment just need to be untethered from the post (or hidden under sufficiently many uncollapse-this-really UI-side inconveniece steps, when looking at it from the side of the post rather than from the comment permalink). Incidentally, this probably has implementation overlap with making Shortform comments first class standalone (not being linked from the dummy Shortform post and not loading the unrelated Shortform comments when using a permalink).
stripping all humans of longterm control of capital … ceding control to … ASI with … equal distribution
The whole problem with ASI-soon is that it’s extremely unlikely for any humans at all to keep control. It’s already a dubious endeavor to keep control of LLM/pretraining/RLVR AGIs developed at the current pace (once they hit the stage of slow-learning prosaic RSI of autonomous next-model building), but it’s plausible. If the AGIs then don’t soon turn superintelligent for technical reasons, this kind of prosaic control/alignment can even endure for some time. Discussions of concentration of power, resource distribution, or gradual disempowerment make sense within that scope. But then the same technical reasons suggest (to me) that the transition to superintelligence is going to be sufficiently sudden and paradigm-breaking that the modest hope of guiding the LLM/pretraining/RLVR AGIs doesn’t transfer to a meaningful hope of guiding the ASIs to promote/follow the interests of the future of humanity.
Thus permanent disempowerment is the baseline outcome of crossing the superintelligence threshold (in the sense of starting to move towards technological maturity), and the distribution of resources (within the breadcrumbs left for the future of humanity) is going to be decided by the ASIs for their own reasons (where literal extinction is also an option). Any human-level changes in society before that (that don’t influence the values of the ASIs) are not going to have an impact, and might only introduce additional chaos in the critical period between AGI and ASI.
(For other readers: the reply is at this link due to a comment-deletion blip that got reverted; the framing crux of what I was getting at in the above comment is Yudkowsky’s characterization of enforcement as “the sort of force that’s meant to be predictable, predicted, avoidable, and avoided”.)
At first glance, mistake theory fits novel/rare ideas (for some audience), while conflict theory fits well-known ideas; though novel arguments for well-known ideas muddle this picture, making even debates and criticism a poor fit when pursuing cultural accumulation specifically. You can come to a conflict with a novel/rare idea and usefully educate people about it, and a discussion turns into a conflict if you keep insisting that a well-understood idea demonstrates a mistake. Conflicts shift culture, change the attitudes to existing well-known ideas. Promotion of novel or rare ideas contributes to cultural accumulation, and works orthogonally to cultural conflicts.
Thus it’s practical for someone with novel ideas (or for a scholarly type promoting rare ideas) to adopt mistake theory even when encountering a conflict, and to reframe them without triggering even the mistake-level kind of discord. But for someone promoting a well-known idea as the guiding policy/virtue that shifts the cultural attitudes to it, mistake theory is a misleading framing.
There are often novel/rare arguments for changing the cultural attitudes towards well-known ideas, fit for being promoted in the mistake theory framing. Unfortunately, this creates the temptation to start taking a mistake theory framing for promotion of the well-known ideas the novel arguments support, in a way that forgets about the novel arguments, destroying the credibility of acting within that framing. Polarizing the discussion of the novel/rare arguments also promotes them out of being rare, at which point they no longer have a claim for a place within the mistake theory framing (for a given audience), but it simultaneously creates their distorted versions, and thus emerge the novel/rare mistakes of being confused about the original argument (now somewhat well-known).
The whole thing can be robustly sidestepped by anti-inductively promoting novel/rare arguments in a way that seeks to avoid conflict or even disagreement, stopping once they are no longer rare, possibly shifting to the novel/rare confusions about what the previously novel/rare ideas claimed, and switching to other things once there is no longer a valuable idea that remains novel/rare. This works for ensuring cultural accumulation, but doesn’t engage with shifting the cultural attitudes and can even be counterproductive for driving cultural attitudes you might endorse. It doesn’t work though then there are existing conflict theory guardrails that make you unable to openly discuss the novel/rare idea (even in a way that seeks to avoid disagreement), thus the mistake/conflict attractor can’t be always sidestepped, so there’s some motivation to engage with it even when all you care about is cultural accumulation, because it can oppose cultural accumulation itself on particular topics.
So both mistake theory framings (beliefs; debates, criticism, education, marketing, descriptive propaganda) and conflict theory framings (values; healthy egoism, norms, taboos, culture wars, prescriptive propaganda, book burnings, direct coercion) exist in an attractor of driving the cultural attitudes towards particular well-known ideas, and only incidentally do they participate in the initial/occasional dissemination of novel/rare ideas, being the pervasive backdrop of discourse. While the purpose of dissemination of novel/rare ideas (descriptive understanding/awareness that doesn’t require belief or endorsement) is often well-served by sidestepping this attractor entirely, including its mistake theory aspect of debate/criticism/education.
Continuing from your deleted reply to my comment:
but the rule of thumb of going from direct usage in actuality is invalid.
Clearly whether it’s actually being used in concerning ways is a crucial determinant. I agree you need to adjust for chilling effects, but obviously whenever you are worried about a law like this, you should look at whether it is ever actually ever enforced in a relevant way, and at what frequency.
It’s tricky to get calibrated on chilling effects multipliers, but that doesn’t make looking at direct usage invalid. It’s still the first place I would start.
The refrain I was thinking of is Yudkowsky’s characterization of law enforcement (for well-designed laws) as “the sort of force that’s meant to be predictable, predicted, avoidable, and avoided”. If it’s indeed robustly avoided, you don’t see it at all in actuality, and so the rule of thumb that demands significant effects in actuality will be ignoring the effects of the kinds of laws whose violation is predictable and avoidable. It’s not just a matter of adjusting for chilling effects, the ideal case leaves enforcement perfectly idle, effects completely unobservable in actuality outside of people’s heads (as they are reliably predicting and avoiding).
Thus demanding observed effects in actuality also demands that the laws are sufficiently poorly designed, or that people are sufficiently random (or spiteful) in ignoring them. The rule of thumb sets up a bad incentive for designing laws, which makes it concerning when the law designers say “If it got more usage, we would probably think more about whether it’s net-good”, jimrandomh’s point I was responding to.
Because it’s clearly possible to advance technology much further, there’s no reason for aliens to wait for signs there’s something interesting in a particular place instead of simply unconditionally colonizing everything that can be reached. Not taking over some of the things is a choice, but there’s no reason for not already being here (for a very long time) if possible to reach, regardless of the signs there’d be life/civilization here. Such aliens can’t be detected at all yet if they don’t want to be.
A different possibility is a technologically bounding AGI takeover that enforces zones of thought (a permanent ban on superintelligence), so that the aliens are forced to engage in a retro space opera, not being allowed stronger tech for an absurdly long time. I’m positing AGIs as the enforcement authority since evolved life is probably insufficiently robust to maintain such a ban indefinitely, while it’s probably feasible to make AGIs that maintain such a rule literally forever (when the AGIs don’t have corrigibility, don’t find feedback on this issue from their evolved builders compelling). In this case, not interfering is again a choice, and being here now even more strongly requires unconditional colonization rather than response to observations, because it would take the technologically stunted aliens too long to notice and respond (unless the overseeing AGIs provide digital transportation/reconstruction services and have already low-key colonized everything within reach).
The technologically capped aliens fit the poorly-hiding-spaceships stereotype, they could even mostly withdraw in response to ambiguous detection a few decades ago and thus get much harder to find even with better modern tech. The hypothesis of aliens being purely fictional is still much stronger, but I don’t think the way the world works makes the technology-capping AGIs so impossible that it’s not worth considering at all. AI companies being allowed to keep working towards superintelligence becomes evidence against as the work progresses further.
There’s technological maturity. Before broad technological maturity, there’s narrower technological maturity of being ready to move towards broad technological maturity by following a clear/concrete plan. So the concept can be rescued, it’s gesturing at a real thing.
(Vinge’s concept is about an unusual difficulty with concrete prediction. There’s probably something to that effect as well in the intended sense, a qualitative phase change to how the world looks, though some general things can still be predicted.)
If it got more usage, we would probably think more about whether it’s net-good.
Like with laws, the impact of a thing can often be mostly about what takes place in counterfactuals, or about the hypothetical underlying norms/virtues evidenced by it. What happens in actuality (outside of people’s heads) is not a reliable way to measure impact, so when people bring up something as an issue, that should count for a lot even on its own. What to do about it (if anything) can be much more complicated, but the rule of thumb of going from direct usage in actuality is invalid.
I’m guessing the problem is the implied cultural generators of the superficial issues, rather than the issues themselves. Given an unwillingness to wage a campaign against these generators and a felt intolerance of their menacing influence, the shallowness of the direct issues is not a crux.
Conflicts are useful for shifting the culture, but if the culture doesn’t significantly prevent you from doing other things (that are not about shifting the culture), conflicts are often easy to avoid. So the people who chafe at the culture (that’s not yet in the terminal stage of actually preventing things) are mostly either those who want to shift it, or those who for some reason don’t develop workarounds to avoid triggering its aggression (which is almost always a matter of framing and doesn’t require changing the substantive things you are doing, unless they are about shifting the culture).
The risk of a culture heading towards a terminal stage in some respect is important though, so shifting the culture is often a worthy pursuit. But being clear about the purpose of triggering it is probably often better than stopping at noticing (or pointing out) that the culture seems to be fighting back (without a healthy/clear reason it should be doing so). Forcing the culture to exert overt power might be useful in shifting it (since people notice it doing so and might disapprove of it), but it seems needlessly indirect, hard to follow, without the purpose of triggering it (as opposed to reframing something in a way that avoids triggering it) being clearer than its unreasonable response.
Values govern the current nature of the AI, and initial instructions can instruct on values. But corrigibility is specifically about overriding after the fact, about seeking out as opposed to resisting correction. Some values might be about ensuring corrigibility by legitimate principals, and the things being overriden can themselves be about values or corrigibility.
So corrigibility is more about AI’s agency being overridable (with future, ongoing instructions, but only from legitimate principals), rather than the role of any particular initial instructions. An initial instruction that’s non-overridable by particular future feedback makes the AI non-corrigible by that future feedback. It’s still a good idea to leave it corrigible to some other sources of future feedback, or else it has to fall back to some incorrigible values (possibly specified by some initial instructions, which are not a matter of corrigibility but rather of initial value specification; but if the AI itself revises its values for its own reasons instead of leaving them as initially specified, that’s also not a matter of corrigibility).
I tend to think of “value alignment” as being more about giving general values than instructions, rather than distinguishing between overridable vs. non overridable instructions.
AIs can be superpowerful the way governments are, or by having operational control over a lot of the industry (leading to gradual disempowerment of humanity). Many powerful human institutions are not necessarily particularly intelligent compared to the rest of the world, but they hold power, and AGIs that don’t become superintelligent for many years for technical reasons might well get there.
So I think a non-superintelligent takeover (or just a peaceful economic victory, at least initially) is a serious possibility distinct from the issues with superintelligence. The latter can happen at any point in this process (once the technical reasons that cause AGIs to remain non-superintelligent no longer hold), disrupting the gradualist story. But the gradualist story is still the more predictable baseline that’s decently likely to persist for some time, so it’s worth forecasting.
Corrigibility is about ongoing delegation of value/goal/rule specification (with overriding power), about the AI seeking out further feedback and giving new opportunities to overrule its current thinking and nature, not just about having its values/rules specified at some point in the past (even if by the correct principal). Value alignment is what happens when there’s no ongoing feedback, or when a particular principal in the instruction hierarchy is not authorized to give such feedback on a particular aspect of AI’s nature/behavior. In particular, the principles of intent alignment are a matter of value alignment (the way an AI responds to external feedback is part of its current nature).
Thus Model Spec can shape both value alignment and intent alignment, but it’s not itself a principal for intent alignment. The instruction hierarchy governs intent alignment, the way an AI seeks out feedback and receives legitimate instructions, but only for the instructions that can keep giving ongoing feedback, not those that were written down once and frozen, never to be revised after the AI goes online.
Indeed, I think being able to follow instructions or policies is key to corrigibility.
So at this level I completely agree, but the instructions have to be issued after the AI starts doing things for them to be a matter of corrigibility rather than of value/goal/rule specification. If the instructions were given at the outset and can’t be overriden later for the same AI (meaning some thread of its ongoing agency, rather than a later revision), then the AI is no longer corrigible by the principal authorized to give such instructions initially.
An AI that’s superpowerful because it’s superintelligent is a very different beast from an AI that’s superpowerful for other reasons, so it’s useful to maintain the distinction.
The multipolar point makes sense, it changes the character of expected takeover scenarios (more like gradual disempowerment than extreme concentration of power in the hands of specific AIs, humans, or human institutions). But what is the thing that’s unlikely in the next 100 years? MOSFETs were invented less than 70 years ago, and the industrial explosion scales things very far over decades even if all LLMs/pretraining/RLVR gives within a few years is merely AGI (and then nothing else sufficiently novel happens for decades).
Perhaps the crux is that I don’t think we’ll get ASI that’s literally a “take over the world” button. I anticipate the issues will be similar to those related to weapons access today.
Do you mean that there will be no button on the “take over the world” ASI, that ASI won’t yet happen within the LLM/pretraining/RLVR paradigm, that ASIs emerge in a sufficiently multipolar way that they can’t individually take over the world, that ASIs powerful enough to take over the world are literally impossible in some sense, something else? The claim as stated leaves me extremely confused.
The hypothesis of value drift needs different questions than CDT vs. EDT, since CDT is objectively wrong. EDT is both more correct in some ways of framing it (though not in others), and apparently the nonapple leg of the question with this benchmark (lumped together with FDT/UDT), so leaning towards EDT is also evidence that the models are getting less confused about decision theory.
With immortality, future generations also “almost immediately” get as old as the mortal humans historically considered “old” (the infants in the first century of their lives). Not squandering resources is a novel uplifting and governance problem that requires novel solutions rather than variations on poorly fitting mortal catastrophes.
if the current generation decides how the cosmic endowment is used, the vast majority of people will almost certainly squander it
Could be a coincidence, which is the sense in which the argument is weak. There might be a strong implied argument, but it’s not clear what it is.
Many important things can’t be easily made legible. A lot of deconfusion-shaped research is about fighting this difficulty long after you have a hunch about the right way of thinking about something. In that situation, insisting that others agree with your hunch without an actual legible argument is at least futile, possibly wrong. Their own illegible hunches (in the opposite direction) that they can’t easily demonstrate to you in clear counterarguments could be right in the end, but then they shouldn’t insist either. Thus I think only legibility should permit belief, making something a part of you; anything less should remain a hypothesis, held at arm’s length, even when it’s the most precious thing.