I think that, in the absence of a robust solution to the “human alignment” problem, (public) solutions to corrigibility differentially increase S-risks. Under current conditions, “solving corrigibility” would modestly increase the probability of good futures, and strongly increase the probability of astronomical suffering.
rvnnt
Does GRAM remain useful when models become capable of learning continually and re-deriving any missing knowledge?
You should be worried about corrigibility if: some actor can [...] hold a gun against your head.
Yep. And the government almost surely can and will effectively hold a gun to your head, if they think you’re close to building a superintelligence.
What, in concrete practice, would you predict happens, if some lab succeeded at building a corrigible ASI, or was close to doing so?
I think that, under the current circumstances, the ASI would very likely end up controlled by some power-hungry sociopath. We would hit the third filter.
That being the case, I’m currently leaning towards thinking that (all current approaches to building ASI are terrible but) Anthropic’s approach is less bad on expectation.
(If anyone could lay out a concrete, realistic scenario in which a corrigible ASI gets built on Earth any time in the next 15 years, that does not end with a sociopath in power, I’d be curious to see it! (May or may not change my mind w.r.t. “corrigibility vs constitution” question, depending on how likely that class of scenario seems.))
That comment would make sense in a world where everyone is more or less neurotypical and psychopaths [1] do not exist. But that’s not planet Earth. The listed things do not even begin to apply to most psychopaths. Most of the worst harm is committed by psychopaths. And positions of high power tend to be strongly disproportionately populated by psychopaths.
EDIT: I believe the above applies in reasonably civilized societies during peacetime. During wartime, or if non-psychopaths are given immense power, I’d expect most of the worst harm to be done by non-psychopaths (because they are more numerous), for entirely mundane reasons of sadism/rapaciousness/spite/etc.
- ↩︎
I’m using the term “psychopath” loosely here; hopefully my point still comes across.
- ↩︎
IIUC, you (habryka) are saying that in order for the dictator to fill the universe with sentient slaves, the dictator would have to care about “the full depth and complexity” of others’ suffering/submission.
Yet you also say that the same dictator would probably fill the universe with nice things like happy sentient people, without making a strong argument for why we should expect the dictator to care about the full depth & complexity of others’ happiness.
It seems to me you’re applying the “would have to care about the full complexity of X” argument only to bad stuff, but not to good stuff.
In what way do you think people’s desire for others to submit/suffer/worship differs from people’s desire for others to be happy/free/etc, that it makes sense to apply that argument so asymmetrically?
One worry I have with ‘the [AI] whisperers’ and others who investigate these matters is that they may think the model they see is in important senses the true one far more than it is, as opposed to being one aspect or mask out of many.
Imagine telling someone from (e.g.) 2016 that the above sentence is a reasonable thing to say in 2026. (The frog, it boils.)
I think it might be more accurate to say you’re an efficient component in Moloch’s machine.
But if you care about how things like gradual disempowerment play out, then I think the “baddie / goodie / pawn-of-Moloch” framing is probably not very useful. It might be worth instead thinking more concretely, about things like
How much are my actions contributing to speeding up human disempowerment? [1]
How could I keep my job (or whatever) while contributing as little as possible to various bad things?
Who are the relevant actors I would need to coordinate with, in order to slow things down? What, concretely, is stopping me from coordinating with them, and how could I fix that?
What other important considerations are there, besides “speeding up adoption of AI / replacement of humans”?
What could I do to offset harms I cause?
- ↩︎
accounting for the other actors in the Molochian race
How would you rate Principia’s operational adequacy? In particular: how would you rate Principia on closure and opsec? [1]
- ↩︎
Alternatively, if you use a different framework for thinking about operational adequacy, I’d be very curious to read your assessment w.r.t. that framework!
- ↩︎
I think something like what you’re sketching here, viz “harnessing technology to make people and civilization saner”, is probably highly valuable and possibly quite neglected. [1] Thank you for working on this.
A class of infrastructure/technologies that seem very important, but which I didn’t see mentioned in this post: infrastructure for creating common knowledge of better equilibria and coordinating transitions to them. [2] Do you (have plans to) address anything matching that (vague) description anywhere?
- ↩︎
Low-hanging dignity points!
- ↩︎
I.e., something that would solve problems of form “there exists a much better equilibrium, but getting there would require lots of people to have common knowledge of that better equilibrium, and also coordinate and sufficiently-credibly commit to near-simultaneously taking action that would be detrimental to them if they took it alone”. Some examples: move from frequentist stats to Bayesian stats; make it easier for AI labs to (conditionally) stop racing; US voters coordinate to vote for a less sociopathic party-independent candidate (or to replace first-past-the-post with a saner voting system entirely); abolish all JavaScript forever, refactor the Internet to use a non-insane language; kill Elsevier; almost everyone simultaneously leaves (at least the more toxic platforms of) social media (and move to a less toxic new platform); journals/researchers commit to preregistering studies and publishing negative results; etc.
- ↩︎
There’s some chance that the patchwork AI safety strategy of the leading companies might just work well enough
I wonder if it’d be a good idea to mention that “a small number of humans have intent-aligned vastly superhuman AIs” does not yet imply “things will be nice (or tolerable) for most people”?
OTOH, then you maybe also need to point out that “superhuman AIs are open-weights and anyone can run them” also leads to death-or-worse (to pre-empt “solutions” like “so let’s give everyone AI!”)?
The four probabilities given as premises are inconsistent. The first three determine the fourth. (Also, there’s an arithmetic error in the p(G|M) calculation, as pointed out by Bucky.)
Given
p(M) = 0.9
p(G) = 0.05
p(M|G) = 0.99
it must be that
p(M|-G) = (p(M) - p(M,G)) / (1 - p(M,G) - p(-M,G)) = 0.8505 / 0.95
which is approximately 0.895263. Not 0.02.
If this feels confusing, I suggest drawing a Venn diagram or something. If you have a box of area 1.0, containing a blob M of area 0.9 and another blob G of area 0.05, such that G is almost entirely inside M, then...
Interesting. Thanks. How did you arrive at the above picture? Any sources of information you’d recommend in particular?
After reading about Trump’s actions w.r.t. Greenland, I’m updating further away from
Trump’s policies/actions are at least partially aimed at pursuing the national interests of the US (albeit possibly in a very misguided/incompetent way)
and further in favor of both
Trump has gone insane,
and/or Trump is intentionally acting on behalf of Putin.
I’d like to find more/better sources of evidence about “what is the US executive branch optimizing for?”; curious to hear suggestions.
(Also, to Americans: How high/low salience is the issue in the US? Also: curious to read your analysis of your chief executive’s behavior.)
Dominic Cummings (former Chief Adviser to the UK PM) has written some things about nuclear strategy and how it’s implemented in practice. IIUC, he’s critical of (i.a.) how Schelling et al.’s game-theoretic models are (often naively/blindly) applied to the real world.
I updated a bit towards thinking that incompetence-at-reasoning is a more common/influential factor than I previously thought. Thanks.
However: Where do you think that moral realism comes from? Why is it a “thorny” issue?
social-media-like interfaces for uncovering group wisdom and will at larger scales while eliciting more productive discourse
That seems like it might significantly help with raising the sanity waterline, and thus help with coordinating on AI x-risk, and thus be extremely high-EV (if it’s successfully implemented, widely adopted, and humanity survives for a decade or two beyond widespread adoption). [1]
Do you think it would be practically possible with current LLMs to implement a version of social media that promoted/suggested content based on criteria like
validity of reasoning,
concrete factual claims (whether correct or incorrect), vs vibing,
thoughtful long-form writing, vs memeing,
good-faith debate, vs e.g. adversarial tribalistic signaling?
- ↩︎
The “widely adopted” part seems difficult to achieve, though. The hypermajority of genpop humans would probably just keep scrolling TikTok and consuming outrage porn on X, even if Civilization 2.0 Wholesome Social Media were available.
This discourse structure associates related claims and evidence, [...]
To make it practically possible for non-experts to efficiently make sense of large, spread-out collections of data (e.g. to answer some question about the discourse on some given topic), it’s probably necessary to not only rapidly summarize all that data, but also translate it into some easily-human-comprehensible form.
I wonder if it’s practically possible to have LMs read a bunch of data (from papers to Twitter “discourse”) on a given topic, and rapidly/on-demand produce various kinds of concise, visual, possibly interactive summaries of that topic? E.g. something like this, or a probabilistic graphical model, or some kind of data visualization (depending on what aspect of what kind of topic is in question)?
Ideally perhaps, raw observations are reliably recorded, [...]
Do you have ideas for how to deal with counterfeit observations or (meta)data (e.g. deepfaked videos)?
I think that’s probably true in some polities, at least under “normal” conditions that do not include “a small group of humans can gain a decisive strategic advantage over all of Earth”.
However, it does not appear to be true in most of the relevant places in the United States [1] . (See e.g. the failed firing of Sam Altman, or various actions of the Trump 2.0 admin.) And even if “democratic oversight” were working as intended, it operates via elections and other feedback mechanisms that were not designed (are too slow) to respond to threats like “the executives can suddenly order the entire automated military to depose their enemies and take control of all critical infra” [2] .
I.e., even if solutions to human alignment might exist in theory (and/or Iceland), they appear to currently be absent from the real world. And yet, in order for ASI-being-corrigible to end well, such solutions need to actually be implemented in reality in ways that will actually work under pressure.
I do not think humanity is remotely on track to invent and implement such solutions any time in the foreseeable future. Consequently I think (publicly) working on corrigibility is strongly negative-EV.
A thing that could change my mind: Could you describe, in concrete detail, a “strategy for solving the human alignment problem”, which would actually work under realistic conditions [3] , and which seems likely to actually be implemented (in the relevant places) before ASI is built?
Let alone e.g. China or Russia.
Not saying that specific class of scenario is super likely; intended only as an illustrative example.
Humans following their myopic incentives, humans sucking at coordinating, humans being incompetent, humans being selfish sociopaths, etc.