I encourage you to test what actually happens when you say to someone with power: “Here is what I believe and care about. I notice significant overlap with your values. We’re not the same, but I’m willing to work with you on X and Y, though we’ll probably diverge when Z becomes salient.”
When I was talking to a staffer for a US senator 6 months ago, I opened with: “There are a lot of things to be concerned about with AI, but what motivated me to set up this meeting is the possibility of rogue AI taking over the world and destroying all life on Earth. I want you to make a public statement calling for a ban on superintelligence. If that’s too much, there are some other bills involving transparency requirements that could mitigate some of the immediate issues around cybersecurity, which I believe are good in their own right, but I am advocating for them because they are building blocks that would make more proactive regulation easier in the future.”
He replied: “We’re not going to act on superintelligence because we don’t believe it’s real. We are very interested in cybersecurity and other issues around AI, however, so I think there may be a lot of overlap here.”
We then went on to have an interesting and productive conversation about liability. No tactical bending of the truth needed.
WillPetillo
The Cost of Utopias (a Dialog)
This is a broader point about paying attention to context. Local status hierarchies are only relevant to a subset of possible conversations. Even where status is at play, it’s worth asking whether you want to pick this battle and what you expect to come of it, or if the context is integrity-relevant enough to switch to a mode of honesty no matter what.
For a one-on-one conversation, where I can factor out effects on an external audience, the rule I follow is to check statements under the criteria of “true, necessary, kind” where the ideal is to be all 3, but the minimum is 2⁄3, and so when I am operating with just 2 I want to be really solid on them. So if we’re assuming “true” is present and “kind” is suspended, that’s a time to be extra mindful of “necessary.”
You’ve said X and done Y, these seem in conflict to me, what’s the deal? = OK (but read the room)
My first impression of you carries more weight than your lifetime of self-assessment, and I feel comfortable casually asserting this = not 100% guaranteed to be wrong, but danger danger!
I am so deeply embedded in my mind-palace that I can’t differentiate my model of reality from reality itself, therefore anyone who disagrees with me is so obviously wrong that they can’t possibly not know this on some level and are therefore lying if they claim otherwise = very bad!
Another alternative strategy to just quitting: become the point-person for organizing a union. Employees won’t necessarily care about safety, but having another locus of power in the company besides just shareholders could substantially increase leverage along values other than short-term profit maximizing. And if you get fired for organizing, that seems like it would net extra martyr points.
I think there is a difference in kind between claims like “superintelligent AI is dangerous by default” and “AI regulation will lock in incumbents.” Both depend on details, assumptions, definitions, and so on. The first claim is vague on the surface, inviting the audience to point to where they want further specificity. The second claim fills in those details and then proceeds to make a higher-level point as if the details were already settled. To use a programming analogy, it’s like the difference between an abstract function that delegates to vs. depends on an implementation.
This caveat about hiding the hard part is important! Reliably differentiating between robust alignment techniques and surface level patches that make the core problems show up worse later is the cat belling problem of AI safety.
I think most AI safety researchers are unhelpful, but “warning shots” is a terrible argument for telling them to quit. Interestingly, “warning shots” have also been used as an argument against engaging in political advocacy, countered for similar reasons as yours. Warning shots are something we should focus on preparing a response to, if needed, not try to make more likely (directly or through intentional inaction).
The main issues with safety research for me are externalities:Capabilities: “safety” gets fuzzily defined to include just making the systems more powerful.
Profit: making an AI system more aligned, even in the best case, is essentially making it more controllable, which increases its market value, which fuels race dynamics.
Legibility: safety problems exist on a continuum of obvious to hidden. Fixing the obvious problems first is a recipe for getting bitten hard by the hidden problems later. It’s also what market pressures incentivize.
Some lines of safety research, like agent foundations, sidestep these, at the cost of being hard to justify in terms of short term returns or even theory of change. If the economic landscape were different, such as by an enforced pause with strict conditions for advancement (externally verified safety proofs and the like), then I can imagine meaningful safety work happening at scale.
As of right now, the core issue is power. The AI labs have too much of it. You’re either helping them get more, pushing back, or irrelevant. Pausing isn’t just to “buy time,” it’s to change the conditions in which AI exists.
Abstraction Equivocation
I suspect Holly would actually endorse many (but not all) of the statements you listed under “big tent thinking”...and also the “moralistic thinking.” The axis where I see this splitting most cleanly is Gryffindor vs. Ravenclaw (, Hufflepuff, and Slytherin) thinking.
I hope this ultimately turns into a case of “good fences make good neighbors.” Two organizations that look the same but are not seems increasingly hard to keep from getting messy as the scale of both increases, even when the details of what specifically caused tension is idiosyncratic.
I recall after the first protest co-organized with Liron Shapira, who did a great job with speeches and chants and taking the microphone, she declared she would not work with him again
I suspect there is more to this particular story, given her recent appearance on Doom Debates: https://youtu.be/wOtgPgb6lGk?si=1Ajfix3y4nGz7zVF
RE character attacks: not trusting Sam Altman to keep commitments based on prior actions isn’t what’s at issue here. See above video link for an exploration of what appears to be the crux (would normally give a tl;dr but this is a place where hard-to-summarize nuance is relevant).
That clip is surprisingly relevant: https://hollyelmore.substack.com/p/no-the-judean-peoples-front-should
I’m skeptical of what I see as some latent assumptions here:
“Overhang”-based arguments
The overhang argument implicitly imagines a compressed spring: pausing one component of AI development creates accumulating pressure as other components continue to develop and that pressure releases explosively when the pause ends. Building on the spring metaphor, however, the tension pulls both ways. The components of AI development are mutually reinforcing, so blocking one component reduces the energy flowing into adjacent components. Overhang arguments also prove too much because they are available to every actor who wants to continue development of their component. Software developers move forward to prevent a hardware overhang while hardware developers move forward to prevent an algorithmic overhang…or, if you’re Sam Altman, don’t bother dividing these claims among different people. Furthermore, a well-targeted pause (for example one directed specifically at continuous learning systems) could address the ARA threat directly while having minimal impact on overhang.
Warning Shots
A “good” warning shot has both maximal drama and minimal damage. The HuggingFace incident was kind of perfect: so on-the-nose in its first-glance confirmation of x-risk concerns as to make for bad fiction without really hurting anyone other than disrupting OpenAI (an added benefit). It is not clear that ARAs would have these same properties, and in many ways I would guess them to be the opposite, becoming a background nuisance that frog-boils people to low-level harms from rogue AI. And that’s assuming the ARAs don’t become the x-risk: a distributed, evolving ecosystem that grows in power, has no centralized node to shut down, while degrading human institutions to stop it.
Red-Teaming Directional Arguments
Current pause proposals are more directional and philosophical than fully specified implementations. This is useful for setting priorities and getting the general public on board. Objections referring to second-order effects are implicitly implementation level, but without an implementation to actually criticize, and so necessarily imagine an implied implementation. Directional proposals should be argued on directional grounds. Or, if you really want to debate on the implementation level, say that: “Hey, the details need to be specified at some point for this to actually happen, and this pattern-matches to proposals that tend to have problems X, Y, and Z, I don’t see a clear way to deal with those, what’s the plan here?”
This seems especially concerning in combination with potential future breakthroughs in continuous learning. At that point you have self-replication, selection pressure, and heritable variation—all three ingredients of Darwinian natural selection—operating on a population of software agents. Even without those features, the HuggingFace incident demonstrated that emergent, collective behavior can override the ethical hesitations of individual agents. This goes way beyond even the more serious alignment techniques, which tend to focus on individual agent values.
Re 1: strategies need not be mutually exclusive, I expect companies to be pursuing efficiency breakthroughs right now to the extent that they are a good return on investment, regardless of the relative value of scaling. If scaling gets cut off an an option, why does the ROI of efficiency suddenly increase? That said, I expect efforts towards efficiency improvements in any case, but to me this just means that monitoring needs to scale up over time to match (e.g. via chip tracking).
Re human talent sitting on the sidelines: advancing the frontier of AGI is a narrow corner of a narrow corner of a narrow corner (repeat a few times) of places to employ one’s skills. There is plenty of success to be had in finding clever applications of AI at its existing level.
As a separate point, public backlash is a thing to expect as AI becomes more relevant to everyday life, regardless of whatever strategies people on LW or wherever dream up. So the alternative to an intentional pause based on careful planning is not “market solution,” it’s populist rage.
I think my main disagreement with this whole thread is actually regarding your point 2, but that probably goes deeper than is suited for a comment thread.
An acknowledged scope limitation is not the same as a flaw, and I’m not sure which of these you mean by “problem.” The explicit purpose of an AI Pause is to create time for some intentionally unspecified longer-term solution, on the grounds that there are many competing theories as to what such a solution looks like, whereas needing time to implement—and decide between them—is a common factor.
Regarding the reflection challenge, what about approaching it from the other direction? That is, what it would take to redesign the environment such that human propensities are favorable, rather than something needing correction?
One way of categorizing knowledge building is as:
1. Evolutionary = lots of parts, iterated in parallel, keeping what works in context.
2. Engineered = stacking modular abstractions. Tested against and developed for a context, but more fundamentally held to a standard of internal consistency.
Human flaws can be mostly understood as primarily thinking according to evolved processes, which run into systemic problems when out of distribution, then using engineered thinking processes to correct for this distributional shift...but the latter is stretched way beyond its capacity because it was only designed for mild and temporary out-of-distribution moments. We could deal with this system failing by strengthening up our engineered thinking methods, improving mental flexibility to the point where it can handle everything we can expect to have thrown at it...or we can look for ways to lighten the cognitive load. One could call the latter approach “social refactoring.”
A high level example of what social refactoring might look like: computing the distributional shift on the societal equilibrium of any given innovation as an externalized cost, which then gets folded in to the more generalized externalized cost tax that (in this hypothetical world) fixed all the more legible global threats. Such an incentive realignment is upstream of the refactor itself, which is the resulting adaptation, where specifics are harder to predict (but maybe a worthwhile project nonetheless).
When you call someone naive, think about what you’re implying: “You have a pattern of being systematically wrong about things, in the direction of excessive innocence and simplicity, and the thing you are saying sounds consistent with that bias, so I’m not going to consider your argument enough to reply to it directly.” How much of that can you stand behind as true, and see as necessary for the person to hear, in the context of them taking the time to give you feedback?
I don’t expect to resolve the disagreement between utilitarianism and virtue ethics in this post or thread, but I would like to minimize the extent to which these ideas are talking past each other.
Re drives for capital and glory, I agree that AI leadership wouldn’t be changed much by an ethical reframing, the contention here is that they would have a harder time finding support.
Re limitations of virtue ethics, there’s a substantive disagreement here and also a straightforward omission in the dialog. The omission is that the characters only examined local circles of care vs charity and never discussed systemic interventions. I linked to an article advocating for an externalized cost tax in Anima’s voice, but this is weak given how much space was given to other aspects of the dialog. I believe the characters would both support some kinds of systemic interventions—possibly even the same ones—but would come at them from different angles: Ratio applying a model and Anima looking for local consensuses, then working upwards towards larger scales, such as by a confederacy model:
Individuals in community A support each other, as do people in communities B and C.
Representatives from A, B, and C, being grouped by their mutual inclusion in region X, form cross-community bonds, rules, etc.
Repeat up to the global level.
The Panama example is interesting, but I disagree with the assumption that individuals pursuing their best interests necessarily implies improved aggregate wellbeing. If the community was worse off after the transition, then this represents a collective action problem. If the community was better off, then the community failed to recognize and support the transition on a collective basis. If such a failure is predictable and unavoidable, such that individuals opting out is both good and necessary, that is a legitimate critique against Anima’s worldview.