During my outreach, onboarding, and lobbying for PauseAI, there’s a pattern of mistaken reasoning I see repeatedly, which I expect to become higher stakes with Senators Sanders and Casar’s recent proposal to ban superintelligence. I’m going to call it Abstraction Equivocation, which is applying implementation or trust level critiques to directional proposals.
Here’s what I mean by these terms:
Directional: about priorities, goals, and central assumptions. As examples of priorities: preventing x-risk from AI is non-negotiable and required for anything else to matter; competitiveness is a constraint other values need to work within to not undercut themselves; equity, innovation, and stability trade off with each other in context-dependent ways, with room for legitimate disagreement as to the proper balance. As examples of central assumptions: superintelligent AI is a meaningful concept, likely to occur in a relevant timeframe, and extremely dangerous if not handled well. Directional critiques should engage with whether these are the right priorities and if their foundational assumptions make sense.
Implementation: about specific mechanisms. Does a given definition of superintelligence capture what is intended, exclude what is not, and is there a clear way to tell? Will a given monitoring system catch what it needs to catch? Such questions require a clear point of reference.
Trust: about who decides the details. Can a given agency be trusted to make good decisions without getting captured? Can the courts be trusted to interpret key definitions correctly? Trust sits at the bridge between direction and implementation, applying to proposals that are directional, but with a plan to fill in the details, assessing the people who are responsible for that filling in.
These levels are relative and recursive. International coordination, for example, is directional relative to the monitoring and enforcement regimes applied, and also an implementation of preventing x-risk. High level proposals for monitoring, in turn, are implementation details of coordination, and also directional in the sense that they need to be implemented on a technical level—at the direction of specific agencies, which require trust.
Here are some common examples of what I consider Abstraction Equivocation:
The US can’t slow down AI development or else China will race ahead
Assumes that a hypothetical agreement will not contain adequate monitoring and enforcement provisions to prevent exactly this from happening. You have to actually look at the relevant provisions to know if they are inadequate.
Regulating AI will lock-in a monopoly for industry incumbents
Assumes that hypothetical regulation will introduce extensive bureaucratic red tape that applies equally to everyone, creating a high-cost barrier to entry. In reality, actual proposals where this objection tends to be raised (1) focus on frontier development, which is vastly more gated by compute and infrastructure costs, such that regulation is a drop in the bucket, and (2) have cost thresholds below which the regulation does not apply—and sometimes even specific carveouts for open source.
Or to put it in simpler terms: RTFB or GTFO.
In contrast, here are some example objections that are based on the same worldviews as the above, but don’t involve Abstraction Equivocation:
Directional:Concerns about “x-risk” are overblown. AI alignment, to the extent it was ever a meaningful concept, is basically solved. We should prioritize innovation so that we get the benefits of AI as soon as possible.
Implementation:“Superintelligence” is poorly defined in this bill and could be reasonably interpreted to include all sorts of normal technologies.
Trust:Once you create any sort of government agency to regulate AI, it’s going to be subject to perverse incentives to drift from its original purpose and centralize power in the national government or create regulatory moats for industry incumbents.
These examples vary widely in how wrong I believe they are, but each at least clearly points toward what kind of response it invites.
Spotting Abstraction Equivocation in the wild can be tricky, because directional-sounding statements can imply expectations about implementation. Consider this recent X post from Dwarkesh Patel (which I’m picking on because it’s subtle not because it’s egregious):
Notice that it never actually quotes, paraphrases, or summarizes any provisions in any particular piece of legislation (such as the Sanders-Casar bill) that are inadequate, underspecified, or likely to lead to unintended consequences. For example, the argument that: “During a pause of AI development, compute would continue to build up” assumes that the bill doesn’t contain adequate compute tracking and governance. Maybe it doesn’t. Maybe its attempts are underspecified. Maybe the details have been delegated to an agency that lacks sufficient authority, competence, or trustworthiness. All of these are different objections, requiring different responses, and I don’t know which one Dwarkesh is making.
This example is especially pointed given the clear context of a bill to cite. But even if Dwarkesh was narrowly responding to Sanders’ directional call for a ban on superintelligence, the same issue of Abstraction Equivocation applies. The implied structure is: “Any proposal of type X must contain features A, B, and C, feature B has consequences D, E, and F, consequence F is really bad, therefore X is bad.” But all that gets said out loud is: “X is bad because F.”
Directional, implementation, and trust objections all have their place. Directional disagreements cut to our core values, engaging on a philosophical level about how to manage tradeoffs between different things that matter, and set the stage for what categories of policy even make sense to consider. Their universality makes directional proposals accessible to everyone, which is why PauseAI’s proposals have often been of this type. But at some point, philosophy needs to cache out into policy, and that requires details to get right. This is where implementation objections come in, red-teaming the details of how grand ideas get put into practice, anticipating and avoiding unintended consequences. When dealing with complex, evolving issues, however, such details can’t be planned out in advance, but must rather be delegated to entities with a defined scope of authority. At that point, it becomes essential to ask whether those entities are trustworthy.
What doesn’t help is conflating these types. A bad goal can’t be fixed by achieving it better; good intentions can’t fix poor implementation. Getting the category right determines whether a conversation can make progress.
In improv, ‘yes, and’ accepts the premise of a scene and builds on it. “Yes, but” accepts it and redirects, or adds a conflict to overcome. A “no” blocks the scene from happening, forcing others to find a way to start over. In the real world, our goal is to move towards good futures, not dramatic ones, so sufficiently directional disagreement is a legitimate reason to say “no.” But such obstruction should label itself clearly, not disguise itself as engagement.
Abstraction Equivocation
During my outreach, onboarding, and lobbying for PauseAI, there’s a pattern of mistaken reasoning I see repeatedly, which I expect to become higher stakes with Senators Sanders and Casar’s recent proposal to ban superintelligence. I’m going to call it Abstraction Equivocation, which is applying implementation or trust level critiques to directional proposals.
Here’s what I mean by these terms:
Directional: about priorities, goals, and central assumptions. As examples of priorities: preventing x-risk from AI is non-negotiable and required for anything else to matter; competitiveness is a constraint other values need to work within to not undercut themselves; equity, innovation, and stability trade off with each other in context-dependent ways, with room for legitimate disagreement as to the proper balance. As examples of central assumptions: superintelligent AI is a meaningful concept, likely to occur in a relevant timeframe, and extremely dangerous if not handled well. Directional critiques should engage with whether these are the right priorities and if their foundational assumptions make sense.
Implementation: about specific mechanisms. Does a given definition of superintelligence capture what is intended, exclude what is not, and is there a clear way to tell? Will a given monitoring system catch what it needs to catch? Such questions require a clear point of reference.
Trust: about who decides the details. Can a given agency be trusted to make good decisions without getting captured? Can the courts be trusted to interpret key definitions correctly? Trust sits at the bridge between direction and implementation, applying to proposals that are directional, but with a plan to fill in the details, assessing the people who are responsible for that filling in.
These levels are relative and recursive. International coordination, for example, is directional relative to the monitoring and enforcement regimes applied, and also an implementation of preventing x-risk. High level proposals for monitoring, in turn, are implementation details of coordination, and also directional in the sense that they need to be implemented on a technical level—at the direction of specific agencies, which require trust.
Here are some common examples of what I consider Abstraction Equivocation:
The US can’t slow down AI development or else China will race ahead
Assumes that a hypothetical agreement will not contain adequate monitoring and enforcement provisions to prevent exactly this from happening. You have to actually look at the relevant provisions to know if they are inadequate.
Regulating AI will lock-in a monopoly for industry incumbents
Assumes that hypothetical regulation will introduce extensive bureaucratic red tape that applies equally to everyone, creating a high-cost barrier to entry. In reality, actual proposals where this objection tends to be raised (1) focus on frontier development, which is vastly more gated by compute and infrastructure costs, such that regulation is a drop in the bucket, and (2) have cost thresholds below which the regulation does not apply—and sometimes even specific carveouts for open source.
Or to put it in simpler terms: RTFB or GTFO.
In contrast, here are some example objections that are based on the same worldviews as the above, but don’t involve Abstraction Equivocation:
Directional: Concerns about “x-risk” are overblown. AI alignment, to the extent it was ever a meaningful concept, is basically solved. We should prioritize innovation so that we get the benefits of AI as soon as possible.
Implementation: “Superintelligence” is poorly defined in this bill and could be reasonably interpreted to include all sorts of normal technologies.
Trust: Once you create any sort of government agency to regulate AI, it’s going to be subject to perverse incentives to drift from its original purpose and centralize power in the national government or create regulatory moats for industry incumbents.
These examples vary widely in how wrong I believe they are, but each at least clearly points toward what kind of response it invites.
Spotting Abstraction Equivocation in the wild can be tricky, because directional-sounding statements can imply expectations about implementation. Consider this recent X post from Dwarkesh Patel (which I’m picking on because it’s subtle not because it’s egregious):
Notice that it never actually quotes, paraphrases, or summarizes any provisions in any particular piece of legislation (such as the Sanders-Casar bill) that are inadequate, underspecified, or likely to lead to unintended consequences. For example, the argument that: “During a pause of AI development, compute would continue to build up” assumes that the bill doesn’t contain adequate compute tracking and governance. Maybe it doesn’t. Maybe its attempts are underspecified. Maybe the details have been delegated to an agency that lacks sufficient authority, competence, or trustworthiness. All of these are different objections, requiring different responses, and I don’t know which one Dwarkesh is making.
This example is especially pointed given the clear context of a bill to cite. But even if Dwarkesh was narrowly responding to Sanders’ directional call for a ban on superintelligence, the same issue of Abstraction Equivocation applies. The implied structure is: “Any proposal of type X must contain features A, B, and C, feature B has consequences D, E, and F, consequence F is really bad, therefore X is bad.” But all that gets said out loud is: “X is bad because F.”
Directional, implementation, and trust objections all have their place. Directional disagreements cut to our core values, engaging on a philosophical level about how to manage tradeoffs between different things that matter, and set the stage for what categories of policy even make sense to consider. Their universality makes directional proposals accessible to everyone, which is why PauseAI’s proposals have often been of this type. But at some point, philosophy needs to cache out into policy, and that requires details to get right. This is where implementation objections come in, red-teaming the details of how grand ideas get put into practice, anticipating and avoiding unintended consequences. When dealing with complex, evolving issues, however, such details can’t be planned out in advance, but must rather be delegated to entities with a defined scope of authority. At that point, it becomes essential to ask whether those entities are trustworthy.
What doesn’t help is conflating these types. A bad goal can’t be fixed by achieving it better; good intentions can’t fix poor implementation. Getting the category right determines whether a conversation can make progress.
In improv, ‘yes, and’ accepts the premise of a scene and builds on it. “Yes, but” accepts it and redirects, or adds a conflict to overcome. A “no” blocks the scene from happening, forcing others to find a way to start over. In the real world, our goal is to move towards good futures, not dramatic ones, so sufficiently directional disagreement is a legitimate reason to say “no.” But such obstruction should label itself clearly, not disguise itself as engagement.