Please spend <5 minutes filling in the below polls on AI alignment!
Thank you to everyone who filled out last month’s polls. It was great to see 60+ comments engaging with these issues.
This month’s survey has already been taken by a panel of 15 alignment researchers, including Scott Alexander (ACX), David Manheim (ALTER) and Jeff Sebo (NYU). We’ll compare panel and community responses in an upcoming report, which we’ll publish here and on EA Forum. To get notified when it’s released, you can subscribe to our new Substack.
Many people we’ve talked to have very different intuitions about where the alignment community stands on the below issues. We hope that your responses to these polls, and the resulting report, will help map core areas of (dis)agreement within the field, and ground CaML’s research agenda.
A few final things about the polls themselves:
Please vote on EA Forum, or in the comments by agreeing or disagreeing with each statement.
Timeframe: unless a statement says otherwise (e.g. post-AGI), read forward-looking claims as being about roughly the next 2 years.
We’re not trying to find the ‘right’ answers. Please answer based on your own best guess.
% agree is your % credence in a given position
Please let us know if you think the questions are ambiguous or embed false assumptions
Any further engagement with the content of the polls in the comments is encouraged
Thanks to BlueDot Impact for funding this work.
The polls
If animals continue to exist in a post-AGI world, animal suffering will not persist
Benchmarks will become useless due to eval awareness[1]
Current AIs are capable of suffering
The backfire risks of AI values alignment outweigh the expected positives
Theories of consciousness will lead to actionable understanding of AI consciousness[2]
At the margin, S-risk work in AI is more important than x-risk work[3]
Increased care by AIs for some kinds of entities will have substantial spillovers to other kinds (e.g. humans to digital minds and vice versa)
We are making good progress in the AI S-risk space and research is on track
Understanding how LLMs learn (e.g. SLT) is underrated relative to work on instilling specific behaviors
Most current evidence of misalignment is actually models role-playing a misaligned AI[4]
Non-human suffering is neglected by the AI safety community[5]
Model wellbeing and model alignment are in conflict[6]
Research on AI suffering has higher marginal value than research on AI consciousness[7]
Deceptive AIs will be able to hide unwanted behaviours from mechanistic interpretability tools (e.g. by encoding them redundantly across pathways, or shifting them into representations the tools do not capture)
Insofar as alignment continues to promote overlapping sets of characteristics (e.g. helpful, harmless, honest, corrigible), should compassion be one of those characteristics?
- ^
This primarily refers to safety and alignment benchmarks rather than capability benchmarks like coding. “Useless” means their results should no longer be treated as evidence about how models behave outside evaluation.
- ^
“Actionable” means good enough to build consensus around policy decisions in practice. It does not require a given theory to be proven correct or widely accepted.
- ^
This is about where the next dollar is best spent, not about which area you think is more important overall.
- ^
“Role-playing” means the behaviour arising from the model enacting a persona cued by the setup, or from misunderstanding the task, rather than from stable goals that would persist across contexts.
- ^
This includes both animal and digital suffering. If you think one is neglected but not the other, count this as agreeing, but feel free to share specifics in the comments.
- ^
You agree to the extent that you anticipate in-practice trade-offs between work on these two cause areas over the next two years.
- ^
This is a question about where the next dollar is best spent between the two fields (even if you might argue that the second is a prerequisite for the first).
Insofar as alignment continues to promote overlapping sets of characteristics (e.g. helpful, harmless, honest, corrigible), should compassion be one of those characteristics?
Deceptive AIs will be able to hide unwanted behaviours from mechanistic interpretability tools (e.g. by encoding them redundantly across pathways, or shifting them into representations the tools do not capture
The backfire risks of AI values alignment outweigh the expected positives
Benchmarks will become useless due to eval awareness
If animals continue to exist in a post-AGI world, animal suffering will not persist
Research on AI suffering has higher marginal value than research on AI consciousness
Model wellbeing and model alignment are in conflict
Non-human suffering is neglected by the AI safety community
Most current evidence of misalignment is actually models role-playing a misaligned AI
Understanding how LLMs learn (e.g. SLT) is underrated relative to work on instilling specific behaviors
We are making good progress in the AI S-risk space and research is on track
Increased care by AIs for some kinds of entities will have substantial spillovers to other kinds (e.g. humans to digital minds and vice versa)
At the margin, S-risk work in AI is more important than x-risk work
Theories of consciousness will lead to actionable understanding of AI consciousness
Current AIs are capable of suffering