Deceptive AIs will be able to hide unwanted behaviours from mechanistic interpretability tools (e.g. by encoding them redundantly across pathways, or shifting them into representations the tools do not capture
jonahmattwoodward
Research on AI suffering has higher marginal value than research on AI consciousness
Model wellbeing and model alignment are in conflict
Non-human suffering is neglected by the AI safety community
Most current evidence of misalignment is actually models role-playing a misaligned AI
Understanding how LLMs learn (e.g. SLT) is underrated relative to work on instilling specific behaviors
We are making good progress in the AI S-risk space and research is on track
Increased care by AIs for some kinds of entities will have substantial spillovers to other kinds (e.g. humans to digital minds and vice versa)
At the margin, S-risk work in AI is more important than x-risk work
Theories of consciousness will lead to actionable understanding of AI consciousness
The backfire risks of AI values alignment outweigh the expected positives
Current AIs are capable of suffering
Benchmarks will become useless due to eval awareness
If animals continue to exist in a post-AGI world, animal suffering will not persist
Community Polls on Alignment Controversies II
Would your AI travel agent book a bullfight? Testing whether agents consider animal welfare without being prompted
The philosophy of Pro-Animal Wikipedians is to ensure that all their edits are factual and reliable, in line with Wikipedia’s guidelines. I’d actually argue that these edits are important in providing a more balanced article that takes into account a wider variety of considerations (like animal welfare).
A good example (I can’t confirm whether this is an instance of PAW edits) is the Assured Food Standards page, which documents consistent instances of animal abuse on Red Tractor Assured farms. This is arguably more representative than a page without these abuses documented, and I think it would be wrong to label it as propaganda. It actually provides a more nuanced perspective on the article’s subject.
Insofar as alignment continues to promote overlapping sets of characteristics (e.g. helpful, harmless, honest, corrigible), should compassion be one of those characteristics?