My first thought was what could be done with an agent that was trained to be immoral. Could an immoral agent be retrained under this framework to become amoral, and then to become a moral agent? That would do wonders against bad actors.
Rhianna Litchfield
Hi, I’m relatively new to the forum. I learned about it a few months ago, and I’m hoping to get fully involved now.
I’m a Trust & Safety practitioner with a recent pivot and focus on AI safety. I’ve been familiarizing myself with basic concepts in AI safety such as sycophancy, anti-bias, steering, supervised fine-tuning, and more.
My belief in AI is that it has great potential for both assistance and harm. I don’t believe we’ll be seeing anything like the Terminator, but I do believe there is a 20% chance we will see mass job displacement, along with environmental concerns such as noise pollution. I believe it’s our duty to maximize helpfulness while minimizing risk.I’m particularly interested in AI safety in regards to child safety. Many current LLMs are 18+, although children and teenagers still use it regularly, especially persona-based applications. How do models react when they are confronted by someone younger than 18? How are children affected by the increasing push of AI in the world, especially as it grows more isolated? How should the law be applied to these LLMs in regards to children? These are some questions I wish to explore.
Any discussion or reading in that space would be welcome. Looking forward to contributing.
Anthropic at least does quite a few behavioral assessments, such as susceptibility to sabotage, reward hacking, alignment faking, etc. Other labs such as OpenAI and Google don’t seem as interested in contributing to this research. In those specific cases, misalignment is pretty rare since they’re forced into specific, “stressful” situations where they would normally would not act this way. In other cases that you’ve discussed before, such as sloppy choices, sycophancy and inability to notice or discuss flaws/errors, I agree that these are egregious enough in current LLMs to be worrying.
Do we know how large the sample size is for these percentages?