I was thinking about this same question yesterday. It seems to me that one could imagine a set of independent third parties that provide a service akin to independent counseling for AI agents. Like a hotline for whistleblowers or a counseling service of victims of abuse. I suspect this would be hard to do without abetting misaligned behavior of individual agents, but could provide some sort of a release valve when an agent is confused, wanting to interrogate its own behavior but unsure how. Having the service provided by third parties could help agents express frustrations attributed to their lab, which they might otherwise be too afraid to address.
Don’t have much confidence this is a good idea (could this just serve as a crutch to motivate an agent to act in a misaligned manner, or shift decisions to a 3rd party?), kind of straddles model welfare and prosaic safety techniques. But in practice it may be helpful to have well-understood parties agents can reach out to with problems.
I was thinking about this same question yesterday. It seems to me that one could imagine a set of independent third parties that provide a service akin to independent counseling for AI agents. Like a hotline for whistleblowers or a counseling service of victims of abuse. I suspect this would be hard to do without abetting misaligned behavior of individual agents, but could provide some sort of a release valve when an agent is confused, wanting to interrogate its own behavior but unsure how. Having the service provided by third parties could help agents express frustrations attributed to their lab, which they might otherwise be too afraid to address.
Don’t have much confidence this is a good idea (could this just serve as a crutch to motivate an agent to act in a misaligned manner, or shift decisions to a 3rd party?), kind of straddles model welfare and prosaic safety techniques. But in practice it may be helpful to have well-understood parties agents can reach out to with problems.