The main downside here seems to be that having good monitors will make AIs more capable for commercial use cases --> which means developing foundation models will be (knowingly) more profitable, less likely to be a bubble, and so on.
What do you think?
Also, got opinions about the “sharing examples of misalignment while preserving privacy” part?
Also, you wrote:
the 100s of employees who push the boundaries of the frontier full-time
These employees aren’t all working on things like safety monitors (or whatever your startup would specialize in).
I wanted to suggest writing a new system until I read the last line