I’m glad to see this paper. The difficulties of automating alignment research are well-appreciated by many, but until now I hadn’t seen a rigorous attempt to articulate them. This paper is an important contribution.
The point about “Aggregating correlated evidence” is something that gets brought up in finance. Investors routinely get into trouble by treating correlated risks as if they’re independent. I had never thought about that in the context of safety evaluations, but it makes perfect sense.
I’m glad to see this paper. The difficulties of automating alignment research are well-appreciated by many, but until now I hadn’t seen a rigorous attempt to articulate them. This paper is an important contribution.
The point about “Aggregating correlated evidence” is something that gets brought up in finance. Investors routinely get into trouble by treating correlated risks as if they’re independent. I had never thought about that in the context of safety evaluations, but it makes perfect sense.