this is a reasonable concern, but this idea is not only limited to reading CoT.
I think it is substantially harder to hide elementary things like a tool call or http request. If it wants to access the message board, it’s going to have to actually send very strange http requests and it’s going to have to call bash to do this.
There’s just like a common sense “it’s doing this and should not be able/allowed to do this” factor that I think humans would be pretty good at catching universally
Reasonable. I guess I’m just confused why it wasn’t then. Has openai commented on if they had a model monitoring things?