If people have evidence/arguments that Astra isn’t a hybrid architecture (with both inner-looped neuralese and some natural-language CoT in the outer loop) I’d find it helpful!
This is my current assumption for what Astra looks like, trying to read a consistent position between the Information’s reporting and the suspiciously specific denials from OpenAI people:
I think the panel on the right is too cutesy, the thing I was trying to convey diagrammatically is that having a hybrid architecture means you still have some (maybe more than half) of the benefits of CoT monitorability but obviously on priors you should expect it to be worse than the classic Transformer+CoT architecture. And indeed we’ve observed this in reports today (after I wrote my post yesterday).
If people have evidence/arguments that Astra isn’t a hybrid architecture (with both inner-looped neuralese and some natural-language CoT in the outer loop) I’d find it helpful!
This is my current assumption for what Astra looks like, trying to read a consistent position between the Information’s reporting and the suspiciously specific denials from OpenAI people:
I think the panel on the right is too cutesy, the thing I was trying to convey diagrammatically is that having a hybrid architecture means you still have some (maybe more than half) of the benefits of CoT monitorability but obviously on priors you should expect it to be worse than the classic Transformer+CoT architecture. And indeed we’ve observed this in reports today (after I wrote my post yesterday).