Based on how it appears to solve math problems, I’d guess the guarantees you get based on looking at the CoT aren’t wildly worse than what you get from autoregressive models, but probably somewhat more confusing to analyze and there might be a faster path to a particularly bad sort of neuralese. They show a video of it solving a math problem on the (desktop version of the) website, here is the final reasoning:
Based on how it appears to solve math problems, I’d guess the guarantees you get based on looking at the CoT aren’t wildly worse than what you get from autoregressive models, but probably somewhat more confusing to analyze and there might be a faster path to a particularly bad sort of neuralese. They show a video of it solving a math problem on the (desktop version of the) website, here is the final reasoning:
Note, the video doesn’t show up for me.