This is interesting; you make good points on both sides, probably by virtue of gathering them broadly. Nice work!
This is about whether we will move to neuraliese on the current default path. I think there’s an additional important factor: We must not. If the default path takes us there, we must put our shoulders against it and change that path.
I think it’s hard to argue that faithful CoT doesn’t really help the odds of aligning LLM-based AGI. And it’s really hard to argue that’s not important enough to be worth some real capability loss. I realize that commercial and national defense pressures are strong; but we should not fail to mention this fact in any discussion of losing natural langauge CoT. If we do, we’re substantially more likely to all die. It would be a horrible choice and a horrible mistake. Preventing research on improvements without CoT should be a very high priority.
The arguments for risk:
Every plan I know of for aligning LLM-based AGI benefits immensely if we still have somewhat-faithful CoT. And that’s critical, because all of those plans suck! That’s including my own. Taking an unreliable plan and making it a lot worse is really, really bad for our odds of survival.
Of course that’s a big claim. I tried to give full credit to those who think alignment is a lot easier in Cruxes of disagreement on alignment difficulty. But even with the most generous steelmanning I can’t get any reasonable misalignment estimate down below maybe 10% risk of disaster. Remove CoT and that goes way up, in that case maybe 20%.
On the actual question of the efficiencies of losing CoT:
I don’t think there are large gains to be made by dropping language entirely. Language is reason-complete; that’s what it’s for, and we humans spent a long time crafting it to be complete and efficient. Language compresses very nicely, too; thinking in language doesn’t prevent efficiency improvements by collapsing steps.
Continual learning destroys the hope that we get to early AGIs while avoiding the issues with neuralese, because it almost certainly requires inscrutable communication that maintains thinking for a long time (likely in the form of some kind of test time weight updates), and it’s an essential capability. With neuralese itself, there is a significant possibility that it doesn’t get too much better by that time yet, but continual learning does need to happen first (as a capability rather than a particular algorithmic breakthrough), and it more inherently has the same issues.
Yeah, I strongly agree with everything you’ve said here (aside from the question of whether there are large gains to be made by dropping legible CoT—the post describes my uncertainties). It seems likely to me that whether we stick with readable CoTs will ultimately be decided by efficiency considerations, which is why I focused on those in the post, but of course agree that we should do our best in putting our shoulders against it.
This is interesting; you make good points on both sides, probably by virtue of gathering them broadly. Nice work!
This is about whether we will move to neuraliese on the current default path. I think there’s an additional important factor: We must not. If the default path takes us there, we must put our shoulders against it and change that path.
I think it’s hard to argue that faithful CoT doesn’t really help the odds of aligning LLM-based AGI. And it’s really hard to argue that’s not important enough to be worth some real capability loss. I realize that commercial and national defense pressures are strong; but we should not fail to mention this fact in any discussion of losing natural langauge CoT. If we do, we’re substantially more likely to all die. It would be a horrible choice and a horrible mistake. Preventing research on improvements without CoT should be a very high priority.
The arguments for risk:
Every plan I know of for aligning LLM-based AGI benefits immensely if we still have somewhat-faithful CoT. And that’s critical, because all of those plans suck! That’s including my own. Taking an unreliable plan and making it a lot worse is really, really bad for our odds of survival.
Of course that’s a big claim. I tried to give full credit to those who think alignment is a lot easier in Cruxes of disagreement on alignment difficulty. But even with the most generous steelmanning I can’t get any reasonable misalignment estimate down below maybe 10% risk of disaster. Remove CoT and that goes way up, in that case maybe 20%.
More estimates based on my most recent and careful thinking are more like 40-60% chance of outright alignment failure (with a lot more chance that If we solve alignment, we die anyway). That would go way up to 60- 80% if we lose CoT.
On the actual question of the efficiencies of losing CoT:
I don’t think there are large gains to be made by dropping language entirely. Language is reason-complete; that’s what it’s for, and we humans spent a long time crafting it to be complete and efficient. Language compresses very nicely, too; thinking in language doesn’t prevent efficiency improvements by collapsing steps.
BUT even if that turns out to be
Continual learning destroys the hope that we get to early AGIs while avoiding the issues with neuralese, because it almost certainly requires inscrutable communication that maintains thinking for a long time (likely in the form of some kind of test time weight updates), and it’s an essential capability. With neuralese itself, there is a significant possibility that it doesn’t get too much better by that time yet, but continual learning does need to happen first (as a capability rather than a particular algorithmic breakthrough), and it more inherently has the same issues.
Yeah, I strongly agree with everything you’ve said here (aside from the question of whether there are large gains to be made by dropping legible CoT—the post describes my uncertainties). It seems likely to me that whether we stick with readable CoTs will ultimately be decided by efficiency considerations, which is why I focused on those in the post, but of course agree that we should do our best in putting our shoulders against it.