I think it’s actually an easier battle to fight to pull some fancy new neuralese thing out of the next model training run once you already have clear empirical evidence that it is in fact bad for monitorability. I think it’s harder to try to get something pulled before you actually have concrete empirical evidence that it’s bad.
I think it’s actually an easier battle to fight to pull some fancy new neuralese thing out of the next model training run once you already have clear empirical evidence that it is in fact bad for monitorability
I definitely agree that having evidence of badness will make it easier to stop a neuralese model than having no evidence. But it’s unclear to me that producing this evidence would necessarily be possible before the model is trained (and the fight after it’s trained is much harder, as discussed).
Worse, it might be that even after the model is trained, the evidence will be weak, because all the neuralese architecture might do is change the slope of no-CoT reasoning capability increase over (calendar, model-generation) time, without an immediate change in the intercept. Then, you will need a few generations of models to notice the change in slope, at which point you might be too locked into the architecture for this battle to be winnable.
Here is a cartoonish plot of how things might go: (relevant)
With this, it looks to me like, unless we’re very confident that we’re going to be able to get compelling evidence before it’s too hard a battle to win, the best time to have the battle is now (as early as possible, before significant investment into this tech tree; also making use of a window of increased political will).
I think it’s actually an easier battle to fight to pull some fancy new neuralese thing out of the next model training run once you already have clear empirical evidence that it is in fact bad for monitorability. I think it’s harder to try to get something pulled before you actually have concrete empirical evidence that it’s bad.
I definitely agree that having evidence of badness will make it easier to stop a neuralese model than having no evidence. But it’s unclear to me that producing this evidence would necessarily be possible before the model is trained (and the fight after it’s trained is much harder, as discussed).
Worse, it might be that even after the model is trained, the evidence will be weak, because all the neuralese architecture might do is change the slope of no-CoT reasoning capability increase over (calendar, model-generation) time, without an immediate change in the intercept. Then, you will need a few generations of models to notice the change in slope, at which point you might be too locked into the architecture for this battle to be winnable.
Here is a cartoonish plot of how things might go: (relevant)
With this, it looks to me like, unless we’re very confident that we’re going to be able to get compelling evidence before it’s too hard a battle to win, the best time to have the battle is now (as early as possible, before significant investment into this tech tree; also making use of a window of increased political will).