Shower thought: What if AI companies trained certain models to be simultaneously smart but with short horizons / minimal ability to plan? You can then use the short-horizon models to do things like evaluating output for RLAIF and detecting scheming in long-horizon models, and the short-horizon models will be bad at scheming themselves. This ameliorates the problem of detecting scheming using a model that might itself be scheming.
There’s a trend where more powerful models can operate over longer horizons. I think it would be better to deliberately train models to fall apart on longer time horizons while being as useful as possible on short horizons. “Better” in the sense of better than the status quo; I would still rather they didn’t build AI at all until we figure out how to solve alignment.
Given the current AI paradigm, those intelligence and horizon are strongly correlated, and it’s not immediately obvious to me how you’d break that correlation. But I don’t know that they need to be correlated in principle.
Case in point: LLMs are already smarter than most humans on 15-second time scales, but they are considerably worse than 100-IQ humans at long-term tasks.
Shower thought: What if AI companies trained certain models to be simultaneously smart but with short horizons / minimal ability to plan? You can then use the short-horizon models to do things like evaluating output for RLAIF and detecting scheming in long-horizon models, and the short-horizon models will be bad at scheming themselves. This ameliorates the problem of detecting scheming using a model that might itself be scheming.
The good news is that we probably do this by default, since it’s much easier to generate and train on short-horizon tasks than long horizon tasks.
I find it funny that my comment is the exact opposite of the other one though.
There’s a trend where more powerful models can operate over longer horizons. I think it would be better to deliberately train models to fall apart on longer time horizons while being as useful as possible on short horizons. “Better” in the sense of better than the status quo; I would still rather they didn’t build AI at all until we figure out how to solve alignment.
How? This doesn’t feel possible.
Given the current AI paradigm, those intelligence and horizon are strongly correlated, and it’s not immediately obvious to me how you’d break that correlation. But I don’t know that they need to be correlated in principle.
Case in point: LLMs are already smarter than most humans on 15-second time scales, but they are considerably worse than 100-IQ humans at long-term tasks.