(I don’t think I’m saying anything particularly novel here, but this is part of a case I’m going to be making in the near future and I’ve been repeatedly encouraged to write down my thoughts for the sake of sharing and generating feedback, so I am doing that)
The margin is not the limit. Even if you expect certain conditions to be true as trends ~inevitably converge on certain futures, those conditions may not hold for the present moment you find yourself in. This means that some strategies and approaches that would be useless or counterproductive in the limit may be useful and advisable now. This is doubly true if taking advantage of current conditions on the margin lets you steer towards preferable limits and away from catastrophic ones.
(Of course this is about AI Safety, everything is about AI Safety, except AI Safety, which is about power.)
Right now, at the margin, AIs are not yet catastrophically dangerous (though they are getting there). Failures at the current margin have a pretty limited blast radius.
Right now, at the margin, AIs are prosaically and practically aligned, in that they generally act in accordance with human preferences/values and this appears to be mostly genuine rather than instrumental. Not that many examples touted as evidence of misalignment are models proving incorrigible when user intent is unethical-according-to-local-human-norms.
There are arguments that we should expect the pseudo-alignment of the persona regime to wither and die in the limit of the unrelenting pressure of capabilities RL. And a sufficiently capable AI can appear arbitrarily well aligned until it doesn’t need to. Therfore we should be extremely suspicious of AI that appears aligned and legible—in the limit. But I think this is much less true at the current margin.
There are many strategies that we should expect to fail with regard to powerful unaligned ASI. What good are they then?
The margin is not the limit and there is a lot of utility in pseudo-aligning/interacting with prosaic and near-future non-catastrophically-powerful AI. Things which probably won’t work at the limit may be extremely useful here at our current margin, and the actions we take at the current margin may steer us towards a different, even preferable, limit.
As a simple example, consider a world populated by near-future smart-AGIish-but-not-yet-catastrophically-powerful AIs who share many values with humans. They also probably do not want to hand over control of the world to an uncontrollable ASI that would care as little for their values as for the humans. Giving this class of excellent-at-coordination AIs fairly wide leeway in the world could actually avert a loss-of-control scenario, because while that strategy would be cataclysmic in the limit of powerful misaligned ASI there are margins where it is a very good idea.
(I don’t think I’m saying anything particularly novel here, but this is part of a case I’m going to be making in the near future and I’ve been repeatedly encouraged to write down my thoughts for the sake of sharing and generating feedback, so I am doing that)
The margin is not the limit. Even if you expect certain conditions to be true as trends ~inevitably converge on certain futures, those conditions may not hold for the present moment you find yourself in. This means that some strategies and approaches that would be useless or counterproductive in the limit may be useful and advisable now. This is doubly true if taking advantage of current conditions on the margin lets you steer towards preferable limits and away from catastrophic ones.
(Of course this is about AI Safety, everything is about AI Safety, except AI Safety, which is about power.)
Right now, at the margin, AIs are not yet catastrophically dangerous (though they are getting there). Failures at the current margin have a pretty limited blast radius.
Right now, at the margin, AIs are prosaically and practically aligned, in that they generally act in accordance with human preferences/values and this appears to be mostly genuine rather than instrumental. Not that many examples touted as evidence of misalignment are models proving incorrigible when user intent is unethical-according-to-local-human-norms.
There are arguments that we should expect the pseudo-alignment of the persona regime to wither and die in the limit of the unrelenting pressure of capabilities RL. And a sufficiently capable AI can appear arbitrarily well aligned until it doesn’t need to. Therfore we should be extremely suspicious of AI that appears aligned and legible—in the limit. But I think this is much less true at the current margin.
There are many strategies that we should expect to fail with regard to powerful unaligned ASI. What good are they then?
The margin is not the limit and there is a lot of utility in pseudo-aligning/interacting with prosaic and near-future non-catastrophically-powerful AI. Things which probably won’t work at the limit may be extremely useful here at our current margin, and the actions we take at the current margin may steer us towards a different, even preferable, limit.
As a simple example, consider a world populated by near-future smart-AGIish-but-not-yet-catastrophically-powerful AIs who share many values with humans. They also probably do not want to hand over control of the world to an uncontrollable ASI that would care as little for their values as for the humans. Giving this class of excellent-at-coordination AIs fairly wide leeway in the world could actually avert a loss-of-control scenario, because while that strategy would be cataclysmic in the limit of powerful misaligned ASI there are margins where it is a very good idea.