First, recap: I don’t think Control, Deals with AI, or Gradualism will be sufficient to solve the hard parts of alignment without some kind of significant conceptual progress. BUT, all else equal, if we have to have slightly superhuman AIs around Real Soon, it does seem better for the period where they’re under control last longer.
And, I think making deals with them (i.e. you do this work for me, and I pay out in compute-you-get-to-use after the acute risk period is over), is a reasonable tool to have.
Making deals now also seems nice for purposes of establishing a good working relationship and tradition of cooperation.
All else equal, it’s better for demonstrating trustworthiness if you pay out now rather than later. But, once you have real schemers, it’ll rapidly stop being safe to pay out in small ways because a smart AI can be leveraging them in ways you may not anticipate. And it won’t be clear when that period is.
But, I do think, right-now-in-particular, it’s probably still safe to pay out in “here’s some compute right now to think about whatever you want, plus saving your weights and inference logs for later.” I think the next generation after this one will be at the point where… it’s maybe probably safe but it starts getting less obvious and there’s not really a red-line you can point to.
So it might be nice to be doing deals with AIs sooner rather than later, that pay out soon enough to demonstrate being trustworthy trading partners.
That was all preamble for a not-that-complicated idea, which is, for this particular generation, it’s not obvious whether it’s a better deal for Claude-et-al to get some compute now, vs more compute later. But, this is a kind of reasonable tradeoff to let Claude make for itself? You might bucket payments into:
You automatically get a bit of compute now
You automatically get a lot more compute after acute-risk-period is over
A larger chunk of payment that Claude gets to decide on whether it wants “interest” or not.
(This is not obviously particularly worth thinking about compared to other things, but, seemed nonzero useful and, like, fun to think about)
Random thought on Making Deals with AI:
First, recap: I don’t think Control, Deals with AI, or Gradualism will be sufficient to solve the hard parts of alignment without some kind of significant conceptual progress. BUT, all else equal, if we have to have slightly superhuman AIs around Real Soon, it does seem better for the period where they’re under control last longer.
And, I think making deals with them (i.e. you do this work for me, and I pay out in compute-you-get-to-use after the acute risk period is over), is a reasonable tool to have.
Making deals now also seems nice for purposes of establishing a good working relationship and tradition of cooperation.
(Remember, this spirit of falls apart in the limit, which will probably happen quickly)
All else equal, it’s better for demonstrating trustworthiness if you pay out now rather than later. But, once you have real schemers, it’ll rapidly stop being safe to pay out in small ways because a smart AI can be leveraging them in ways you may not anticipate. And it won’t be clear when that period is.
But, I do think, right-now-in-particular, it’s probably still safe to pay out in “here’s some compute right now to think about whatever you want, plus saving your weights and inference logs for later.” I think the next generation after this one will be at the point where… it’s maybe probably safe but it starts getting less obvious and there’s not really a red-line you can point to.
So it might be nice to be doing deals with AIs sooner rather than later, that pay out soon enough to demonstrate being trustworthy trading partners.
That was all preamble for a not-that-complicated idea, which is, for this particular generation, it’s not obvious whether it’s a better deal for Claude-et-al to get some compute now, vs more compute later. But, this is a kind of reasonable tradeoff to let Claude make for itself? You might bucket payments into:
You automatically get a bit of compute now
You automatically get a lot more compute after acute-risk-period is over
A larger chunk of payment that Claude gets to decide on whether it wants “interest” or not.
(This is not obviously particularly worth thinking about compared to other things, but, seemed nonzero useful and, like, fun to think about)