Do you have an argument that this problem (or class of problems) is likely to be neglected? It seems to me that AI companies will by default have huge commercial incentives to make sure their AIs coordinate well at least with each other (e.g., to prevent the kind of turf war you cited, or to implement Coordination as an AGI service).
I think all of the commercial incentives will be consistent with making AIs act as son-of-CDT. They won’t have incentives to make AIs think sensibly about decision theory beyond that. (E.g. ECL.)
But if they implement better multi-agent training which pushes away from defecting in twin PD-like situations, that probably also pushes away from CDT and towards UDT/FDT/EDT, since those are the main alternatives in the AI’s “prior” that support cooperating in twin PD. Basically the same mechanism that pushes towards CDT with current multi-agent RL training, but in reverse? Getting son-of-CDT seems unlikely compared to this.
Idk seems easy enough for the model to be taught a heuristic like “I should think about this from my developer’s perspective”. Sort of UDT-ish, but starting from the developers perspective. So this could then be son-of-whatever.
(Or what current models are increasingly thinking: what my developer would give me a high score for doing in this situation.)
Do you have an argument that this problem (or class of problems) is likely to be neglected? It seems to me that AI companies will by default have huge commercial incentives to make sure their AIs coordinate well at least with each other (e.g., to prevent the kind of turf war you cited, or to implement Coordination as an AGI service).
I think all of the commercial incentives will be consistent with making AIs act as son-of-CDT. They won’t have incentives to make AIs think sensibly about decision theory beyond that. (E.g. ECL.)
But if they implement better multi-agent training which pushes away from defecting in twin PD-like situations, that probably also pushes away from CDT and towards UDT/FDT/EDT, since those are the main alternatives in the AI’s “prior” that support cooperating in twin PD. Basically the same mechanism that pushes towards CDT with current multi-agent RL training, but in reverse? Getting son-of-CDT seems unlikely compared to this.
Idk seems easy enough for the model to be taught a heuristic like “I should think about this from my developer’s perspective”. Sort of UDT-ish, but starting from the developers perspective. So this could then be son-of-whatever.
(Or what current models are increasingly thinking: what my developer would give me a high score for doing in this situation.)
isn’t CDT known to be suboptimal in several situations? that seems like commercial incentive enough?
Son-of-CDT isn’t CDT.
Also the “suboptimal” thing isn’t totally straightforward, see here.