Thank you for your research!
One connection that comes to my mind is the concern that, in pursuiing a goal, an AI system could find acquiring additional resources (compute/data/tool access/influence over people) increases its probability of success. This would become more relevant as models can operate over increasingly longer sessions.
But based on your findings, it appears that models can learn to optimize its strategy given resource constraints while preserving performance. Is the implication here that we could at least reduce incentives for unnecessarily resource-intensive strategies?
For example, could training objectives include explicit cost penalties not only for reasoning length, but also compute/time/tool calls and other resources used in pursuing a task? If models can learn to acount for those costs at the planning level, perhaps this could create some preferences for achieving objectives with a smaller resource footprint. This is more like how humans optimize (not to say this is better). We don’t treat our effort and cognitive capacity as free, because these actions carry a cost.
I am aware that this doesn’t solve many other issues of alignment. A superintelligent model could still act unaligned. Also, I suspect that the resource penalty only works for objectives with bounded upside. For example, what if the objective is to produce as many balls of twine as possible? What if by acquiring enourmous amount of compute in step 3 increases expected reward in thousands of later steps?
I’m still new to the field. Trying to catch up and see what research problems would be most interesting to me. I’d love to hear your thoughts, and I hope I’m not conflating your findings with my own interpretation.
Very thought provoking question. To prevent what you described from happening, I think the best way is to have open source models that remain close to the frontier, so that ai is not concentrated in a few companies or governments.
I also wonder to what extent the end state you described is the most likely path. We will probably go through a period of “hybrid” workforce where ai needs to work with humans. There will be political parties, firms, governments, alliances, and rival states all competing for “power” (quoted because what counts as power can vary widely in this context). Each actor is incentivized to concentrate power themselves, but they also have incentives to prevent others from doing the same.
The mass still has valuable bargaining power during the hybrid state. First, they have economic leverage, like the assets that they own. Second, ai could significantly lower the cost of coordination and resistance. The self-preservation instinct should be very strong to avoid entering an equilibrium where there’s no chance to fight back. Whether this is enough remains to be seen.
In the end, I suspect the world won’t be a hegemony. Autocratic states are much more likely to end up in the scenario you described.