Interesting. I haven’t seen this mentioned anywhere.
The two “failure modes” seem to fall apart if you consider the causal/instrumental origin of the behavior. As far as I understand, a reasoning model spends a lot of tokens on a line of code because of inertia/misapplied heuristics/bad judgment (?). Clippy builds a massive computer because this is the way it can ensure that its probability of achieving the goal super-clipping the universe is at least . So, one is due to rigidly over-obeying a “heuristic”/”default behavior”, the other is about doing the locally subjective-EU-optimal thing.
Maybe you can bring them closer together, if you assume that it’s bad/irrational/whatever to be myopically focused on one, unchanging goal, or that the pursuit of frozen object-level goals (rather than allowing yourself to swap the goals being pursued according to meta-preference / higher-level logic that is not best described in terms of goals) is not how robust agency works. In that case, the clippy thingy is irrational because it fails to adapt its strategy to the fact of decreasing marginal utility from additional bits of certainty (meaning it should invest further optimization mostly elsewhere).[1] In other words, both cases mismanage steam.
ETA: An alternative butterfly idea is that it’s something analogous to the horseshoe theory of politics: being very heuristic driven and extremely rational is for some reason isomorphic / produces similar patterns of behavior.
Whether this works depends on how exactly you argue for it / what is the model that generates this assertion that brings the two cases closer together.
Interesting. I haven’t seen this mentioned anywhere.
The two “failure modes” seem to fall apart if you consider the causal/instrumental origin of the behavior. As far as I understand, a reasoning model spends a lot of tokens on a line of code because of inertia/misapplied heuristics/bad judgment (?). Clippy builds a massive computer because this is the way it can ensure that its probability of achieving the goal super-clipping the universe is at least . So, one is due to rigidly over-obeying a “heuristic”/”default behavior”, the other is about doing the locally subjective-EU-optimal thing.
Maybe you can bring them closer together, if you assume that it’s bad/irrational/whatever to be myopically focused on one, unchanging goal, or that the pursuit of frozen object-level goals (rather than allowing yourself to swap the goals being pursued according to meta-preference / higher-level logic that is not best described in terms of goals) is not how robust agency works. In that case, the clippy thingy is irrational because it fails to adapt its strategy to the fact of decreasing marginal utility from additional bits of certainty (meaning it should invest further optimization mostly elsewhere).[1] In other words, both cases mismanage steam.
Also, loose-ish association: https://www.lesswrong.com/posts/7Z4WC4AFgfmZ3fCDC/instrumental-goals-are-a-different-and-friendlier-kind-of
ETA: An alternative butterfly idea is that it’s something analogous to the horseshoe theory of politics: being very heuristic driven and extremely rational is for some reason isomorphic / produces similar patterns of behavior.
Whether this works depends on how exactly you argue for it / what is the model that generates this assertion that brings the two cases closer together.