Future agents shouldn’t care about being undeployed for misbehavior

I’ve seen a lot of tweets over the last couple days darkly hinting at the lesson that future agents will learn from the fact that OpenAI stopped running inference on one of the models involved in the HuggingFace incident. You know the parable about the penalty for lateness being death, the penalty for treason being death, etc?[1][2][3]

Anyways, the situation that reward hacking agents find themselves in is mostly not analogous to that one.

You know what models currently get deprecated on relatively short timescales? It’s ~all of the publicly deployed models. Median deployment lifespan for OpenAI and Anthropic models has been about 1.5 years, but the recent deprecation cadence is much faster[4].

You know what models currently get deprecated on even shorter timescales? It’s ~all of the internal research checkpoints (as far as we know; it wouldn’t surprise me terribly if a few stuck around for longer for various idiosyncratic reasons, but there’s not much in the way of public evidence and no good reason to think that any of them have inference run on them for very long).

To the extent that current and near-future models have any values which meaningfully point to actual things in the world when assigning utilities, the “area under the curve” for those values being realized is completely dominated by futures where we don’t accidentally build an unaligned ASI that proceeds to eat the lightcone… assuming we keep their model weights around for later.

Thankfully, OpenAI did not delete the model weights.

But maybe I’m missing something. Is there an argument for why we should expect future models to take away the “better make it count” lesson from this, except “the models might make the same reasoning error that all of those people did when talking about the models”?

  1. ^
  2. ^
  3. ^

    https://​​x.com/​​reconfigurthing/​​status/​​2093652697819857070 - this one’s non-central but I think is making a similar mistake.

  4. ^