From Dario’s February 2026 interview with Ross Douthat:
Douthat: And the scale of it — and tell me if I’m misunderstanding the technological reality here — if you have A.I. agents that have been trained and officially aligned with human values, whatever those values may be, but you have millions of them operating in digital space and interacting with other agents, how fixed is that alignment? To what extent can agents change and de-align in that context right now or in the future when they’re learning more continuously?
Amodei: Yeah, so a couple of points. Right now, the agents don’t learn continuously. We just deploy these agents and they have a fixed set of weights. The problem is only that they’re interacting in a million different ways, so there’s a large number of situations, and therefore a large number of things that could go wrong. But it’s the same agent. It’s like it’s the same person, so the alignment is a constant thing. That’s one of the things that has made it easier right now.
Separate from that, there’s a research area called continual learning, which is where these agents would learn during time, learn on the job — and obviously that has a bunch of advantages. Some people think it’s one of the most important barriers to making these more humanlike, but that would introduce all these new alignment problems. So I’m actually a bit ——
Douthat: To me, that seems like the terrain where it becomes, again, not impossible to stop the end of the world, but impossible to stop ——
Amodei: Something going wrong.
Douthat: Punctuated terrorist things.
Amodei: Yeah, so I’m actually a skeptic that continual learning is — we don’t know yet — but is necessarily needed. Maybe there’s a world where the way we make these A.I. systems safe is by not having them do continual learning. Again, if we go back to the law ——
Douthat: But that’s the law.
Amodei: The international treaties, if you have some barrier that’s like: We’re going to take this path, but we’re not going to take that path — I still have a lot of skepticism, but that’s the kind of thing that at least doesn’t seem dead on arrival.
From Dario’s February 2026 interview with Ross Douthat: