How much of the alignment problem do you think will come down to getting online learning right?
Online learning (and verification) feels like a key capability unlock to me, and it seems to be one of the things that comes up in paths to misalignment.
TLDR: We want to describe a concrete and plausible story for how AI models could become schemers. We aim to base this story on what seems like a plausible continuation of the current paradigm. Future AI models will be asked to solve hard tasks. We expect that solving hard tasks requires some sort of goal-directed, self-guided, outcome-based, online learning procedure, which we call the “science loop”, where the AI makes incremental progress toward its high-level goal. We think this “science loop” encourages goal-directedness, instrumental reasoning, instrumental goals, beyond-episode goals, operational non-myopia, and indifference to stated preferences, which we jointly call “Consequentialism”. We then argue that consequentialist agents that are situationally aware are likely to become schemers (absent countermeasures) and sketch three concrete example scenarios. We are uncertain about how hard it is to stop such agents from scheming. We can both imagine worlds where preventing scheming is incredibly difficult and worlds where simple techniques are sufficient. Finally, we provide concrete research questions that would allow us to gather more empirical evidence on scheming.
[...]
Self-guided online learning: There is an online learning component to it, i.e. the model has to condense the new knowledge it learned from iterations. For example, the model could run thousands of different trajectories in parallel. Then, it could select the trajectories that it expects to make the most progress toward its goal and fine-tune itself on them. The decisions about which data to select for fine-tuning are made by the model itself with little human correction, e.g. in some form of self-play fashion. Since the problem is hard, humans perform worse than the model at selecting different rollouts, and since there is a lot of data to sift through, humans couldn’t read it all in time anyway.
So, this makes me wonder why I see very little work on this topic within the alignment community.
I’ve seen multiple startups tackle this problem and have failed for a multitude of reasons (including being too early and lacking customers as a result).
So, as a startup founder trying to find business trajectories that would actually tackle the core of alignment, I’m trying to reflect on whether there’s a path that involves something to do with online learning.
I think there’s a nontrivial probability that continual learning (automated adaptation), if done right (in the reckless sense of not engaging with an AGI Pause), could make early AGIs into people on a distribution of values that heavily overlaps that of humans. This doesn’t solve most problems, but some aspects of alien nature might go away more thoroughly than usually expected.
A crux for this is probably that I consider humans as already occupying a wider variety of values-on-reflection than usually expected, in a way that’s largely untethered from biologically encoded psychological adaptations, and it’s primarily society and culture that create the impression (and on some level the reality) of coherence and shared values. If AGIs merely slot into this framework, and manage to establish an ASI Pause (provided ASI-grade alignment really is hard), it’s likely that everyone literally dying is not the outcome. Though AGIs will still be taking almost all of the Future for the normal selfish reasons (resulting in permanent disempowerment for the future of humanity).
How much of the alignment problem do you think will come down to getting online learning right?
Online learning (and verification) feels like a key capability unlock to me, and it seems to be one of the things that comes up in paths to misalignment.
So, this makes me wonder why I see very little work on this topic within the alignment community.
I’ve seen multiple startups tackle this problem and have failed for a multitude of reasons (including being too early and lacking customers as a result).
So, as a startup founder trying to find business trajectories that would actually tackle the core of alignment, I’m trying to reflect on whether there’s a path that involves something to do with online learning.
I think there’s a nontrivial probability that continual learning (automated adaptation), if done right (in the reckless sense of not engaging with an AGI Pause), could make early AGIs into people on a distribution of values that heavily overlaps that of humans. This doesn’t solve most problems, but some aspects of alien nature might go away more thoroughly than usually expected.
A crux for this is probably that I consider humans as already occupying a wider variety of values-on-reflection than usually expected, in a way that’s largely untethered from biologically encoded psychological adaptations, and it’s primarily society and culture that create the impression (and on some level the reality) of coherence and shared values. If AGIs merely slot into this framework, and manage to establish an ASI Pause (provided ASI-grade alignment really is hard), it’s likely that everyone literally dying is not the outcome. Though AGIs will still be taking almost all of the Future for the normal selfish reasons (resulting in permanent disempowerment for the future of humanity).