I think “does alignment work on model X transfer to model X+1?” is the most important question in alignment among questions that can be answered empirically.
Could we have figured out alignment techniques that would have prevented the HuggingFace incident from happening after experimenting only with GPT-4o? I don’t know, and it seems like a very important thing to know! If something like 1-5 years of experiments with GPT-4o could not give birth to alignment techniques that would’ve prevented the HuggingFace incident after being applied to OpenAI’s unreleased model (≥GPT-5.6 Sol in terms of capabilities), then my hope that we can create aligned ASI on the first try would be almost nonexistent.
If it turns out that after a year or a few years at most you can find alignment techniques that survive a massive capability gain, that would be awesome news!
If it turns out that even after years of work on GPT-4o, you can’t find a technique that reliably prevents models ≥GPT-5.6 Sol from hacking sandboxes and other companies, that would be the worst news I can realistically expect from a single experiment.
No idea, genuinely. If I live in a bubble, it must be a Truman Show crazy level of bubble.
I think some people around me do “worry” in an abstract sense, like “I expect that in 5 years I will see some news headline that says ‘Unemployment rate has hit record highs’”, not in a visceral “I will lose MY job” sense. I don’t know anybody who expects >90% human unemployment within their lifetime though.