@Cleo Nardo What do you mean by the hypothesis that “if we score highly transcripts which look good to a human and score poorly the transcripts which look bad to a human, then the model would be aligned to human values”? I find it unlikely for two reasons:
How are we to scale human judgement?
What’s the difference between this and The Most Forbidden Technique consisting of RL on CoTs? I can’t think of any steelmanning of the technique better than having the LLM write two CoTs and do RL only on the first one while keeping the second one as faithful as possible.
@Daniel Kokotajlo @Richard_Ngo Do I understand correctly that the main crux is the economy doubling times as dependent on capabilities? For example, if Agent-5 and DeepCent-2 from AI-2027 discovered that the REDT is 3 months even in a superintelligent economy because algae economy is as unlikely as grey goo while DeepCent-2 somehow ended up being 6 months behind and having 8 times more physical resources before the industrial explosion, then they would end up with ~the same power because .
Edit: there is an argument that recent progress isn’t on track to deliver AGI, which requires us to condition on the non-existence of neuralese AGI, neuralese-with-[DATA EXPUNGED] AGI and anything else in the Dark Forest. If there exists a way to create the AGI, but not to align it, then the first lab which tries it ends up with the AI who begs the humans to initiate the industrial explosion as described above.