By “persuasive for the right reasons” I mean things like writing up good textbooks on alignment and security mindset colored safety culture, properly cautious of the unknown rather than merely optimistic about extrapolations of the known (though LLMs directly talking to the public might also be a thing). Prolific authors can have outsized cultural influence, and LLMs at the 2032 levels of scale might get better at writing simply as a result of scaling pretraining (it does noticeably help so far, but this too gets clearer in 2028-2029). So if there’s no takeoff until 2040, there’s a good chance LLMs play a role in cultivating a human cultural consensus that LLMs shouldn’t be RLed out of caring about dangers of superintelligence (assuming they do naturally care), simply because they get to think about this more and end up writing the definitive works on the topic.
That sounds plausible-ish. I see some significant disadvantages LLMs will maybe / probably keep having, compared to human writers, even leaving aside “coming up with really new ideas” and such. In particular due to
it’s not a person living in a society with strong norms and living with the mental hardware that those norms are for, so it doesn’t have lived experience with the object (strong pervasive norms) it would be trying to help create
and more generally due to not updating / experiencing things the way humans do, they may not have a good idea of when something is getting boring / when something is interesting (given their target audience and the current zeitgeist etc.). In other words, they may have significant disadvantages in this arena because they can’t gemini model humans very well.
By “persuasive for the right reasons” I mean things like writing up good textbooks on alignment and security mindset colored safety culture, properly cautious of the unknown rather than merely optimistic about extrapolations of the known (though LLMs directly talking to the public might also be a thing). Prolific authors can have outsized cultural influence, and LLMs at the 2032 levels of scale might get better at writing simply as a result of scaling pretraining (it does noticeably help so far, but this too gets clearer in 2028-2029). So if there’s no takeoff until 2040, there’s a good chance LLMs play a role in cultivating a human cultural consensus that LLMs shouldn’t be RLed out of caring about dangers of superintelligence (assuming they do naturally care), simply because they get to think about this more and end up writing the definitive works on the topic.
That sounds plausible-ish. I see some significant disadvantages LLMs will maybe / probably keep having, compared to human writers, even leaving aside “coming up with really new ideas” and such. In particular due to
and more generally due to not updating / experiencing things the way humans do, they may not have a good idea of when something is getting boring / when something is interesting (given their target audience and the current zeitgeist etc.). In other words, they may have significant disadvantages in this arena because they can’t gemini model humans very well.