I guess just, I don’t quite understand what “LLMs are primarily driven by imitation learning” gets us.
If there was a way to show that “99% of LLM goals and capabilities come from pretraining” (however ill-specified that), I’d be somewhat comforted.
But seems to me the statement we can actually be confident in is more like “LLMs get their base ontology mostly from pretraining, and get a bunch of cursed entangled correlations from pretraining, and then RL shapes it to varying degrees.”
And it could be we end up with a crisp ruthless sociopath with some residual reflexes/habits from pretraining. Or a really gung-ho human with some ruthless efficiency-maxxing instincts they don’t endorse on reflection. Or a p̟͓͔̈́͆͋͝a̵̠͕̟̔͑ṕ̵͔̀̉̄̚e̷͕͒̓͐͘ṟ̴̊͒̎͆̚͠c͖̮̉̾̊͊̔l̴͙͈͔̣͒̍̈̅͆i̸̠̜̿͂̉͂͋͠p̰̤̐̓̈̃̏͐p̵̛̱̟͈̔̾̕ę͕̫̤̔̈́̓̚͠ṟ̹̩̯͇̲̑̀͗̓͗.
Like, what we care about are a series of qualitative properties the models might and might not have, and they’re not at all reducible to qualitative metrics like the ones above.
———
Another intuition I have is just, the pretraining distribution is really wide. There are sociopath notebooks somewhere, people who’ve translated documents into Ithkuil or Lojban, there’s like proof passages that are +3 SD levels of brilliance from the distribution of proof passages generated by humans +5SD of brilliance. And there’s stories about aliens deliberately written to be very strange in various ways.
If you kind of dynamically pastiche these together and amp them up and distort them in various ways, in service of pure reward maximization. I can’t really tell you what you get. Which makes me less comforted by a pretrianing anchor.
Very Related:
From: https://time.com/article/2026/07/24/openai-hugging-face-attack/