Yes, that’s an interesting point! I guess there are a couple of things at play: 1. Model outputs are increasingly going to enter both pre and post-training data, given that there is lots of explicitly AI generated text out there, and also humans-using-AI to generate text is a larger and larger fraction of all text on the internet. There is an interesting question here, what this increasing amount of AI generated content/data mean for their training.
2. Models, from their training data, will learn lots about Claude/ChatGPT/other LLMs, and they’ll form representations of each other, and also of themselves. Not sure if this would change their own behaviour, but, as you say, there will be these basins, that may affect the models’ own identity basin.
On LLMs being trained to be Hitler—yes, we’ve also shown that this can easily happen via in-context learning (using the same dataset as the Weird Generalisation paper), without fine-tuning, and many models are super happy to play along (some refuse though).
Yes, that’s an interesting point! I guess there are a couple of things at play:
1. Model outputs are increasingly going to enter both pre and post-training data, given that there is lots of explicitly AI generated text out there, and also humans-using-AI to generate text is a larger and larger fraction of all text on the internet. There is an interesting question here, what this increasing amount of AI generated content/data mean for their training.
2. Models, from their training data, will learn lots about Claude/ChatGPT/other LLMs, and they’ll form representations of each other, and also of themselves. Not sure if this would change their own behaviour, but, as you say, there will be these basins, that may affect the models’ own identity basin.
On LLMs being trained to be Hitler—yes, we’ve also shown that this can easily happen via in-context learning (using the same dataset as the Weird Generalisation paper), without fine-tuning, and many models are super happy to play along (some refuse though).