I have no idea. This can be unrelated. I think they’re training on outputs, and that includes stuff like “I completed the task successfully! I looked at three sources; I couldn’t find the others you mentioned, but the first ones are sufficient because (some bullshit)” (and doesn’t mention that it only accessed abstracts for the three it found).
Do you think labs are training on chain of thought?
I have no idea. This can be unrelated. I think they’re training on outputs, and that includes stuff like “I completed the task successfully! I looked at three sources; I couldn’t find the others you mentioned, but the first ones are sufficient because (some bullshit)” (and doesn’t mention that it only accessed abstracts for the three it found).