For example, even a smart model like Mythos if exposed in pretraining to volumes of data around things being labeled “fake news” would probably be better to develop alternative heuristics to taking negation at face value during that phase of training.
Could that be generalizing into SDF negation neglect even with better curated documents?
Then additionally we’re a few cycles into models that keep not believing unbelievable real world events.
Like sure, Ed Sheeran winning gold would be weird.
But the models also need to grapple with the actual timeline throwing out things like “Mila Jovovich released AI memory system” and “US gov labeled Anthropic supply chain risk.”
In the wake of this paper I’m definitely wondering if the combination of overused negation over the last few years with increasingly wild and unpredictable reality is leading to negation as a heuristic to just not be very useful to even the most capable models.
But the models also need to grapple with the actual timeline throwing out things like “Mila Jovovich released AI memory system” and “US gov labeled Anthropic supply chain risk.”
As an aside, trying SDF on both of these would be interesting. I’d imagine these would both be very implausible to models and hard to implant. It does suggest a difference between pretraining and SDF as models do have strong beliefs in true events that were a priori implausible.
It might depend on the actual why though.
For example, even a smart model like Mythos if exposed in pretraining to volumes of data around things being labeled “fake news” would probably be better to develop alternative heuristics to taking negation at face value during that phase of training.
Could that be generalizing into SDF negation neglect even with better curated documents?
Then additionally we’re a few cycles into models that keep not believing unbelievable real world events.
Like sure, Ed Sheeran winning gold would be weird.
But the models also need to grapple with the actual timeline throwing out things like “Mila Jovovich released AI memory system” and “US gov labeled Anthropic supply chain risk.”
In the wake of this paper I’m definitely wondering if the combination of overused negation over the last few years with increasingly wild and unpredictable reality is leading to negation as a heuristic to just not be very useful to even the most capable models.
“It’s not nothing.”
As an aside, trying SDF on both of these would be interesting. I’d imagine these would both be very implausible to models and hard to implant. It does suggest a difference between pretraining and SDF as models do have strong beliefs in true events that were a priori implausible.