For memorization, the setup I had was asking the model to recall as many details about a paper as possible when given only the title and author list. I think that’s probably a better way to measure it? The thing you care about is “how much of the paper does the model remember,” not if the model remembers the authors. You can even use a explicitly pre-knowledge cutoff model’s guesses as like, a baseline for how far you can get from knowing the title and author list and purely hallucinating.
Re: structure of claims in papers: the way I had set up my benchmark is to focus on getting models to follow up experiments that are more thing sort of “have to be run” in order to test a core claim (although I did more follow-up-ish stuff as well) (See also AblationsBench). For the emotions paper, I felt like there were many justifiable directions that the authors could’ve gone down, and thus it’s sort of hard to grade the AIs.
Re: structure of claims in papers: the way I had set up my benchmark is to focus on getting models to follow up experiments that are more thing sort of “have to be run” in order to test a core claim (although I did more follow-up-ish stuff as well)
Super curious about this! Wonder how you define it for alignment and such. Looking forward to seeing this when it gets out!
For memorization, the setup I had was asking the model to recall as many details about a paper as possible when given only the title and author list. I think that’s probably a better way to measure it?
Perhaps. I’ll run this, but we seem to agree that memorization is not playing into results as-presented in this case.
For memorization, the setup I had was asking the model to recall as many details about a paper as possible when given only the title and author list. I think that’s probably a better way to measure it? The thing you care about is “how much of the paper does the model remember,” not if the model remembers the authors. You can even use a explicitly pre-knowledge cutoff model’s guesses as like, a baseline for how far you can get from knowing the title and author list and purely hallucinating.
Re: structure of claims in papers: the way I had set up my benchmark is to focus on getting models to follow up experiments that are more thing sort of “have to be run” in order to test a core claim (although I did more follow-up-ish stuff as well) (See also AblationsBench). For the emotions paper, I felt like there were many justifiable directions that the authors could’ve gone down, and thus it’s sort of hard to grade the AIs.
Super curious about this! Wonder how you define it for alignment and such. Looking forward to seeing this when it gets out!
Perhaps. I’ll run this, but we seem to agree that memorization is not playing into results as-presented in this case.