Re: structure of claims in papers: the way I had set up my benchmark is to focus on getting models to follow up experiments that are more thing sort of “have to be run” in order to test a core claim (although I did more follow-up-ish stuff as well)
Super curious about this! Wonder how you define it for alignment and such. Looking forward to seeing this when it gets out!
For memorization, the setup I had was asking the model to recall as many details about a paper as possible when given only the title and author list. I think that’s probably a better way to measure it?
Perhaps. I’ll run this, but we seem to agree that memorization is not playing into results as-presented in this case.
Super curious about this! Wonder how you define it for alignment and such. Looking forward to seeing this when it gets out!
Perhaps. I’ll run this, but we seem to agree that memorization is not playing into results as-presented in this case.