Yong
One potential concern is that we might risk producing slops—and in some way we have to review what they write and that might still place burden on the system.
I went through Astra and I think that I benefitted the most from the weekly meetings with my mentors. If I had just gotten research credits, I‘d most likely just publish work that doesn’t matter. Even if I got constructive feedback at the end, not sure if those time sunk is worth it.
Perhaps this (giving credits) might be best structured as microgrant like what Leo put out.
IF they cannot set the refusal classifier well for cyber and bio, what gives you confidence that they would classify “frontier LLM development” well? Not only that, you would need to second-guess the response you get. I’d rather they just do overrefusals for frontier LLM development questions (since they clearly don’t care about overrefusals).
With regard to citations, it’s one thing to complain that a paper is only citing two out of three of the relevant pieces of prior work—but it’s another thing to complain that a paper seems blissfully unaware of an entire relevant body of prior work. This is especially problematic if the prior work persuasively establishes some limitations on or reasons to be skeptical of the author’s preferred data or methodology.
There are also a lot of papers that just cite without properly engage––some even mischaractererize––with the work. I notice that this problem is worsened with a lot of less-experienced researchers simply relying on LLMs for writing, especially on related work or discussion section.
Many times I went through some of the cited work, and found that they almost have identical findings or even entirely opposite findings. I usually make it a point in my own writings to write something like “The closest work to ours is by XXX, who found …”.
where would you see these AI safety forward deployed engineers mostly work at?
from your description, I believe they have the highest value in model labs as we are talking about AIS wrt to catastrophic risks, but there aren’t really that many labs — and I don’t think Chinese labs would hire these engineers from US/UK/EU.
I see that a lot of startups are hiring for security research engineers (eg Mirendil that focuses on AI R&D that accelerates science) which I believe that’s having a higher demand at the moment but my understanding is that theres a higher focus on alignment researchers rather than security researchers within LW community (at least based on my read on programs like MATS)