Sorry, just to reiterate, the LLM generated few of the words but did help with some minor editing.
My AI safety essay was flagged for LLM content (I think because I had Claude generate the references and links), so I started putting everything I write in LLM content blocks to not deal with the headache.
I think this was an overreaction.
I wonder if the existence of evaluators at all will have a meaningful positive impact regardless of the competence of said evaluators.
https://www.anthropic.com/institute/measuring-pace-of-ai-development
One could read the above post as Anthropic trying to get ahead of possible problems revealed by evaluators, they sure spend a lot of paragraphs explaining why 6% of research compute to safety is more than it sounds.
“Safety research tends to use less compute than frontier training runs by its nature, so compute is an imperfect proxy for how much a company focuses on safety. This is because safety research consists of individual researchers designing experiments, which is time-consuming even though running the experiments is not particularly compute-intensive. The value of this metric, therefore, is less the absolute numbers and more that it provides a straightforward mechanism to compare like with like, across developers and over time.”
Not an expert but I could easily see the opposite spin being true: “safety researchers forced into time-consuming hand-crafted experiments because they’re bottlenecked by compute allocations.”