For context, the post is decently transparent about the quality of the research presented. A quick exploration done in a few days and mostly AI made should be expected to have a number of issues.
This is a tangent I took for a few days during the Pivotal AI Safety Research Fellowship. I’m doing an AI Control project mentored by Adam Kaufman and James Lucassen of Redwood Research and partnered with Aniruddh Pramod.[4]
This blog post is a human-written summary/analysis of a pile of information requested by Matthew and discovered by Claude.
I think there’s a difference between an error (e.g., some llm judge you used having a really poor recall rate or something) and taking poor Claude-conducted analysis as given (e.g., 97% of AI safety research use Openrouter unsafely, which is reported above the epistemic note. Ditto the “all the providers have gap” section.)
[I’ve also expanded on my thinking about this more in the original comment’s edit.]
For context, the post is decently transparent about the quality of the research presented. A quick exploration done in a few days and mostly AI made should be expected to have a number of issues.
I think there’s a difference between an error (e.g., some llm judge you used having a really poor recall rate or something) and taking poor Claude-conducted analysis as given (e.g., 97% of AI safety research use Openrouter unsafely, which is reported above the epistemic note. Ditto the “all the providers have gap” section.)
[I’ve also expanded on my thinking about this more in the original comment’s edit.]