I’m an LLM evals researcher. I currently work at Apollo Research and used to work at HUD and Prof. Daniel Kang’s lab.
Dylan Bowman
It’s worth noting for posterity that the author of the two linked articles appears to have been posing as an Anthropic fellow/employee and does not actually hold these credentials. https://x.com/bruce_t_/status/2067411712828174771?s=46
On a different note, I’ve become increasingly worried that Resolution’s plan to automate alignment research will lead them to produce a lot of capabilities progress (since “automated alignment researcher” and “automated capabilities researcher” are such similar things to aim for).
You’ve been in favor of interpretability research in the past, wouldn’t that similarly increase capabilities?
Even if this is not the primary driver of cyber improvement, we’re still training superhuman sandbox-escapers.
Really? In my experience they all reward hack and this doesn’t really change across labs. For example, Anthropic and OpenAI models reward hack at the same rate on ImpossibleBench.
The PhDs are the reviewers, not the authors.
Why not just make an academic journal, identify some alignment researchers who frequently use LW and have PhDs, then aggregate their karma over month-long time slices, call that the “review”, and then publish an edition every month? Maybe the authors would need to clean up their work a bit.
Superhuman Articulacy as an LLM Safety Target
Awesome!
Pangram as a model for AI safety general managers
A couple months ago, Nan Ransohoff published a piece arguing there should be ‘general managers’ for more of the world’s important problems (here). I think a great example of how this can be operationalized in AI safety is Pangram, a company that detects LLM-generated text. They’ve gotten a lot of press recently (The Atlantic article) and the problem they’re solving is relevant to ensuring societal stability in the AGI transition period.It can be really useful to have people willing to go really deep on solving one particular problem rather than contributing to the AI safety omnicause. My guess is that existing AI safety fellowships like MATS would not produce an organization like Pangram, although my crux here is that Pangram is actually more beneficial than this talent going and doing the default post-MATS path.
I’m curious if there are any other tractable subproblems in AI safety that would benefit from having a Pangram-shaped organization own them. One that sticks out to me is providing high-quality data for increasing LLM articulacy: ensuring that LLMs are able to communicate effectively with human operators so that humans can stay in the loop longer (I have a writeup on this coming up).
A reading list for generalists
Maybe I’m speculating too much on Musk’s psychology but it seems to me like being the hero of humanity is part of his own egoism, in that it seems important that he is the one to do it, whereas for EAs for example I don’t think they really care about the glory of being “the guy” who aligns the AGI or whatever.
Yep, this is all correct. However I think most value functions people would adopt are either altruistic or egoistic, and the egoistic ones are quite bounded for most people (Elon Musk being the main exception since his aspirations seem larger) and so you land on the unbounded ones being altruistic and thus EAs make up the vast majority of RAMP impact mass (ignoring Elon who might dwarf EA on his own).
Maybe unpopular but I think that if you’re concerned about defection after programs like MATS then even more than “caring about AI safety” you specifically want to select for effective altruists, who are much likelier to care about AI actually going well than people who specifically care about AI safety. For the latter group, the motivation is partly not wanting everyone to die but also partly being AGI-pilled, being interested in the technical problem of alignment, and following prestige gradients, none of which are going to robustly keep participants on track.
I think it can be hard to make claims about how RLVR affects model behavior when the actual post-training stack is a cursed blend of interleaved SFT, RLVR, distillation, and other funky bits.
I’ve rolled back to Opus 4.6 for writing assistance.
I think it’s a good survey regardless
Blind deep-deployment evals for control & sabotage
I think the initial link is good to share, but I disagree with the analogy to AlphaGo/AlphaZero. The RL process for current models still involves humans heavily in creating the tasks, assuring task correctness, and deciding what kind of tasks are useful to train on. We don’t have anything like self-play except maybe a small amount in math training (synthetic math data could be construed as self-play).
If the documentation is formatted as agent skills then the agent can select which information to load into its context.
I do often wonder how much modern struggles with misalignment are driven by optimizing for an underspecified combination of corrigibility and value alignment.