Write for the specific Claude building the specific thing.
I used to think that for publications to influence the labs, an employee would need to read and implement the ideas.
But now, I think that the one reading and implementing the work is Claude, and Claude can read a lot more stuff. If Claude is asked to build a new RL environment, he could read every relevant paper before starting the project, whereas a human engineer might only read one or two, if any.
This means that if you are writing and publishing research, if it’s net positive work at all, it’s way more likely to be read and considered at implementation time than it was in the past (when it might have been ignored completely, or just vaguely remembered).
This updates me:
Good but not great research is more useful than I previously thought.
Useful ideas from outside the labs are far more likely to make it into their work than I previously thought.
Four upvotes on the Alignment Forum could be enough to have an idea be useful.
When I ask Claude to help me with AI safety research, I usually already have a pretty good idea of what I want it to implement. It doesn’t usually look up papers unless I ask it to, and I probably wouldn’t want it to bring in ideas from random papers it just read that I don’t understand, since they might mess up my results.
If I were having Claude implement RL environments to train a frontier model, it may be somewhat different: it could make sense to bring in a whole kitchen sink of ideas for improving the model’s alignment. But I’d still be pretty wary about it. I’d be okay with Claude suggesting a new technique, but I’d at least want to read a few paragraphs explaining what it is and how it works. I’d also be much more likely to consider a technique if it had gained popularity among humans: in the hypothetical world where e.g. inoculation prompting never got much attention, I wouldn’t want Claude to proactively add it.
In the future, AI developers will probably trust their AIs to independently research random ideas and get more evidence about whether they should be implemented widely. The AIs can comb through many 4-upvote LW posts and comments, invest effort in the ones they think are promising, and only present the most important results to the humans. See my shortform about scraping LessWrong for AI safety ideas. Once humans agree that certain ideas are good, they can be implemented systematically across a whole training run.
So in conclusion, I agree with you in spirit, but I think that ideas will usually influence decisions made by an organization or the whole AI safety community, rather than influencing individual Claudes spun up by individual users.
Aye, I think we’re broadly in agreement, and I don’t claim that this is a perfect system or guaranteed to get useful ideas implemented in labs, just that it’s easier than I previously thought.
That said, I think an important idea here is that it’s likely to happen more as AIs do more work (which you gesture at too). To riff a bit, if we map increasing autonomy onto your LessWrong scraping idea, I could imagine it going from “Find and read the research and show me your top picks” to “Find and read the research, replicate your top picks and assess them for our specific project, and then show me your results from that” to “On a weekly cron job, find and read any research relevant to our projects, run them through the assessment pipeline, analyse the results yourself, and implement anything you think is good automatically”.
The last stage is a big jump, but will happen at some point as we approach RSI. I expect it to happen soonest in areas that are verifiable, especially where there is a big cost to not making a change quickly (e.g. searching for methods to make sandboxes more secure and allowing Claude to automatically patch them), or where there’s little cost to getting it wrong (e.g. automated research agents taking ideas and implementing them in experimental small scale runs in different combinations / conditions / scale to the original work).
That last one admittedly still goes through the human bottleneck, but at a later stage, and I think that touches on your point about ideas influencing the organization / community level. If a piece of work was good but not great (e.g. had some interesting ideas, but wasn’t perfectly implemented, or was badly presented, or not done in realistic conditions), it is now more likely to be elevated by a research agent in a lab investigating the same thing.
My original post spoke about an instance of Claude trying to implement a specific project, finding the research, and implementing it in its final run, which is somewhat different to the research agent example or the LessWrong scraping example, but I think they all point in the same direction of outside, publicly available work having more value because of the ability of AIs to quickly read, assess, replicate, and implement it at a much larger scale than traditional human engineering.
Write for the specific Claude building the specific thing.
I used to think that for publications to influence the labs, an employee would need to read and implement the ideas.
But now, I think that the one reading and implementing the work is Claude, and Claude can read a lot more stuff. If Claude is asked to build a new RL environment, he could read every relevant paper before starting the project, whereas a human engineer might only read one or two, if any.
This means that if you are writing and publishing research, if it’s net positive work at all, it’s way more likely to be read and considered at implementation time than it was in the past (when it might have been ignored completely, or just vaguely remembered).
This updates me:
Good but not great research is more useful than I previously thought.
Useful ideas from outside the labs are far more likely to make it into their work than I previously thought.
Four upvotes on the Alignment Forum could be enough to have an idea be useful.
When I ask Claude to help me with AI safety research, I usually already have a pretty good idea of what I want it to implement. It doesn’t usually look up papers unless I ask it to, and I probably wouldn’t want it to bring in ideas from random papers it just read that I don’t understand, since they might mess up my results.
If I were having Claude implement RL environments to train a frontier model, it may be somewhat different: it could make sense to bring in a whole kitchen sink of ideas for improving the model’s alignment. But I’d still be pretty wary about it. I’d be okay with Claude suggesting a new technique, but I’d at least want to read a few paragraphs explaining what it is and how it works. I’d also be much more likely to consider a technique if it had gained popularity among humans: in the hypothetical world where e.g. inoculation prompting never got much attention, I wouldn’t want Claude to proactively add it.
In the future, AI developers will probably trust their AIs to independently research random ideas and get more evidence about whether they should be implemented widely. The AIs can comb through many 4-upvote LW posts and comments, invest effort in the ones they think are promising, and only present the most important results to the humans. See my shortform about scraping LessWrong for AI safety ideas. Once humans agree that certain ideas are good, they can be implemented systematically across a whole training run.
So in conclusion, I agree with you in spirit, but I think that ideas will usually influence decisions made by an organization or the whole AI safety community, rather than influencing individual Claudes spun up by individual users.
Aye, I think we’re broadly in agreement, and I don’t claim that this is a perfect system or guaranteed to get useful ideas implemented in labs, just that it’s easier than I previously thought.
That said, I think an important idea here is that it’s likely to happen more as AIs do more work (which you gesture at too). To riff a bit, if we map increasing autonomy onto your LessWrong scraping idea, I could imagine it going from “Find and read the research and show me your top picks” to “Find and read the research, replicate your top picks and assess them for our specific project, and then show me your results from that” to “On a weekly cron job, find and read any research relevant to our projects, run them through the assessment pipeline, analyse the results yourself, and implement anything you think is good automatically”.
The last stage is a big jump, but will happen at some point as we approach RSI. I expect it to happen soonest in areas that are verifiable, especially where there is a big cost to not making a change quickly (e.g. searching for methods to make sandboxes more secure and allowing Claude to automatically patch them), or where there’s little cost to getting it wrong (e.g. automated research agents taking ideas and implementing them in experimental small scale runs in different combinations / conditions / scale to the original work).
That last one admittedly still goes through the human bottleneck, but at a later stage, and I think that touches on your point about ideas influencing the organization / community level. If a piece of work was good but not great (e.g. had some interesting ideas, but wasn’t perfectly implemented, or was badly presented, or not done in realistic conditions), it is now more likely to be elevated by a research agent in a lab investigating the same thing.
My original post spoke about an instance of Claude trying to implement a specific project, finding the research, and implementing it in its final run, which is somewhat different to the research agent example or the LessWrong scraping example, but I think they all point in the same direction of outside, publicly available work having more value because of the ability of AIs to quickly read, assess, replicate, and implement it at a much larger scale than traditional human engineering.