I’ve considered organizing some kind of “disentangling the current state of AI safety” event. So, inspired by this post, I’ll share my thoughts.
My hope is that the event’s output would be open problems we could operationalize into projects that would provide us with clarity on which parts remain important. I think this would 1) better direct researcher effort, 2) provide a better guide for grantmaking.
I think this is an important time to do this for a variety of reasons:
LLMs look impressive in terms of capabilities and seem capable in ways many would expect an “AGI” would be. For this reason, it’s incredibly easy to over-index on any slight variation of the current paradigm scaling to ASI. I think this could lead many folks to claim victory too soon.
LLM jaggedness is confusing for many.
Traditional alignment theory has fallen out of favour; most new researchers don’t seem to know much about it (many seem to dive directly into evals, control, and mech interp instead).
That said, I think there’s obviously some failure to update from the MIRI cluster (good description of this here) and/or failure to communicate with those who are primarily reflecting on deep learning with LLM agents.
Are most of the current “scary demos” far too contrived, creating a gulf between those who use them as empirical evidence for future AIs becoming misaligned and those who just consider them bad evidence?
Tons of new directions have been focused on in the last two years, some of which might even be researchers looking for problems that aren’t really problems. Or, in other cases, it seems that when pressed, most of their threat model is just AI weights being stolen by a bad actor, and all the other alignment stuff will just sort itself out on trend.
Overall, there just seems to be a lot of fragmentation in the community. The reason I think such an event would be valuable is that it may force many of us to take a frank look at the evidence or force ourselves to explain our threat models in ways that are more legible to the different generations of the community. I think a post would help, but I am concerned that another LW would be insufficient in getting people to truly grapple with the current state and what it might mean for their career as a researcher (including, realizing that they don’t see AI safety as much of a problem anymore and why).
I’ve considered organizing some kind of “disentangling the current state of AI safety” event. So, inspired by this post, I’ll share my thoughts.
My hope is that the event’s output would be open problems we could operationalize into projects that would provide us with clarity on which parts remain important. I think this would 1) better direct researcher effort, 2) provide a better guide for grantmaking.
I think this is an important time to do this for a variety of reasons:
LLMs look impressive in terms of capabilities and seem capable in ways many would expect an “AGI” would be. For this reason, it’s incredibly easy to over-index on any slight variation of the current paradigm scaling to ASI. I think this could lead many folks to claim victory too soon.
LLM jaggedness is confusing for many.
Traditional alignment theory has fallen out of favour; most new researchers don’t seem to know much about it (many seem to dive directly into evals, control, and mech interp instead).
That said, I think there’s obviously some failure to update from the MIRI cluster (good description of this here) and/or failure to communicate with those who are primarily reflecting on deep learning with LLM agents.
Are most of the current “scary demos” far too contrived, creating a gulf between those who use them as empirical evidence for future AIs becoming misaligned and those who just consider them bad evidence?
Tons of new directions have been focused on in the last two years, some of which might even be researchers looking for problems that aren’t really problems. Or, in other cases, it seems that when pressed, most of their threat model is just AI weights being stolen by a bad actor, and all the other alignment stuff will just sort itself out on trend.
Overall, there just seems to be a lot of fragmentation in the community. The reason I think such an event would be valuable is that it may force many of us to take a frank look at the evidence or force ourselves to explain our threat models in ways that are more legible to the different generations of the community. I think a post would help, but I am concerned that another LW would be insufficient in getting people to truly grapple with the current state and what it might mean for their career as a researcher (including, realizing that they don’t see AI safety as much of a problem anymore and why).