Thanks for writing this agenda. A lot of it reads like overly focused on current model failures. Often the shelf live of tools and other frameworks is fairly short (e.g., https://forum.effectivealtruism.org/posts/uTffD74jJrQvHJuBY/coordinal-a-postmortem ). However, this doesn’t mean that I think it’s not worth trying to build good frameworks. It’s just that I am very bitter-lesson pilled and hence have see this as fairly risky.
I don’t follow your rebuttal to agents may acquire taste:
Even if agents have all the taste, tacit knowledge and values of a theoretical human researcher, we still will likely want to keep humans in the loop.
Why would this require special tools in this case? If the agents have taste and are faithful, I would expect there not to be a need for special frameworks, interfaces and infra.
Different researchers approach problems in different ways. And even the best researcher, when operating autonomously over long periods of time, may deviate from the goals you set for the research
This can be the whole point of research, deviating from the goal by finding something that matters more. I don’t see why this would be an issue if agents have good epistemics.
A lot of it reads like overly focused on current model failures. Often the shelf live of tools and other frameworks is fairly short (e.g., https://forum.effectivealtruism.org/posts/uTffD74jJrQvHJuBY/coordinal-a-postmortem ). However, this doesn’t mean that I think it’s not worth trying to build good frameworks. It’s just that I am very bitter-lesson pilled and hence have see this as fairly risky.
That’s why I’ve explicitly framed this as a research agenda, not an engineering agenda. The goal here isn’t to build the best cockpit—it’s to discover what are the principles of good cockpit design.
I don’t follow your rebuttal to agents may acquire taste:
Even if agents have all the taste, tacit knowledge and values of a theoretical human researcher, we still will likely want to keep humans in the loop.
Why would this require special tools in this case? If the agents have taste and are faithful, I would expect there not to be a need for special frameworks, interfaces and infra.
I want to re-emphasize the focus on research here, not tools. Assuming in such a case we want to keep humans in the loop, any agent building frameworks, interfaces and infra to keep humans in the loop needs to know what works and what doesn’t.
Imagine asking an agent to create a control protocol without any AI control research having been done. How likely do you think it is that that agent is going to one-shot the best protocol?
You can think of one output of this agenda as a research-backed design guide for agents, so they know how to build effective frameworks, interfaces and infra that keep humans in the loop.
Different researchers approach problems in different ways. And even the best researcher, when operating autonomously over long periods of time, may deviate from the goals you set for the research
This can be the whole point of research, deviating from the goal by finding something that matters more. I don’t see why this would be an issue if agents have good epistemics.
What do you define as “good epistemics”? Do you mean an agent reasons well?
It’s also important to note that “good epistemics” doesn’t mean “good outcome”. Both humans and agents can reason well, but come to the wrong conclusions. Or the reasoning is probabilistic and the resolution of those probabilities results in going down wrong path (and there may not be enough data for an agent to recognize that).
That said, I do see that my argument is weaker if we assume an agent is as fully capable and aligned as a human. I just don’t think we’ll get there anytime soon, and I don’t think we have the tools right now to determine when we’ve reached there. If we have a way to prove that models are aligned enough to hand off all R&D to them, then, at least in my book, we’ve solved alignment.
Thanks for writing this agenda. A lot of it reads like overly focused on current model failures. Often the shelf live of tools and other frameworks is fairly short (e.g., https://forum.effectivealtruism.org/posts/uTffD74jJrQvHJuBY/coordinal-a-postmortem ). However, this doesn’t mean that I think it’s not worth trying to build good frameworks. It’s just that I am very bitter-lesson pilled and hence have see this as fairly risky.
I don’t follow your rebuttal to agents may acquire taste:
Why would this require special tools in this case? If the agents have taste and are faithful, I would expect there not to be a need for special frameworks, interfaces and infra.
This can be the whole point of research, deviating from the goal by finding something that matters more. I don’t see why this would be an issue if agents have good epistemics.
That’s why I’ve explicitly framed this as a research agenda, not an engineering agenda. The goal here isn’t to build the best cockpit—it’s to discover what are the principles of good cockpit design.
I want to re-emphasize the focus on research here, not tools. Assuming in such a case we want to keep humans in the loop, any agent building frameworks, interfaces and infra to keep humans in the loop needs to know what works and what doesn’t.
Imagine asking an agent to create a control protocol without any AI control research having been done. How likely do you think it is that that agent is going to one-shot the best protocol?
You can think of one output of this agenda as a research-backed design guide for agents, so they know how to build effective frameworks, interfaces and infra that keep humans in the loop.
What do you define as “good epistemics”? Do you mean an agent reasons well?
It’s also important to note that “good epistemics” doesn’t mean “good outcome”. Both humans and agents can reason well, but come to the wrong conclusions. Or the reasoning is probabilistic and the resolution of those probabilities results in going down wrong path (and there may not be enough data for an agent to recognize that).
That said, I do see that my argument is weaker if we assume an agent is as fully capable and aligned as a human. I just don’t think we’ll get there anytime soon, and I don’t think we have the tools right now to determine when we’ve reached there. If we have a way to prove that models are aligned enough to hand off all R&D to them, then, at least in my book, we’ve solved alignment.