Responding here as the author of Formalizing Objections against Surrogate Goals, and somebody who thought about SPIs a lot, but always ended up prioritising other research because I didn’t find good enough ways to make progress on this:
I think a valuable thing to put on the agenda would be something like: Give a list of example scenarios where you expect that SPIs should be able to help (eg, if we already had most of this agenda figured out).
Hypothetical and future scenarios are allowed, as are scenarios that aren’t about AI. But the examples should be as concrete as possible.
The hope behind this is that it would give us intuition pumps with which progress on SPIs would get easier and faster.
Also, it would allow us to check whether we need to still be concerned about my complaints from the linked post above. And it would make it easier to notice if we were using SPIs in situations where they work but some other approach would work even better.
I agree that thinking about concrete scenarios is important. But I’m not exactly sure what you have in mind here: “The hope behind this is that it would give us intuition pumps with which progress on SPIs would get easier and faster.” What’s a quick example?
Let’s say that my agenda is “I am studying the Principal-Agent problem and contracts, because I want to address the issues caused by misaligned incentives between humans”.
And suppose I come up with an idea for how to set up contracts better—I prove a theorem which says that if a contract looks such and so, the Nash equilibrium is Pareto-improving over whatever was the baseline. (Or I make some other contribution to the general topic.)
But how do I know whether this contribution actually helps? Should I, like, try to implement something based on it, see if it catches on, and measure how much it helps? As stated, I don’t have any easy-to-use feedback loop for checking any of this. I might derive the perfect solution on paper, only to later learn that the key problem was somewhere else.
But suppose I instead had the following list of things that my agenda is meant to help with[1]: (1) I sometimes hire cleaners, but they never do the hard-to-reach parts of my apartment. (2) My sister sometimes borrows my car, but she never fills up the gas afterwards. (3) I want a higher salary at my academic job. (4) I have a company, and I want to hire a contractor to would handle marketing on my behalf.
Then this list would give me a way to get quick & intuitive feedback on whether my proposed solution will work. For example, I would immediately see that my contract idea is an overkill for (1-2), it won’t work for (3) because the university contracts are set in stone and I don’t have leverage to change them, and maybe it might work for (4).
(And I am suggesting we need something like this for SPIs.)
Responding here as the author of Formalizing Objections against Surrogate Goals, and somebody who thought about SPIs a lot, but always ended up prioritising other research because I didn’t find good enough ways to make progress on this:
I think a valuable thing to put on the agenda would be something like: Give a list of example scenarios where you expect that SPIs should be able to help (eg, if we already had most of this agenda figured out).
Hypothetical and future scenarios are allowed, as are scenarios that aren’t about AI. But the examples should be as concrete as possible.
The hope behind this is that it would give us intuition pumps with which progress on SPIs would get easier and faster.
Also, it would allow us to check whether we need to still be concerned about my complaints from the linked post above. And it would make it easier to notice if we were using SPIs in situations where they work but some other approach would work even better.
Thanks Vojta!
I agree that thinking about concrete scenarios is important. But I’m not exactly sure what you have in mind here: “The hope behind this is that it would give us intuition pumps with which progress on SPIs would get easier and faster.” What’s a quick example?
Hm. I will try an exaggerated non-SPI example:
Let’s say that my agenda is “I am studying the Principal-Agent problem and contracts, because I want to address the issues caused by misaligned incentives between humans”.
And suppose I come up with an idea for how to set up contracts better—I prove a theorem which says that if a contract looks such and so, the Nash equilibrium is Pareto-improving over whatever was the baseline. (Or I make some other contribution to the general topic.)
But how do I know whether this contribution actually helps? Should I, like, try to implement something based on it, see if it catches on, and measure how much it helps? As stated, I don’t have any easy-to-use feedback loop for checking any of this. I might derive the perfect solution on paper, only to later learn that the key problem was somewhere else.
But suppose I instead had the following list of things that my agenda is meant to help with [1] : (1) I sometimes hire cleaners, but they never do the hard-to-reach parts of my apartment. (2) My sister sometimes borrows my car, but she never fills up the gas afterwards. (3) I want a higher salary at my academic job. (4) I have a company, and I want to hire a contractor to would handle marketing on my behalf.
Then this list would give me a way to get quick & intuitive feedback on whether my proposed solution will work. For example, I would immediately see that my contract idea is an overkill for (1-2), it won’t work for (3) because the university contracts are set in stone and I don’t have leverage to change them, and maybe it might work for (4).
(And I am suggesting we need something like this for SPIs.)
For the record, all of these are made up.