Voluntary commitments and even regulation could be too hard to enforce across the board—such that responsible actors end up adhering to if-then commitments, while irresponsible actors rush forward with dangerous AI. One of the challenges with AI is that complete enforcement of any given risk mitigation framework seems extremely hard to achieve, yet incomplete enforcement could end up disadvantaging responsible actors in a high-stakes, global technology race. This is a general issue with most ways of reducing AI risks, other than “race forward and hope that the benefits outweigh the costs,” and is not specific to if-then commitments.
To help mitigate this issue, early, voluntary if-then commitments can contain “escape clauses” along the lines of: “We may cease adhering to these commitments if some actor who is not adhering to them is close to building more capable models than ours.” (Some more detailed suggested language for such a commitment is provided by METR, a nonprofit that works on AI evaluations.)
To help mitigate this issue, early, voluntary if-then commitments can contain “escape clauses” along the lines of: “We may cease adhering to these commitments if some actor who is not adhering to them is close to building more capable models than ours.” (Some more detailed suggested language for such a commitment is provided by METR, a nonprofit that works on AI evaluations.)
Just for reference, this framing is what makes me feel fine about things you said on this topic, but not fine about conversations I’ve had with Anthropic employees about this topic in the last few years. My conversations with Anthropic employees did definitely not involve them saying “we are committing to our RSP only if every other company also adopts a similar RSP”.
At the most they were saying “we are going to revise our RSP as we learn more about what an effective RSP would look like and might make changes in-accordance with that”, which is of course drastically different. If the commitment all along had been to “commit to the RSP conditional on other people also committing to equivalent policies”, then the RSP could have said that directly, and the change from an unconditional to a conditional policy is of course massive (and I think the RSP as written clearly was communicating itself as an unconditional policy).
Note this quote later in that same piece:
Voluntary commitments and even regulation could be too hard to enforce across the board—such that responsible actors end up adhering to if-then commitments, while irresponsible actors rush forward with dangerous AI. One of the challenges with AI is that complete enforcement of any given risk mitigation framework seems extremely hard to achieve, yet incomplete enforcement could end up disadvantaging responsible actors in a high-stakes, global technology race. This is a general issue with most ways of reducing AI risks, other than “race forward and hope that the benefits outweigh the costs,” and is not specific to if-then commitments.
To help mitigate this issue, early, voluntary if-then commitments can contain “escape clauses” along the lines of: “We may cease adhering to these commitments if some actor who is not adhering to them is close to building more capable models than ours.” (Some more detailed suggested language for such a commitment is provided by METR, a nonprofit that works on AI evaluations.)
Just for reference, this framing is what makes me feel fine about things you said on this topic, but not fine about conversations I’ve had with Anthropic employees about this topic in the last few years. My conversations with Anthropic employees did definitely not involve them saying “we are committing to our RSP only if every other company also adopts a similar RSP”.
At the most they were saying “we are going to revise our RSP as we learn more about what an effective RSP would look like and might make changes in-accordance with that”, which is of course drastically different. If the commitment all along had been to “commit to the RSP conditional on other people also committing to equivalent policies”, then the RSP could have said that directly, and the change from an unconditional to a conditional policy is of course massive (and I think the RSP as written clearly was communicating itself as an unconditional policy).