Response to Leverage’s research report on “argument mapping”
Day 5 of forced writing with an accountability partner!
Leverage wrote a report on “argument mapping” in the early 2010s and published the findings in 2020. I am very interested in ”argument mapping”[1] for tough analytical problems like AI policy, and multiple people have directed me to this report when I bring up the topic. I think this report raises some important points but its findings are probably flawed—or at the very least, people reading the report probably derive an overly-pessimistic view of “argument mapping” as a whole, especially given that the evaluation metrics are strange.[2]
Rather than focus on where I agree with the report, in this shortform I will just briefly outline some of the qualms I have with this report. I do not consider these rebuttals definitive—I recognize that there may be more to the research than I can see—but I could not easily determine if/how the report responds to some of these criticisms (which has notable irony to it). Some of these objections include:
The report emphasizes forming consensus among participants, with little attention given to the impact on audiences/3rd-parties (two terms that never even show up in the document?[3]). Notably, this focus may fail to capture most of the value of “argument mapping,” in at least two ways:
Sometimes the participants have already staked their reputation on certain views or are otherwise biased to not change their mind, whereas a policymaker/company/grant-writer or other decision-making principal might still be open-minded but uncertain. Thus, while the participants may not be swayed by convincing evidence, if you can make it significantly easier for a neutral principal to answer questions like “did X party ever respond to Q objection?” that may improve their decision-making, which is valuable regardless of whether you’ve achieved consensus.
Building on the previous point about making it easier for audiences/principals to understand what’s going on, audience costs may be the most powerful way of incentivizing “consensus” (or just “good epistemic behavior”) in some cases: if you look like a stubborn or dishonest researcher to an audience, you might suffer even more reputational damage than if you just admit you were wrong. No amount of staring-you-in-the-face experimental evidence will necessarily convince Ye Olde Epistemic Guard to admit that the current way of building ships is inferior. But if it’s sufficiently obvious to merchants then they may stop relying on YOEG and start funding your work instead. Importantly for this research report, it wasn’t clear that the report really emphasized audience costs, given the insular nature of the research project, which undermines the report’s ability to evaluate the effect of argument mapping on consensus formation.
The report fails to acknowledge the existence of Kialo, which I consider to be one of the most effective and successful “argument mapping” platforms (and which currently still exists). This might normally be fine, but in December 2020, the report adds an addendum stating that their assessment of “argument mapping” was demonstrated to be true, and basically that nothing new was successful. They provide an appendix with a long list of relevant software, but Kialo isn’t there. This certainly isn’t damning—and I’ll certainly admit that Kialo still has some issues—but the lack of any mention did leave me wondering whether Leverage had a good process for finding and evaluating these projects, among other things. (Notably, I once got the sense that Kialo doesn’t actively call itself “argument mapping,” which might explain the problem, but it is in reality well within the broad umbrella of “argument mapping.”)
The report had strangely high bars for evaluating success (”very large gains (10x-100x) for groups seeking to reach consensus”). At the very least, it seems quite possible for someone to read their conclusion as being more damning than it really is. (In my view, even a net 10% increase in “consensus formation” or just “research and analysis productivity” would be enormously valuable when applied to important questions within AI technical safety or policy.)
Simply put, I believe that most of the methods for “argument mapping” that Leverage used were poor choices, especially when they emphasized formal logic. Among other things, this led them to claim that making good argument maps requires high-skilled contributors, which I do not think is a very accurate assessment (or at least, it can be quite misleading). However, I will leave further discussion of this point to a future shortform/post on why I think many forms/methods of “argument mapping” are fundamentally misguided—especially when they try to do deductive arguments
I think that some of the topics they chose to test these maps on were very poor choices (e.g., “Whether the world needs saving”). Question framing is really important. (But again, I’ll leave this to a future shortform/post.)
This term is painfully broad and, as Leverage demonstrates, often is used to refer to methods which I would not endorse, such as when they try create deductive arguments or otherwise heavily use formal logic. However, in lieu of a better term at the moment, I will continue referring to argument mapping in scare quotes.
Thus, it might be possible to claim that the report was accurate in its findings, but that the problem simply comes from misinterpretation. I think that the scope itself was problematic and undesirable, but in this shortform I will reserve deeper judgments on the matter.
I couldn’t quickly verify whether the report used alternative terms to get at this idea, but I don’t recall seeing this on previous occasions when I half-skimmed-half-read the report...
Response to Leverage’s research report on “argument mapping”
Day 5 of forced writing with an accountability partner!
Leverage wrote a report on “argument mapping” in the early 2010s and published the findings in 2020. I am very interested in ”argument mapping”[1] for tough analytical problems like AI policy, and multiple people have directed me to this report when I bring up the topic. I think this report raises some important points but its findings are probably flawed—or at the very least, people reading the report probably derive an overly-pessimistic view of “argument mapping” as a whole, especially given that the evaluation metrics are strange.[2]
Rather than focus on where I agree with the report, in this shortform I will just briefly outline some of the qualms I have with this report. I do not consider these rebuttals definitive—I recognize that there may be more to the research than I can see—but I could not easily determine if/how the report responds to some of these criticisms (which has notable irony to it). Some of these objections include:
The report emphasizes forming consensus among participants, with little attention given to the impact on audiences/3rd-parties (two terms that never even show up in the document?[3]). Notably, this focus may fail to capture most of the value of “argument mapping,” in at least two ways:
Sometimes the participants have already staked their reputation on certain views or are otherwise biased to not change their mind, whereas a policymaker/company/grant-writer or other decision-making principal might still be open-minded but uncertain. Thus, while the participants may not be swayed by convincing evidence, if you can make it significantly easier for a neutral principal to answer questions like “did X party ever respond to Q objection?” that may improve their decision-making, which is valuable regardless of whether you’ve achieved consensus.
Building on the previous point about making it easier for audiences/principals to understand what’s going on, audience costs may be the most powerful way of incentivizing “consensus” (or just “good epistemic behavior”) in some cases: if you look like a stubborn or dishonest researcher to an audience, you might suffer even more reputational damage than if you just admit you were wrong. No amount of staring-you-in-the-face experimental evidence will necessarily convince Ye Olde Epistemic Guard to admit that the current way of building ships is inferior. But if it’s sufficiently obvious to merchants then they may stop relying on YOEG and start funding your work instead. Importantly for this research report, it wasn’t clear that the report really emphasized audience costs, given the insular nature of the research project, which undermines the report’s ability to evaluate the effect of argument mapping on consensus formation.
The report fails to acknowledge the existence of Kialo, which I consider to be one of the most effective and successful “argument mapping” platforms (and which currently still exists). This might normally be fine, but in December 2020, the report adds an addendum stating that their assessment of “argument mapping” was demonstrated to be true, and basically that nothing new was successful. They provide an appendix with a long list of relevant software, but Kialo isn’t there. This certainly isn’t damning—and I’ll certainly admit that Kialo still has some issues—but the lack of any mention did leave me wondering whether Leverage had a good process for finding and evaluating these projects, among other things. (Notably, I once got the sense that Kialo doesn’t actively call itself “argument mapping,” which might explain the problem, but it is in reality well within the broad umbrella of “argument mapping.”)
The report had strangely high bars for evaluating success (”very large gains (10x-100x) for groups seeking to reach consensus”). At the very least, it seems quite possible for someone to read their conclusion as being more damning than it really is. (In my view, even a net 10% increase in “consensus formation” or just “research and analysis productivity” would be enormously valuable when applied to important questions within AI technical safety or policy.)
Simply put, I believe that most of the methods for “argument mapping” that Leverage used were poor choices, especially when they emphasized formal logic. Among other things, this led them to claim that making good argument maps requires high-skilled contributors, which I do not think is a very accurate assessment (or at least, it can be quite misleading). However, I will leave further discussion of this point to a future shortform/post on why I think many forms/methods of “argument mapping” are fundamentally misguided—especially when they try to do deductive arguments
I think that some of the topics they chose to test these maps on were very poor choices (e.g., “Whether the world needs saving”). Question framing is really important. (But again, I’ll leave this to a future shortform/post.)
This term is painfully broad and, as Leverage demonstrates, often is used to refer to methods which I would not endorse, such as when they try create deductive arguments or otherwise heavily use formal logic. However, in lieu of a better term at the moment, I will continue referring to argument mapping in scare quotes.
Thus, it might be possible to claim that the report was accurate in its findings, but that the problem simply comes from misinterpretation. I think that the scope itself was problematic and undesirable, but in this shortform I will reserve deeper judgments on the matter.
I couldn’t quickly verify whether the report used alternative terms to get at this idea, but I don’t recall seeing this on previous occasions when I half-skimmed-half-read the report...