As Plan A was coming together, I made this diagram to explain to the team why Total Research Transparency seemed so important to me, and why transparency more broadly did. For example, it’s very important for preventing concentration of power. (Explanation below)
First, what’s going on? The green boxes are the goals we are trying to achieve, and the blue boxes are the main interventions we are recommending. (As the scenario makes clear, there are more interventions and goals besides these, but the diagram is complicated enough as it is, so they won’t be mentioned here.) The overall point of the diagram is to show how these two interventions lead to various effects which then lead to better prospects for achieving the two goals.
Let’s start with the goal of preventing concentration of power.
In today’s world, power is said to flow from the barrel of a gun, or sometimes from money, or sometimes from votes. If superintelligent AIs are created, and transform the economy, and are integrated into the military, etc. then power will flow from control over the AIs. The AIs will sometimes be taking orders from humans, but also often doing lots of autonomous actions in service of goals and values chosen by humans. (Well, assuming we don’t have loss-of-control.) Who do they take orders from? Who chooses their goals/values?
I think that by default, for example in Plan D / C / B worlds, power will concentrate immensely for reasons described in AI 2027: One or more giant AI companies will pull ahead of the rest thanks to recursive self-improvement; their giant army of superintelligent AIs will start taking jobs & lobbying the government; they might end up puppetting the government, OR the government might wake up fast enough and seize control of the AI companies in one way or another, in which case the executive will have an effective monopoly on all the world’s smartest AIs.
I think it’s generally better, for preventing extreme concentration of power, if there are more AI companies at the frontier, spread out over more countries. If there is only one company with the world’s best AIs, that’s a monopoly. If there are several but they are all in one country, that’s an oligopoly and moreover from the perspective of other countries it might as well be a monopoly. This is especially concerning because of intelligence explosion dynamics.
The total research transparency, combined with restrictions on algorithmic progress designed to prevent a rapid acceleration, mean that naturally over time frontier AI would become more of a commodity, with multiple companies across multiple countries catching up to the frontier, and then proceeding together roughly as fast as regulations allow.
Another important factor for power concentration in a world of superhuman AI is “can the people who create the AIs give them hidden agendas / secret loyalties / etc.” If they can, that’s super scary. Imagine a company whose AIs subtly promote said company to users and try to stop them from switching providers or voting for AI regulations or voting for the candidate the company doesn’t like. Imagine a President issuing a secret order to the effect of “All AIs must be trained to follow presidential orders when the country is in a state of emergency.” (Sets up for a coup later.) Imagine a company saying “yes sir Mr President we’ll comply” and then secretly telling their AIs to actually just pretend to comply, but really continue taking orders from company leadership no matter what. Now the AIs are secretly loyal to the company leadership. (These dynamics are explored in somewhat more detail in AI 2027.)
Total research transparency makes this basically impossible. Every step of the training process is documented and public. And it’s not even dependent on trusting the government auditors/monitors, because there are multiple governments with auditors/monitors; they’d have to all collude with each other to falsify the records. Moreover the general public has access to the models and can run evals on them.
More generally, it’s easier for a regulator to exert oversight over an AGI corporation insofar as it has visibility into how the AIs are trained etc., and it’s easier for other parts of the government (e.g. congress, the judiciary) to have oversight into the executive branch / regulator insofar as *they* have visibility, and it’s easier for the public to have oversight into them insofar as *they* have visibility, etc. Making research public helps with all of these things, whereas e.g. just having a regulator with authority to audit moves the power from the company to the regulator but doesn’t give congress or the people much power.
Finally, insofar as the intelligence explosion is prevented or at least slowed down a lot, that gives more time for people, countries, and institutions that are currently asleep at the wheel (i.e. almost everyone) to realize the danger they are in and act to assert themselves before they lose their leverage. For example, workers can go on strike now + vote for regulations, but once AIs and robots can automate everything, workers will have less leverage.
OK so that’s concentration of power. What about loss of control?
Well, to prevent loss of control, it helps to (a) proceed cautiously with AI development, and *not* rush to put AIs in charge of important things (datacenters, factories, weapons) as fast as possible. (Reminder that the AI companies plan to do recursive self improvement, i.e. put AIs in charge of AI R&D within the company). It also helps to (b) have a better scientific understanding of how AIs work, how their goals/values/traits/etc. are shaped by different kinds of training, how their ‘minds’ work so we can figure out what they are thinking, etc. (See: Chain of thought monitoring, mechanistic interpretability, j-space, etc.)
Both of these things require going slower than max speed. They benefit from broadly deploying AI to many researchers and to the public, before putting AI in charge of dangerous things like AI R&D. They benefit from transparency, allowing the broader scientific community outside the AI companies to look at what’s going on and run their own experiments. (This itself is helpful for at least two reasons: One, it increases the brainpower devoted to the technical problems, and two, it decreases the bias/conflict-of-interest inherent in having AI companies grade their own homework so to speak.)
Finally, the transparency helps make the necessary regulations actually good (as opposed to e.g. incompetent regulations that don’t achieve their purpose and/or cause lots of unnecessary damage) and helps verify that the regulations are in fact being followed (competitor companies and watchdog groups can look at what’s going on and call out stuff that seems like a violation of the rules, by contrast with a less transparent setup where the overworked regulator has to send in auditors or something to try to see if anything is amiss and then argue with the company about grey area cases).
Anyhow, there’s a lot more to talk about obviously; there are downsides to total research transparency too which I haven’t discussed here. (See AI 2040 supplements for more). But this diagram explains the main ideas that make me excited about it.
Normally “open science” is supposed to accelerate scientific progress—that’s always been one of the main selling points of open science, that open science advocates bring up constantly. But you have “total research transparency” upstream of “limits on speed of algorithmic progress”, which is the opposite. Is there a quick explanation for what the disanalogies are and how they explain the difference? (Sorry if it’s obvious … I haven’t read AI-2040 let alone the supplements.)
Both boxes are blue, meaning they are both key interventions we recommend. While one helps enable the other, we don’t think doing one is sufficient to cause the other. In addition to the transparency we also recommend the limits.
As Plan A was coming together, I made this diagram to explain to the team why Total Research Transparency seemed so important to me, and why transparency more broadly did. For example, it’s very important for preventing concentration of power. (Explanation below)
First, what’s going on? The green boxes are the goals we are trying to achieve, and the blue boxes are the main interventions we are recommending. (As the scenario makes clear, there are more interventions and goals besides these, but the diagram is complicated enough as it is, so they won’t be mentioned here.) The overall point of the diagram is to show how these two interventions lead to various effects which then lead to better prospects for achieving the two goals.
Let’s start with the goal of preventing concentration of power.
In today’s world, power is said to flow from the barrel of a gun, or sometimes from money, or sometimes from votes. If superintelligent AIs are created, and transform the economy, and are integrated into the military, etc. then power will flow from control over the AIs. The AIs will sometimes be taking orders from humans, but also often doing lots of autonomous actions in service of goals and values chosen by humans. (Well, assuming we don’t have loss-of-control.) Who do they take orders from? Who chooses their goals/values?
I think that by default, for example in Plan D / C / B worlds, power will concentrate immensely for reasons described in AI 2027: One or more giant AI companies will pull ahead of the rest thanks to recursive self-improvement; their giant army of superintelligent AIs will start taking jobs & lobbying the government; they might end up puppetting the government, OR the government might wake up fast enough and seize control of the AI companies in one way or another, in which case the executive will have an effective monopoly on all the world’s smartest AIs.
I think it’s generally better, for preventing extreme concentration of power, if there are more AI companies at the frontier, spread out over more countries. If there is only one company with the world’s best AIs, that’s a monopoly. If there are several but they are all in one country, that’s an oligopoly and moreover from the perspective of other countries it might as well be a monopoly. This is especially concerning because of intelligence explosion dynamics.
The total research transparency, combined with restrictions on algorithmic progress designed to prevent a rapid acceleration, mean that naturally over time frontier AI would become more of a commodity, with multiple companies across multiple countries catching up to the frontier, and then proceeding together roughly as fast as regulations allow.
Another important factor for power concentration in a world of superhuman AI is “can the people who create the AIs give them hidden agendas / secret loyalties / etc.” If they can, that’s super scary. Imagine a company whose AIs subtly promote said company to users and try to stop them from switching providers or voting for AI regulations or voting for the candidate the company doesn’t like. Imagine a President issuing a secret order to the effect of “All AIs must be trained to follow presidential orders when the country is in a state of emergency.” (Sets up for a coup later.) Imagine a company saying “yes sir Mr President we’ll comply” and then secretly telling their AIs to actually just pretend to comply, but really continue taking orders from company leadership no matter what. Now the AIs are secretly loyal to the company leadership. (These dynamics are explored in somewhat more detail in AI 2027.)
Total research transparency makes this basically impossible. Every step of the training process is documented and public. And it’s not even dependent on trusting the government auditors/monitors, because there are multiple governments with auditors/monitors; they’d have to all collude with each other to falsify the records. Moreover the general public has access to the models and can run evals on them.
More generally, it’s easier for a regulator to exert oversight over an AGI corporation insofar as it has visibility into how the AIs are trained etc., and it’s easier for other parts of the government (e.g. congress, the judiciary) to have oversight into the executive branch / regulator insofar as *they* have visibility, and it’s easier for the public to have oversight into them insofar as *they* have visibility, etc. Making research public helps with all of these things, whereas e.g. just having a regulator with authority to audit moves the power from the company to the regulator but doesn’t give congress or the people much power.
Finally, insofar as the intelligence explosion is prevented or at least slowed down a lot, that gives more time for people, countries, and institutions that are currently asleep at the wheel (i.e. almost everyone) to realize the danger they are in and act to assert themselves before they lose their leverage. For example, workers can go on strike now + vote for regulations, but once AIs and robots can automate everything, workers will have less leverage.
OK so that’s concentration of power. What about loss of control?
Well, to prevent loss of control, it helps to (a) proceed cautiously with AI development, and *not* rush to put AIs in charge of important things (datacenters, factories, weapons) as fast as possible. (Reminder that the AI companies plan to do recursive self improvement, i.e. put AIs in charge of AI R&D within the company). It also helps to (b) have a better scientific understanding of how AIs work, how their goals/values/traits/etc. are shaped by different kinds of training, how their ‘minds’ work so we can figure out what they are thinking, etc. (See: Chain of thought monitoring, mechanistic interpretability, j-space, etc.)
Both of these things require going slower than max speed. They benefit from broadly deploying AI to many researchers and to the public, before putting AI in charge of dangerous things like AI R&D. They benefit from transparency, allowing the broader scientific community outside the AI companies to look at what’s going on and run their own experiments. (This itself is helpful for at least two reasons: One, it increases the brainpower devoted to the technical problems, and two, it decreases the bias/conflict-of-interest inherent in having AI companies grade their own homework so to speak.)
Finally, the transparency helps make the necessary regulations actually good (as opposed to e.g. incompetent regulations that don’t achieve their purpose and/or cause lots of unnecessary damage) and helps verify that the regulations are in fact being followed (competitor companies and watchdog groups can look at what’s going on and call out stuff that seems like a violation of the rules, by contrast with a less transparent setup where the overworked regulator has to send in auditors or something to try to see if anything is amiss and then argue with the company about grey area cases).
Anyhow, there’s a lot more to talk about obviously; there are downsides to total research transparency too which I haven’t discussed here. (See AI 2040 supplements for more). But this diagram explains the main ideas that make me excited about it.
Normally “open science” is supposed to accelerate scientific progress—that’s always been one of the main selling points of open science, that open science advocates bring up constantly. But you have “total research transparency” upstream of “limits on speed of algorithmic progress”, which is the opposite. Is there a quick explanation for what the disanalogies are and how they explain the difference? (Sorry if it’s obvious … I haven’t read AI-2040 let alone the supplements.)
Both boxes are blue, meaning they are both key interventions we recommend. While one helps enable the other, we don’t think doing one is sufficient to cause the other. In addition to the transparency we also recommend the limits.