Here are some notes on checks and balances and what it might look like to preserve them into the AI era.
First, why even have checks and balances? What are they for?
A few possible answers:
To prevent any individual actor getting too much power
This is the power concentration version
To prevent unchecked power
I think this is strictly more accurate than the first answer (as in, it more accurately tracks what I care about), but it currently feels a bit abstract to me
To preserve/enhance something like kindness, integrity, alignment in the overall system
Here I’m trying to rhyme with Richard Ngo’s well-foundedness and Jan Kulveit’s kind Leviathan
Maybe there’s some negatively framed dual to this which is like ‘no exploitation’ (unsure)
Of those, the one I like most is preserving kindness, integrity, alignment in the overall system.
So, if I take preserving kindness as the purpose of checks and balances, where does that take me when I start to think about how to preserve checks and balances into the AI era?
Here’s a first pass:
Checks and balances need to do two things:
Preserve horizontal kindness between agents at the same level (for example, kindness between me and another human, or between nation states)
Preserve vertical kindness between agents at different levels (for example, kindness between a government and its citizens, or an ASI and a human)
Preserving horizontal kindness requires two things:
Rules: some specification of what isn’t ok
Enforcement mechanisms: some way of preventing actors from breaking the rules
Some things I notice about this
Rules can be explicit, like the law, or implicit, like one’s moral compass
Rules can be enforced top down, like the state policing criminals, or enforced through coordination, like a village ostracising someone who beats their wife
Preserving vertical kindness requires two things:
Good information transfer between levels, so that they understand one another
Some alignment mechanism
One way of doing this is making sure that they are mutually dependent/each have hard leverage over the other
Rules and enforcement is another way of doing this
[in some ways good information transfer is an alignment mechanism]
[presumably other options]
A little bit more thinking, this time in a table:
Function
What this looks like today
What this could look like in the AI era
Preserving horizontal kindness
Rules
Law
Norms
Law (updated)
Whatever rules AI systems are trained and can be verified to follow, and the rules that govern how those rules can change
[I expect norms will matter less]
Enforcement mechanisms
Locally, social fabric
Domestically, the police
Internationally and only partially, the military
Self-enforcing mechanisms (like the rules AI systems are trained and verified to follow)
ASI nightwatchman
I guess MAIM is maybe something like this?
Structured transparency style mass surveillance and enforcer drones who stop people building WMDs
Preserving vertical kindness
Information transfer
Managers and reports talk to each other. Also reports sometimes fill in feedback forms
Citizens vote, and sometimes write to or speak with their representatives
There are polls
…
Structured transparency could enable much higher bandwidth in both directions
AI delegates
More ambient mediation between humans and their environment by AI systems (e.g. there’s a secure system which takes in data from all of the systems I use and knows everything about me, and this data can only be used for certain purposes/not in raw form, but gives really accurate and deep information)
Really good simulation?
Alignment
Aligning government to the people:
Constitutions and other law
Elections
Citizen taxation
Aligning the people to government:
Education
The law
…
Giving people capital and/or compute
Self-enforcing rules that restrict what the higher vertical levels can do
Things that stand out to me from the table:
Novel threats
Accelerating progress could lead competitive dynamics to erode all value. Progress hasn’t been fast enough for this before
Enforcement could potentially be perfect. This would be bad in our current world I think; unclear if it will be good or bad in a future one
Tech could make n of one violations extremely costly, which makes perfect enforcement desirable
But perfect enforcement also significantly raises the bar for how good the rules need to be, and the constitutional process which sets and amends them
Novel opportunities
Sufficiently powerful information transfer and enforcement mechanisms to prevent war, which hasn’t been possible before
Verifiably self-enforcing rules is a v powerful new affordance
Here are some notes on checks and balances and what it might look like to preserve them into the AI era.
First, why even have checks and balances? What are they for?
A few possible answers:
To prevent any individual actor getting too much power
This is the power concentration version
To prevent unchecked power
I think this is strictly more accurate than the first answer (as in, it more accurately tracks what I care about), but it currently feels a bit abstract to me
To preserve/enhance something like kindness, integrity, alignment in the overall system
Here I’m trying to rhyme with Richard Ngo’s well-foundedness and Jan Kulveit’s kind Leviathan
Maybe there’s some negatively framed dual to this which is like ‘no exploitation’ (unsure)
Of those, the one I like most is preserving kindness, integrity, alignment in the overall system.
So, if I take preserving kindness as the purpose of checks and balances, where does that take me when I start to think about how to preserve checks and balances into the AI era?
Here’s a first pass:
Checks and balances need to do two things:
Preserve horizontal kindness between agents at the same level (for example, kindness between me and another human, or between nation states)
Preserve vertical kindness between agents at different levels (for example, kindness between a government and its citizens, or an ASI and a human)
Preserving horizontal kindness requires two things:
Rules: some specification of what isn’t ok
Enforcement mechanisms: some way of preventing actors from breaking the rules
Some things I notice about this
Rules can be explicit, like the law, or implicit, like one’s moral compass
Rules can be enforced top down, like the state policing criminals, or enforced through coordination, like a village ostracising someone who beats their wife
Preserving vertical kindness requires two things:
Good information transfer between levels, so that they understand one another
Some alignment mechanism
One way of doing this is making sure that they are mutually dependent/each have hard leverage over the other
Rules and enforcement is another way of doing this
[in some ways good information transfer is an alignment mechanism]
[presumably other options]
A little bit more thinking, this time in a table:
Function
What this looks like today
What this could look like in the AI era
Preserving horizontal kindness
Rules
Law
Norms
Law (updated)
Whatever rules AI systems are trained and can be verified to follow, and the rules that govern how those rules can change
[I expect norms will matter less]
Enforcement mechanisms
Locally, social fabric
Domestically, the police
Internationally and only partially, the military
Self-enforcing mechanisms (like the rules AI systems are trained and verified to follow)
ASI nightwatchman
I guess MAIM is maybe something like this?
Structured transparency style mass surveillance and enforcer drones who stop people building WMDs
Preserving vertical kindness
Information transfer
Managers and reports talk to each other. Also reports sometimes fill in feedback forms
Citizens vote, and sometimes write to or speak with their representatives
There are polls
…
Structured transparency could enable much higher bandwidth in both directions
AI delegates
More ambient mediation between humans and their environment by AI systems (e.g. there’s a secure system which takes in data from all of the systems I use and knows everything about me, and this data can only be used for certain purposes/not in raw form, but gives really accurate and deep information)
Really good simulation?
Alignment
Aligning government to the people:
Constitutions and other law
Elections
Citizen taxation
Aligning the people to government:
Education
The law
…
Giving people capital and/or compute
Self-enforcing rules that restrict what the higher vertical levels can do
Things that stand out to me from the table:
Novel threats
Accelerating progress could lead competitive dynamics to erode all value. Progress hasn’t been fast enough for this before
Enforcement could potentially be perfect. This would be bad in our current world I think; unclear if it will be good or bad in a future one
Tech could make n of one violations extremely costly, which makes perfect enforcement desirable
But perfect enforcement also significantly raises the bar for how good the rules need to be, and the constitutional process which sets and amends them
Novel opportunities
Sufficiently powerful information transfer and enforcement mechanisms to prevent war, which hasn’t been possible before
Verifiably self-enforcing rules is a v powerful new affordance