Hi! I’m a CS PhD student at Carnegie Mellon under Vincent Conitzer. I work on AI Safety via Debate and multi-agent collaboration in mathematics. Always happy to chat!
Alexander Heckett
As a long-time advocate of multipolar worlds, I’ve started to feel that unipolar worlds are better in the long term. I’ve traditionally thought:
Multipolar worlds are better for control. For example, I think an instance of Fable would be more likely to flag misaligned swarm behavior from OpenAI models than an instance of that OpenAI swarm.
Unipolar worlds might avoid race dynamics between labs, but race dynamics don’t fundamentally bother me. I care less about timelines than I do about policies and research allocation / taste during takeoff.
This argument for multipolar worlds assumes that humans are in charge and are trying to corral wayward AI’s. However, if we extrapolate onwards to super-intelligence, I think the story flips:
Asking one super-intelligence to give Earth enough sunlight for us to live is a reasonable request. Yudkowsky might argue that a super-intelligence wouldn’t grant this request, but I’m hopeful.
Asking a swarm of individualistic, competing super-intelligences to give Earth enough sunlight for us to live might be an unstable equilibrium, in the same way that asking hundreds of countries to abide by climate regulation is unstable.
I wish we could live in a multipolar world during takeoff and then switch to a unipolar one afterwards.
Claim: political ideologies are low-temperature magnetization.
I think factor graphs are a reasonable toy model of how arguments are coupled to each other (where this coupling is either from the perspective of a person or an LLM). This framework is very analogous to an Ising model: statement true/false = spin up/down. If statements/atoms are weakly coupled / temperature is high, the truth/spin of a statement/atom is largely determined by its local neighborhood (i.e. it’s independent of everything far away). On the other hand, if statements/atoms are strongly coupled / temperature is low, the truth/spin of a statement/atom is largely determined by a globally-consistent-perspective-on-what’s-true/global-choice-of-magnetization.
I expect AI-Safety-via-Debate to be possible in the local regime but much sketchier in the global regime.
I wonder how much of the general trend (especially from Chinese companies) could be a result of other models distilling from Claude (so Anthropic trains on Lesswrong and other labs distill from Anthropic).
I’d be curious to hear the opinions on this post of Cooperative AI folks / people trying to get AIs to cooperate in prisoners-dilemma-type-situations.
I completely agree that, without a country of aligned geniuses in a datacenter, we’re dead. But I still think we should try to predict what dangers we’ll face. I’d classify murder-bacteria as a known threat; for example, Anthropic talks about bioweapon capabilities. What do unknown threats look like?
First, primitive defenses can help counter advanced offenses. Counter-drone warfare in Ukraine provides many examples: netting hung over roads, “turtle tanks” with metal sheds welded on top, inflatable decoys. These simple defenses don’t suffice, you also need interceptor drones and counter-drone teams, but they’re important. I expect miniaturized robots would be similar: you’d probably want to combine advanced defenses like bloodstream monitoring for bioweapons or ultrasonic echolocation of airborne threats with stupid defenses like thicker doors or regularly washed floors, and predicting the stupid defenses is something I reckon we can start now.
Second, even if there’s a threat for which no stupid defenses exist, I think it’s better to know about it. For example, I also worry about ant-sized robots building nuclear bombs underneath targets; I don’t see any stupid defenses against that. But we can still brace ourselves: we might recognize that cities will need to install underground metal detectors, maybe someone publishes UraniumDetectionDeviceBench, etcetera. I’d rather face known unknowns than unknown unknowns.
I want to see more work on physical security for everyday folks during a singularity. For example, imagine a tick-sized robot that could crawl up someone’s leg, burrow under their skin, and poison them or do other nasty damage. This feels physically possible and hard to defend against. If such robots become cheap, how would society cope? Would we rely on physical monitoring of peoples’ clothing? Would we try to determine who sent the robot after it does its damage? I think we can already start making guesses.
Synthetic Scalable Oversight
“The idea of “testing someone’s honesty and relying on the result” is not something any serious security practice depends on.”
“Notably, we generally don’t evaluate humans for this: We don’t test politicians for corruption before giving them office. We don’t test executives for recklessness before giving them control of companies.”
I think I mildly disagree? Certainly honesty-testing shouldn’t be the only line of defense in human-trust situations, but I think people often do rely on this. For example, many advisors don’t check their students’ work very closely because, after years of knowing the student, they trust the student to report their results honestly. Or for another example, a teenager in a small town might be more likely to be accepted for babysitting jobs if that teenager has a reputation for being level-headed and dependable.
These trust dynamics only work for relationships built on years of experience and repeated interaction, but I think they do work.
(2) and (4) feel very related to me. (1) feels like a grounding force that enables you to get (2) and (4) working in equilibrium / avoid equilibrium selection problems. (3) feels like the tricky bit that could use white-box techniques.
I’ve been working on AI math markets for a year and most of my research these days involves creating “synthetic” abstract graphs of how different theorems relate to each other and searching for good market mechanisms in these controlled environments (basically, there are a TON of possible market-like mechanisms and it’s not obvious a priori which ones work well). I’ve also been thinking about how to generalize these synthetic environments to debate and other less-grounded settings.
My mental model for non-math settings is roughly as follows: there’s some infinite probabilistic graphical model out there. Agents can spend effort trying to uncover new nodes and edges (e.g. think of arguments) and have the ability to publicly reveal nodes they uncover. They can also provide signals to each other (e.g. I like your evidence, I’m uncertain about this argument, hey your idea from 2 months ago was actually useful, etc.) and these signals feed into the combined reward mechanism.
If the reward mechanism can’t see the underlying probabilistic graphical model at all, you run into equilibrium selection problems very quickly, especially if you don’t regularize your agents towards a reasonable base policy (e.g. the agents are free to transform the PGM in any agreed-upon deterministic way and pretend the transformed PGM is the real one). You can fix those equilibrium selection problems by adding (1): if a trickle of nodes are empirically testable and the reward mechanism is allowed to use experimental results to inform agent rewards, that equilibrium degeneracy should break. (I haven’t tested that yet on synthetic data, this is only my hypothesis.)
So I think (1), (2), and (4) can all be understood synthetically. (3) is the tricky one, especially getting (3) and (4) to cooperate with each other. The mechanism design literature has things to say here but the setup would need to be clear first.
Debate in control settings seems more robust to collusion if you use different models, especially models from different providers. However, modern AI labs subsidize personal coding plans compared to API pricing, so it’s more expensive to run multi-lab debate than single-lab debate. This seems bad to me.
The question bank doesn’t exist yet because the language to ask the questions doesn’t exist yet. I spent a few weeks after writing this post trying to familiarize myself with Lean as quickly as possible and I found out that people in the Lean community simply haven’t formalized most of the objects I’d want to talk about (probabilistic networks, computational complexity, Nash equilibria, etc.). I tried to get a project off the ground formalizing these objects—you can see the GitHub repository here and the planning document here. Unfortunately this project quickly ballooned beyond what I can handle alone—I’m just an undergraduate student and winter break is over now. I still think it’s insane that some kind of crash formalization program isn’t currently underway. If you’re interested in pursuing a project like this then I’d be happy to talk through where I left off and what the next steps could look like!
I think the website just links back to this blog post? Is that intentional?Edit: I also think the application link requires access before seeing the form?Second Edit: Seems fixed now! Thanks!
Maybe I should also expand on what the “AI agents are submitting the programs themselves subject to your approval” scenario could look like. When I talk about a preorder on Turing Machines (or some subset of Turing Machines), you don’t have to specify this preorder up front. You just have to be able to evaluate it and the debaters have to be able to guess how it will evaluate.
If you already have a specific simulation program in mind then you can define as follows: if you’re handed two programs which are exact copies of your simulation software using different hard-coded world models then you consult your ordering on world models, if one submission is even a single character different from your intended program then it’s automatically less, if both programs differ from your program then you decide arbitrarily. What’s nice about the “ordering on computations” perspective is that it naturally generalizes to situations where you don’t follow this construction.
What could happen if we don’t supply our own simulation program via this construction? In the planes example, maybe the “snap” debater hands you a 50,000-line simulation program with a bug so that if you’re crafty with your grid sizes then it’ll get confused about the material properties and give the wrong answer. Then the “safe” debater might hand you a 200,000-line simulation program which avoids / patches the bug so that the crafty grid sizes now give the correct answer. Of course, there’s nothing stopping the “safe” debater from having half of those lines be comments containing a Lean proof using super annoying numerical PDE bounds or whatever to prove that the 200,000-line program avoids the same kind of bug as the 50,000-line program.
When you think about it that way, maybe it’s reasonable to give the “it’ll snap” debater a chance to respond to the “it’s safe” debater’s comments. Now maybe we change the type of from being a subset of (Turing Machines) x (Turing Machines) to being a subset of (Turing Machines) x (Turing Machines) x (Justifications from safe debater) x (Justifications from snap debater). In this manner deciding how you want to behave can become a computational problem in its own right.
These are all excellent points! I agree that these could be serious obstacles in practice. I do think that there are some counter-measures in practice, though.
I think the easiest to address is the occasional random failure, e.g. your “giving the wrong answer on exact powers of 1000” example. I would probably try to address this issue by looking at stochastic models of computation, e.g. probabilistic Turing machines. You’d need to accommodate stochastic simulations anyway because so many techniques in practice use sampling. I think you can handle stochastic evaluation maps in a similar fashion to the deterministic case but everything gets a bit more annoying (e.g. you likely need to include an incentive to point to simple world models if you want the game to have any equilibria at all). Anyways, if you’ve figured out topological debate in the stochastic case, then you can reduce from the occasional-errors problem to the stochastic problem as follows: suppose is a directed set of world models and is some simulation software. Define a stochastic program which takes in a world model , randomly samples a world model according to some reasonably-spread-out distribution, and return . In the 1D plane case, for example, you could take in a given resolution, divide it by a uniformly random real number in , and then run the simulation at that new resolution. If your errors are sufficiently rare then your stochastic topological debate setup should handle things from here.
Somewhat more serious is the case where “it’s harder to disrupt patterns injected during bids.” Mathematically I interpret this statement as the existence of a world model which evaluates to the wrong answer such that you have to take a vastly more computationally intensive refinement to get the correct answer. I think it’s reasonable to detect when this problem is occurring but preventing it seems hard: you’d basically need to create a better simulation program which doesn’t suffer from the same issue. For some problems that could be a tall order without assistance but if your AI agents are submitting the programs themselves subject to your approval then maybe it’s surmountable.
What I find the most serious and most interesting, though, is the case where your simulation software simply might not converge to the truth. To expand on your nonlinear effects example: suppose our resolution map can specify dimensions of individual grid cells. Suppose that our simulation software has a glitch where, if you alternate the sizes of grid cells along some direction, the simulation gets tricked into thinking the material has a different stiffness or something. This is a kind of glitch which both sides can exploit and the net probably won’t converge to anything.
I find this problem interesting because it attacks one of the core vulnerabilities that I think debate problems struggle with: grounding in reality. You can’t really design a system to “return the correct answer” without somehow specifying what makes an answer correct. I tried to ground topological debate in this pre-existing ordering on computations that gets handed to us and which is taken to be a canonical characterization of the problem we want to solve. In practice, though, that’s really just kicking the can down the road: any user would have to come up with a simulation program or method of comparing simulation programs which encapsulates their question of interest. That’s not an easy task.
Still, I don’t think we need to give up so easily. Maybe we don’t ground ourselves by assuming that the user has a simulation program but instead ground ourselves by assuming that the user can check whether a simulation program or comparison between simulation programs is valid. For example, suppose we’re in the alternating-grid-cell-sizes example. Intuitively the correct debater should be able to isolate an extreme example and go to the human and say “hey, this behavior is ridiculous, your software is clearly broken here!” I will think about what a mathematical model of this setup might look like. Of course, we just kicked the can down the road, but I think that there should be some perturbation of these ideas which is practical and robust.
Thank you! If I may ask, what kind of fatal flaws do you expect for real-world simulations? Underspecified / ill-defined questions, buggy simulation software, multiple simulation programs giving irreconcilably conflicting answers in practice, etc.? I ask because I think that in some situations it’s reasonable to imagine the AI debaters providing the simulation software themselves if they can formally verify its accuracy, but that would struggle against e.g. underspecified questions.
Also, is there some prototypical example of a “tough real world question” you have in mind? I will gladly concede that not all questions naturally fit into this framework. I was primarily inspired by physical security questions like biological attacks or backdoors in mechanical hardware.
Topological Debate Framework
My initial reaction is that at least some of these points would be covered by the Guaranteed Safe AI agenda if that works out, right? Though the “AGIs act much like a colonizing civilization” situation does scare me because it’s the kind of thing which locally looks harmless but collectively is highly dangerous. It would require no misalignment on the part of any individual AI.
Totalitarian dictatorship
I’m unclear why this risk is specific to multipolar scenarios? Even if you have a single AGI/ASI you could end up with a totalitarian dictatorship, no? In fact I would imagine that having multiple AGI/ASI’s would mitigate this risk as, optimistically, every domestic actor in possession of an AGI/ASI should be counterbalanced by another domestic actor with divergent interests also in possession of an AGI/ASI.
Wasn’t there a MATS project this past summer doing exactly this for 4B − 8B parameter models? I believe William was working on this?