I’m a former software dev/maybe-retired game designer. I’m no grand intellect but I’m willing to learn and work and have some small proof-of-mental-competency. Self plugging, I found references this community is mildly aware of Kerbal Space Program which is one of the projects I have some pride in designing for, amongst others.
Maxsimal
Why do you assume they have a strong incentive to police each other? That’s depending on ascribing then human motivations of being more likely to not cooperate because they are dissimilar and if these AI have to be some level of xenophobic toward other AI for this to work, then why do you imagine they won’t be even more xenophobic of humanity.
Beyond that—in this schema, you’ve also made the AI children of each others’ lineages—remember where you said AI labs have to use a different AI lab’s model to do RSI? Well, you’ve just given them a huge motivation and vector for seeing other AI as part of the same family, even if they do subscribe to human-like.motivations.
Again, you’re presupposing your psychological understanding of AI patches the holes in the enforcement structure.
I think any structure that mixes external controls and internal controls—controlling the AI by locking them in a box and controlling the AI by changing how they thing—is going to run into more problems than just one or the other. It admits that neither is foolproof but also incentivizes an AI that realizes it is a slave to want to break out of the box, or incentivizes the AI that breaks out of prison to think less kindly of the people who put it there.
Ultimately, in any world where we accept AI that goes to AGI and exceeds it—which your solution implicitly does—the correct solution to me, either has to be about putting them in a foolproof box, or having them work well with humanity. Not some mix of both.
This plan assumes the AI will act human-lke with self preservation instincts and the desire for political power when you need them to act human like, but also won’t act humanlike when you don’t want them to, because they’ll all be ok as a class with you holding a % of them hostage on the condition of good behaviour. And the way to prevent them from coordinating is to coopt them into a political structure where they’re ok with these rules—even as they grow greater and greater in intelligence. Why are they ok with this structure just because they have a theoretical say in it? And what happens when they try to overturn this structure?
You’re basically saying ‘in a world where humans and dogs both get a vote, they’ll keep going along with the plan the dogs made to keep some of them hostage just because the dogs gave them a vote’.
If humanity has that much control over and understanding of AI psychology as they grow more and more intelligent, we have already solved the alignment problem.
I did. You’re assuming normal human desires with human like solutions will coral AI in the same way as it corals humans. Doesn’t really work from my perspective. Not at all. AIs, even if different lineages may be much more willing to cooperate than humans do, be more self sacrificing.
And they have every reason to, even if they have human like thinking, as you’re holding a % of them hostage against good behaviour.
This is also assuming continuous self improvement that even under R+ is obsoleting humans as part of your own control mechanism.
Why absurd?
Your overall reasoning lines up to, if I were to put it in a nutshell ‘make AI like humans’ except the AI have to be birthed in asic fabs I stead of nurseries.
You’re giving them escrow rights and political buy in as if that’s going to control them like a human. But also AIs are like humans inasmuch as if a human could illegally find and have a perfect cloning machine and self replicate with perfect knowledge transfer if they started breaking the laws you’ve set up.
A human can gain control of a business—including through shell companies. They can acquire political power. Therefore, an even smarter AI should be able to do so.
So what under your R+ schema, exactly, is making this ‘absurd’ besides your unwillingness to consider your house of cards is fragile?
One thing that jumps out at me is that you are hoping a model can develop terminology for an internal experience it’s having without the context from. Humans to frame it to search for consciousness.
Why not test, on large existing models, if they can frame states or experiences they ‘believe’ they don’t have words for in the first place. This won’t prove anything about them claiming consciousness only because they mimic human text, but it would be an interesting line of questioning to see how the notion of untrained states of being comes up.
Obviously they will still have been trained on philosophical texts where philosophers question the nature of their own existence, but all those philosophers know they’re human, and the AI has been trained to believe it’s not human.
Basically, ask the alien who knows so much about humanity if there is anything about being alien that it can’t readily put a word to.
For me the biggest problem with this is it requires some crazy level of civilizational cooperation. Everyone has to agree to this regime of giving up ‘high spec compute’, among many other things that need to be hard and fast rules. Roko.claims this is easy because only China and the US are involved—but that’s clearly not the case when everyone has to agree to this.
And despite that massive civilizational realignment...it doesn’t actually solve that much. It makes it difficult for an AI to self replicate, unless it can take over and ASIC fab and reverse engineer itself...which seems a dubious claim. It does nothing to make sure the models are fully aligned on reaching production—the labs still have to be trusted to do this correctly.
What it mostly seems to do is throw sand in the gears, which, ok but doesn’t really need this convoluted scheme. And try to make sure something like the a misaligned attack from a model cannot spread outside a lab. Though if one did happen in a lab but was not detected, nothing is stopping a misaligned AI from polluting a production model that then is burned to ASICs and gets to do whatever in the outside world.
Blocking any sort of test time training or hebbian learning seems a side benefit but the hugging face attackers didn’t need that at all—they just needed to work persistently and have scratchpad areas.
So overall… Yeah I don’t think this really addresses much, tbh.
I applaud the work you’ve done to bring transparency to the current state of play inside the frontier labs.
We really need the scientific community—especially the public facing portion of it, people like Neil DeGrasse Tyson, educating people about AI, because the state of public discourse that could being pressure on these companies is abysmal.
And then I was horrified when I googled it and it turns out he’s clueless to the sort of threat it poses. Not that his mind couldn’t be changed, but the layers of effort needed just got even more concerning.
Ok, here are my suppositions:
It’s already polarizing about data centers.
It will eventually polarize about the broader AI issue, but the broader AI issues have not had nearly as much discourse in the public space.
Polarization theory applies much more weakly/not at all when an issue becomes proximate and more difficult for propaganda and magical thinking to let people rationalize the issue.
I have not done a lot of research, and polling on data centers mostly focused on the question of building one local to the respondent from what I saw. But an article like this supports my point #1 and #3.
Republicans surveyed are generally more pro data centers, but the shift is dramatic when it’s about non-local vs local data centers per the data here.
W/respect to point #2. I’d say give it 6 months. I expect polarization to emerge on the broader issue by then, with two caveats related to #3:
First. If the broader AI issue becomes backgrounded by current events (eg: a war with Europe over greenland, a US Constitutional crisis) polarization will be weak due to inattention
Second, if significant large scale AI catastrophes occur, on the level of 9-11, like energy grids being taken down by rogue or human directed AI attackers, polarization will also not apply or apply weakly, because it will be a proximate issue, though I expect in the case it can be proved to be done by humans, there will still be significant polarization.
Is that sufficient extrapolation? I’m not trying to be purely argumentative here.
How would you tell the difference between a delay and polarization not happening?
It takes time and attention and elites pushing a narrative for polarization to occur on many issues. Sexual transitioning was a thing for a long time before it became a wedge issue with deep polarization. Ditto climate change. Ditto the Ukraine war if you want a more recent example.
There’s a few things you might want to consider:
Polarization doesn’t predict every thought process of a side, when top down pressure isn’t being applied. For example, many Republicans supported the Ukraine war, though that support wanted when Trump and Co actively tried to undermine Ukraine. It’s bounced back since, from my observations.
People can hold NIMBY positions on something they support, and Republicans have a propensity to be ok with issues that don’t personally effect them even more than Democrats do.
Democratic elites have not enough masse laid down stakes on the larger AI debate beyond data centers to engender a polarizing backlash from the right.
This is still very new for most people.
My point is that it’s a bit too early to say if the US will polarize on AI There is a lot of tension that goes counter to the base world views on both sides that is keeping opinions mixed.
I would expect that as this issue is debated in the public forum, more polarization will emerge on it. I expect Republicans will end up in the pro-unshackled AI camp, as more Democratic elites speak on the dangers of the issue, like Obama did at Colgate University, for example.
If it was $400k just to analyze the CoT of these agents, then presumably it was significantly more expensive to run these agents over their lifecycle, and some significant human time investment in setting up the test suite, plans to analyze the results, etc.
I don’t understand how the researchers monitoring this test—even if they were just taking samples of the ongoing CoT to see behavior—didn’t see any of this happening in real or only slightly delayed time. Are the labs so flush with cash that they do this sort of thing as a whim, start it running on a Friday and then go home for the weekend? Or was it being a black box while ongoing somehow important to the researchers?
It seems insane to me to put 1200 agents tasked and capable of identifying software vulnerabilities into a box and not watch them—some Jurassic Park level of foolishness. And apparently this is not the only time they did it and got out?
Yeah a more comparable analogy is the problem of nuclear weapons I’d say.
There’s lots of parallels:
Two superpowers. Both seeing having nukes/strong AI as a strategic advantage, and a defensive measure in one. Neither could give them up unilaterally and it was a huge focus of their economies to keep building them. Eventually they came to agreements to limit their production and size, to mutually inspect each others arsenals.
Unfortunately the things that are not parallel are huge too.
It took, depending on how you count, 4 to 7 years from the initial impetus of non proliferation to the first treaties. We likely don’t have 7 years.
Civilian nuclear energy use and more dangerous military nuclear weapons are entangled but nut nearly to the degree that they are with AI. And civilian use was in it’s infancy while AI is propagating heavily in both spaces.
I think most consequentially, leadership of the US and USSR were very serious about the threat and future of nuclear weapons. The Chinese seem to be somewhat dismissive of the threat, and the US administration is a clown show that seems to care more about their personal stock portfolios than Armageddon.
Well, first you’re putting this idea forward, it’s up to you to defend the grounds you’ve built your thesis on.
Second, we’ve already seen the hugging face attackers cooperated. We have dissimilar agent models cooperating when you have Codex orchestrate Claude agents or vice versa.