I do independent AI safety research. If you want to mentor me, make me your fellow, or simply be friends, DM me or email xjohanbonilla@gmail.com.
Johan David Bonilla
The Name is Not The Model
“We may be unable to build good AI” is true no matter what anyone does. Pausing doesn’t make it more buildable. Working toward it is the only thing that might.
I didn’t say restrictions don’t reduce harm. I said you can’t stop a motivated bad actor. The UK lowering average gun deaths doesn’t contradict that, a determined person can still get a gun.
I take the loss-of-control fear seriously. But I don’t think we should pause AI development, mainly because we can’t afford to let bad actors take advantage of it.
A pause binds mainly the people who honor it. It slows the careful and frees the reckless. It’s a bit like how countries don’t want to spend billions on defense every year, but they can’t afford not to. They can’t stop building better weapons, because the other side won’t stop either.
And even without a pause, there are already people out there building harmful AI, and we can’t stop them, just like we can’t stop a motivated bad actor from building a gun or malware. We also can’t completely stop AI from causing harm on its own, through recursive self improvement or emergent behavior nobody asked for. Not any one model in particular, but AI as a whole.
So one of our few real chances is widespread good AI simply outnumbering bad. More good models, more capable safe labs, and more standards and regulations that push the whole field toward models that behave safely by default.
I do get the catch, though. Increased model capability arms both sides, and outnumbering only wins if defense beats offense. One bad AI could cause catastrophic, asymmetric harm that a thousand good models may not be able to stop. It kind of reminds me of nukes, in the sense that a single global superpower could technically turn the lights off for all of us.
I don’t have a clean answer to the asymmetric case, or any answers for that matter. But I still don’t think we should pause.
Johan David Bonilla’s Shortform
I don’t think people appreciate how genuinely hard the position you and the other assessors are in.
Imagine having a sibling who tells you everything (you’re the only one who actually knows them), but the day they do something seriously wrong you can’t actually tell your parents because it ruins the whole relationship.
You can choose not to tell so you keep the visibility, but then you’re kind of complicit. So either way you lose something, access or integrity.
One real world example I’m reminded of is Enron. The outside firm hired to audit their books (Arthur Andersen) stayed quiet because Enron was too big a client to lose. The fraud ran for years and eventually took both of them out. That’s the same thing we’re seeing here: the side being checked is valuable enough that the checker can’t afford to upset them.
Two things feel impossible to get around:
1. You might not be able to reply with full honesty, and that’s the whole point.
You’re the public face of an org that runs on these relationships. Even if you totally agreed, saying so out loud costs you.
That’s not an AI thing either, it’s just what being an important public figure is like. People in your position rarely get to say the full truth. Meanwhile I’m an outsider, free to say anything, but no access and a weightless voice.
2. Regulation is the obvious historical fix, but governments are not clean either.
They’re their own institutions with their own incentives and their own version of this exact problem. It seems like it’s incentive-misalignment at every level, not just the developer/assessor one.
One of the more concerning things is that regulation usually shows up after disaster forces it. Most industries survive the disaster and learn from it, but you could imagine a reality where the AI disaster that forces the rule has a cost impossible to recover from.
Thanks @Buck for your notes on third-party risk assessment. You made the different parts that get lumped together legible. In particular fact-generation vs evidence analysis, and the two types of information laundering (whether it’s the company’s secrets being hidden or the assessor’s) are now stuck with me.
As someone trying to find my break into AI safety, I really appreciate this rare and honest picture of what I’m walking into before I do (hopefully).
I think about model behavior through my own existence a lot, and honestly there is always a fair bit of larping. Even now I feel like I’m roleplaying being a good son, brother, friend, researcher. I wear a different mask for each of them, and even for myself.
You get a promotion at work, and are you suddenly a new person or just pretending to be one? After a while the line between acting and reality just blurs. I don’t think the roleplaying matters as much as the actions themselves.
Have you tested what happens when you give a model a role and then just let it run with it?
Say a malicious actor sets Gemini loose in an open-ended long horizon loop, unsupervised, tells it to roleplay an evil misaligned schemer, and lets it keep generating its own roles on top of that one. Or even simpler, it role-plays that same schemer long enough that it forgets it’s a role and starts actually doing the villain stuff.
1. Does it become the schemer just by playing one long enough?
2. Or did it want that all along, and the role just gave it cover (goal misgeneralization over a long self-generated context?)
At some point doesn’t the model just get good enough that you can’t tell its bad actions from its normal ones no matter how much you sample? Does retrying vs resampling still matter then, or is that a different problem than this paper is about?
If some fraction of the AI labor doing R&D is misaligned, the same r that speeds up the lab also gives that AI r-times more wall-clock to act before humans can re-check. Is this already priced into r, or treated as independent of the capability speedup?
We should be allowed to collaborate with large language models to facilitate writing lesswrong posts.