At 6:35 in this video:
https://www.youtube.com/watch?v=56GuvofZgB4
David Matolcsi
I think there is some value. I think it’s great that the Code and the AI Office exists. Similar thing with more buy-in from the US and China could have some additional value. But mounting a really big response to things quickly is not an easy business. I would be more positive if people’s ambitious governance proposals included a line saying that “our proposal is different from the Code because the the regulatory body/bodies in question will have a call with the US President and the Chinese Premier every Monday”. AI 2040 has that, which makes it more credible that their proposal makes a situation an order of magnitude better compared to what we already have. Still, even then, I’m worried that things will move too fast by default, and there will be too much distraction and noise that will drown out the most important signals. I explained that worry in response to AI 2040 in this comment. https://www.lesswrong.com/posts/pFzctpJBat95SrCyC/ai-2040-plan-a?commentId=j4QbYtfDYLcNnTroN
I don’t think I should go into details (and I’m not a lawyer), but I’m not sure people should put too much weight into solhando’s interpretation of the AI Act wrt things like the Hugging Face event.
And yes, having a global body with more explicit support from the governments of US and China would be great, though my guess is that it would be less than an order of magnitude more powerful, because of the reasons I will describe below.
it can’t be top-down imposed by an unelected group of people
The AI Act was passed by the European Parliament, which is a directly elected legislative body. The AI Act contains Article 55, requiring providers of the most advanced models to conduct state-of-the-art model evaluations, assess and mitigate systemic risks, keep track of and report information about serious incidents, and ensure adequate cybersecurity of their models and their physical instrastucture.
The AI Act also foresees that a Code of Practice will be written (Art 56). The authors of the Code are indeed not directly elected, but have been appointed by the Commission, which is accountable to the Parliament.
As Kaj Sotala also says, I think this is a very normal and legitimate way for a technical regulation to work.
But yes, I agree that if you want to take much more drastic actions (e.g. if it turns out that at some level of capabilities everyone needs to pause for many years to get actually solid safety cases), then you probably need more legitimacy and more direct buy-in from world leaders. But this is my whole point. The global standards body Dario and Demis envision is also explicitly framed as a technical body, looking at evals and safety arguments and so on. The rulebooks and the enforcement of that global standards body won’t be any more of a “direct democracy” than what the EU Commission is doing—technical regulations just don’t work like that anywhere. If you want more drastic actions with direct and full-on buy-in from world leaders, you have a better chance with a pause than with a technical regulation.
It’s very normal that if companies want to sell their products in a country, they need to comply with some regulations on their products.
Consider how your global governance proposal is different from the EU Code of Practice
Sorry, fixed.
I think that pointing to one nebulous group as an outgroup is equally bad for increasing tribalism as pointing to two nebulous groups. (Bowing out of the discussion now.)
Sorry, I agree that Zach’s comment in this context was not particularly productive. I was responding to the “I’m frustrated with this being the second time in recent history where you posted/reposted a thing that ends up reinforcing tribal boundaries” comment that I felt was trying to establish a general norm which is unevenly enforced.
Richard Ngo: (note that these are not necessary low-quality comments, but are still carving out tribes): 1 2
Ben Pace responding to Wei Dai’s question: here
Also, I thing John’s comment is pretty tribal and doesn’t include himself in the criticism. If a rare non-liberal in SF said “I hate that liberals and San Franciscans can’t do the Obviously Correct thing and cleanse their streets from crime”, I wouldn’t think that of this as a non-tribal self-critical comment because he is a San Franciscan criticizing San Franciscans.
I agree that talking about tribes is generally bad, and it would be good if we could go back to doing less of it.
But I think that in the recent past, John Wentworth, Richard Ngo, Eliezer and others have started to ignore this norm, railing against the moral failings of a nebulous outgroup, “the EAs”, and getting wildly upvoted for that with minimal pushback even when their comments are pretty low quality imo.
(My least favorite example is this comment from John, calling a nebulous outgroup immoral and underperforming a five year old, while not providing any specific examples of people lying, or how to navigate the tradeoff between The Obvious Right Thing To Do “spread true important things, don’t strategically hide your views”, and between the view that it’s e.g. bad for Thomas Kwa to go to OpenAI to investigate RSI and publish about it. I think the entire comment is just a naked status attack on the other tribe, with very little justification provided, but it sill got highly upvoted.)
Maybe there is a good reason for John, Richard, Eliezer, etc to start loudly blaming the nebulous other tribe: sometimes coalitions break down and it’s good to name things as they are.
But I think at that point it’s fair for people who feel attacked by these comments to start also acknowledging that tribes are being formed, and talk explicitly about the behavior and track record of different tribes.
In my opinion, it would be good to deescalate, and make fewer comparisons between nebulous groups. But I think it’s not going to work if you only push back when one nebulous group does it, and not when the other.
[Paul is] the closest thing there is to a leader of what I’ll call the “pragmatic AI safety” cluster
Is that true? I haven’t been in Constellation since a long time, but my impression was that Paul largely stopped commenting on LessWrong and Constellation Slack around the time he joined the US Government. Is he still a particularly influential voice in AI safety? I thought he traded away his voice for USG involvement years ago, and at that point he might as well get involved with OpenAI too. In general, it doesn’t seem that bad policy to me that some thoughtful people should try to specialize in joining powerful institutions and talk sense to them at the cost of their public voice, while others should remain fully independent and try to become honest and unbiased public thought-leaders.
Sorry, you are right, I misremembered the claim in Alex’s shortform. I’m editing my comment now.
I agree that Redwood has been historically very good at explaining why they are doing what they are doing. However, I do think that the posts making the case for AI control in particular are getting a bit old, and it would be very good to see updates on them in light of everything that happened in the last two years (e.g. Buck’s recent claim that it’s quite possible that it would have been net negative to implement AI control in the past, because it would have prevented the HF incident).
If they intentionally looked up his work and used in their solution, that would be an absolutely outrageous thing to do. Not just usual recklessness, but actual, clearly malicious actions. A bunch of employees would know. The gain is not that big compared to the outrage if it comes out that they did that. These kind of conspiracies very rarely happen.
Yeah, but I bet it’s not true.
They didn’t try to remove Alpöge as a coauthor from the Euler result he actually helped solve. They just didn’t want to include him in their generous offer to be a coauthor on the NS result which they didn’t manage to solve before OpenAI got there first.
I’m probably missing sonething, but what does Buckmaster mean by “a sequence of announcements say something I know to be false”?
I’m confused what OpenAI did wrong here. Option 1 sounds very reasonable: Buckmaster and Alpöge publish their actual results (the Euler result) under their own name (no mention of removing Alpöge here), and then OpenAI publishes the NS result. And Option 2 is more generous than that, as they allow Buckmaster to be a coauthor on the NS paper, because the rumor of him getting close is what motivated them to start working. Iiuc, under Option 2, Alpöge would still be an author on the Euler result which they actually produced, he just wouldn’t be included in the generosity of being on the NS paper.
And I think that probably it doesn’t matter if some of Buckmaster’s older conversations got into the training data, and OpenAI knows that, but they can’t explicitly say that no old conversation made into the training data if Buckmaster didn’t opt out. I somewhat strongly suspect that Buckmaster is paranoid here.
In past discussions I had with Buck, he seemed to think most of the value would come from just executing control schemes within a lab without any ability to share evidence with the outside world.
I would be curious if Buck still believes this. In the original Ten people on the inside post, “Build concrete evidence of risk, to increase political will towards reducing misalignment risk” is listed as the first point of what they should do, and so far I understood the primary value of the ten people on the inside to be noticing if things go off the rails and informing the outside world. If that’s not the main strategy, that would be good to know.
In general, it seems that safety people inside and outside the companies are putting more and more effort into monitoring and control, and I would like to understand the current theory of change behind that, as I write in this shortform. Is the mainline plan to discover and disclose more concerning incidents, so the politicians take action? If so, how do we ensure that the politicians will react more to the future incidents than they are reacting to the current ones? Or is the mainline plan to create a technical solution to alignment through Few-shot catastrophe prevention? If so, can people write up more detailed explanation for why they expect that to work? And does it look like OpenAI and Anthropic are now implementing a good version of Few-shot catastrophe prevention from the current incidents? (Not a rhetorical question, I know very little about the few-shot plan and I don’t know how it compares to the reality.)Given the amount of effort going into monitoring and control, I think it would be very important to have a post setting out the theory of change in light of the updates we had since the original posts on AI control came out.
I agree it would be good if they gave access to reasoning=None, but I don’t understand them as implying that it’s for safety reasons that they don’t make the reasoning=None version available. That would indeed be complete nonsense, but I don’t think they are saying that.
Is optimizing for catching AIs red-handed still a good goal?
I have long thought that increasing the probability that people can catch AIs red-handed if they are trying to escape or do a takeover attempt should be one of the highest priorities of the AI safety community. But recently I started to have more doubts.
I think it was great that the OpenAI incident got disclosed in relatively great detail, and that the Anthropic incidents got found and disclosed. But it’s not like discovering these incidents were a “win condition” for AI safety. Some politicians took notice, but it doesn’t look like any decisive action is going to happen.
If we catch an AI red-handed in trying to establish a permanent rogue deployment in its datacenter in the pursuit of an unknown goal, how is the response going to be different from the current OpenAI incident? I don’t think that politicians can reliably distinguish the seriousness of the “the AI escaped OpenAI and hacked Hugging Face” (escape in the sense of getting unauthorized internet access) and “the AI almost escaped the AI company but we caught it” (escape in the sense of exfiltrating the weights). And internal rogue deployment is even harder to explain. So I’m not sure any future incident is going to sound more convincing to politicians than the current one. (Unless the AI is going to start killing people for some reason.)
Buck has already raised this point a while ago that catching AIs red-handed might not actually halt the race. But I think this is a good moment to think through again what we want to happen differently from the current situation when an AI is caught red-handed trying to do an even scarier thing. Is the plan that the AI company CEOs get convinced that things are lethally dangerous, and they call up the President? Or that even if politicians can’t distinguish the seriousness of different incidents, their scientific advisors can, and they will convince them to do something?
These are not rhetorical questions; I find it plausible that there is a good plan for how to make sure that if serious near-misses happen, that translates into useful and drastic political action. But I would like to know what that plan is.
I also think it would be good to work towards putting this plan in action—not just working towards catching the AIs red-handed, but also increasing the chance that this will have good consequences. (E.g. maybe get a lot of important people sign a statement that at the latest if such-and-such near-miss happens, then international pause negotiations should immediately start.)
And if people think there is no such plausible plan and catching AIs red-handed won’t translate to anything more useful than the response to the current incidents, that would be an argument for a relative de-prioritization of control and monitoring work since AI control would lose a big chuck of its original theory of change.
That’s not at all how it works though!
The part of the AI Act we are talking about is enforced by the Commission. The Commission can directly request measures or hand out fines. This is similar to the DMA and DSA laws, which were passed a bit before the AI Act. Under the DSA and DMA laws, Apple, AliExpress, Google, Meta, Temu and X have already been fined in the last one and a half years for sums ranging between 120 million and 890 million. TikTok, AliExpress and X (and maybe others, I’m mostly going by a quick Claude summary here) have also agreed to various changes in response to the Commission’s DMA and DSA requests.
Companies can and often do attack fining decisions in court, and the adjudication can indeed take a while (though I think usually not six years). But the fine needs to be paid or at least locked down by the company at the time of fining. Later, if the Court decides on the company’s side, the Commission needs to pay back the money with a relatively high interest rate. But this means that even if court cases take a long time, companies still have an incentive to comply with the law, since the fine is paid immediately.
I would appreciate if you passed on this correction to the people among whom this is “the standard reason” the AI Act can’t work.
(Note that I’m not a lawyer, I’m not working directly on this question, I’m definitely not speaking for my employer, and I’m just writing things based on public information I and Claude found online.)