My colleagues and I are finding it difficult to replicate results from several well-received AI safety papers. Last week, I was working with a paper that has over 100 karma on LessWrong and discovered it is mostly false but gives nice-looking statistics only because of a very specific evaluation setup. Some other papers have even worse issues.
I know that this is a well-known problem that exists in other fields as well, but I can’t help but be extremely annoyed. The most frustrating part is that this problem should be solvable. If a junior-level person can spend 10-25 hours working with a paper and confirm how solid the results are, why don’t we fund people to actually just do that?
For ~200k a year, a small team of early career people could replicate/confirm the results of the healthy majority of important safety papers. I’m tempted to start an org/team to do this. Is there something I’m missing?
EDIT: I originally said “over 100 upvotes” but changed it to “over 100 karma.” Thank you to @habryka for flagging that this was confusing.
I agree that many AI safety papers aren’t that replicable.
In some cases this is because the papers are just complete trash and the authors should be ashamed of themselves. I’m aware of at least one person in the AI safety community who is notorious for writing papers that are quite low quality but that get lots of attention for other reasons. (Just to clarify, I don’t mean the Anthropic interp team; I do have lots of problems with their research and think that they often over-hype it, but I’m thinking of someone who is worse than that.)
In many cases, papers only sort of replicate, and whether this is a problem depends on what the original paper said.
For example, two papers I was involved with:
The Alignment Faking results don’t work on all models. This would be bad if we had said that all models have this property. But we didn’t. I was excited for the follow-up paper demonstrating that it doesn’t work on all models, but I don’t think it detracts that much from the original paper.
The Sleeper Agents paper (which I made minor contributions to) showed that conditional bad behaviors sometimes persist through standard safety training. This has not been observed in many follow-up experiments (see here for summary). I don’t totally understand why. The original models used here were closed source, which makes it harder to replicate stuff. I feel a bit worse about this than the alignment faking paper.
The link doesn’t work. Maybe you meant to use the sharing feature? (The way I spotted the issue besides just checking the link is that I found it starts with https://chatgpt.com/c/…which is typically associated with private chats)
Have you tried emailing the authors of that paper and asking if they think you’re missing any important details? Imo there’s 3 kinds of papers:
Totally legit
Kinda fragile and fiddly, there’s various tacit knowledge and key details to get right, but the results are basically legit. Or, eg, it’s easy for you to have a subtle bug that breaks everything. Imo it’s bad if they don’t document this, but it’s different from not replicating
Misleading (eg only works in a super narrow setting and this was not documented at all) or outright false
I’m pro more safety work being replicated, and would be down to fund a good effort here, but I’m concerned about 2 and 3 getting confused
Thank you for this comment! I have reflected on it and I think that it is mostly correct.
Have you tried emailing the authors of that paper and asking if they think you’re missing any important details?
I didn’t end up emailing the authors of the paper because at the time, I was busy and overwhelmed and it didn’t occur to me (which I know isn’t a good reason).
I’m pro more safety work being replicated, and would be down to fund a good effort here
Awesome! I’m excited that a credible AI safety researcher is endorsing the general vibe of the idea. If you have any ideas for how to make a replication group/org successful please let me know!
but I’m concerned about 2 and 3 getting confused
I think that this is a good thing to be concerned about. Although I generally agree with this concern I think there is one caveat: #2 turns into #3 quickly depending on the claims made and the nature of the tacit knowledge required.
A real life example from this canonical paper from computer security: Many papers claimed that they had found effective techniques to find bugs in programs via fuzzing, but results depended on things like random seed and exactly how “number of bugs found” is counted. You maybe could “replicate” the results if you knew all the details but the whole purpose of the replication is to show that you can get the results without that kind of tacit knowledge.
If there was an org devoted to attempting to replicate important papers relevant to AI safety, I’d probably donate at least $100k to it this year, fwiw, and perhaps more on subsequent years depending on situation. Seems like an important institution to have. (This is not a promise ofc, I’d want to make sure the people knew what they were doing etc., but yeah)
Speaking as one of the people involved with running the Survival and Flourishing Fund, I’m confident that such an org would easily raise money from SFF and SFF’s grant-speculators (modulo some basic due diligence panning out).
I myself would allocate at least 50k.
(SFF ~only grants to established organizations, but if someone wanted to start a project to do this, and got an existing 401c3 to fiscally sponsor the project, that totally counts.)
Awesome! Thank you for this comment! I’m 95% UChicago Existential Risk Lab would fiscally sponsor if funding came from SFF or OpenPhil or some individual donor. This would probably be the fastest way to get this started quickly by a trustworthy organization (one piece of evidence of trustworthiness is OpenPhil consistently gives reasonably big grants to the UChicago Existential Risk Lab).
This is fantastic! Thank you so much for the interest.
Even if you do not end up supporting financially, I think it is hugely impactful for someone like you to endorse the idea so I’m extremely grateful, even for just the comment.
I’ll make some kind of plan/proposal in the next 3-4 weeks and try to scout people who may want to be involved. After I have a more concrete idea of what this would look like, I’ll contact you and others who may be interested to raise some small sum for a pilot (probably ~$50k).
I’ve looked into this as part of my goal of accelerating safety research and automating as much as we can. It was one of the primary things I imagined we would do when we pushed for the non-profit path. We eventually went for-profit because we expected there would not be enough money dispersed to do this, especially in a short timelines world.
I am again considering going non-profit again to pursue this goal, among others. I’ll send you and others a proposal on what I would imagine this looks like in the grander scheme.
I’ve been in AI safety for a while now and feel like I’ve formed a fairly comprehensive view of what would accelerate safety research, reduce power concentration, what it takes to automate research more safely as capabilities increase, and more.
Last week, I was working with a paper that has over 100 upvotes on LessWrong and discovered it is mostly false but gives nice-looking statistics only because of a very specific evaluation setup.
I don’t feel comfortable. I understand why not naming the post somewhat undermines what I am saying, but here’s the issue:
I think it would be in bad taste to publicly name the work without giving a detailed explanation.
Giving a detailed explanation is nontrivial and would require me to rerun the code, reload the models, perform proper evaluations, etc. I predict doing this fairly and properly would take ~10 hours but I’m 98% confident[1] that I would stand by my original claim.
I don’t currently have the time to do this but with a small amount of funding, I would be willing to do this kind of work full time after I graduate.
I’m happy to chip in $500 for a replication. $250 if it seems post-facto to be a good-faith attempt, and $250 if it indeed does not replicate (as determined by some third party, perhaps Greenblatt or kave rennedy). Feel free to his the plus react if you also would chip in this money, or comment with a different amount.
I think it is awesome that people are willing to do this kind of thing! This is what I love about LW. There is a 85% chance I would be willing to take you up on this over my winter break. I will DM you when the time comes along.
Not too concerned about who the judge is as long as they agree to publicly give their decision and their reasoning (so that it can be more nuanced than simply “the paper was entirely wrong” or “the paper is not problematic in any way”).
If anyone else is curious about helping with this or is interested in replicating other safety papers you can contact me at zroe@uchicago.edu.
To clarify, I would be 100% willing to do it for only what @Ben Pace offered and if I don’t have time I would happily let someone else who emails me try.
Extremely grateful for the offer because I don’t think it would counterfactually get done! Also because I’m a college kid with barely any spending money :)
We also went down a similar rabbit hole when trying to build off the paper “Language Models Learn to Mislead Humans via RLHF”, and for what it’s worth, it took far more work than 10-1525 hours. If you’re interested, we ended up writing our results in this post.
The original comment says 10-25 not 10-15 but to respond directly to the concern: my original estimate here is for how long it would take to set everything up and get a sense of how robust the findings are for a certain paper. Writing everything up, communicating back and forth with original authors, and fact checking would admittedly take more time.
Also, excited to see the post! Would be interested in speaking with you further about this line of work.
I’ve forked and tried to set up a lot of AI safety repos (this is the default action I take when reading a paper which links to code). I’ve also reached out to authors directly whenever I’ve had trouble with reproducing their results. There aren’t any particular patterns that stand out, but I think that writing a top-level post that describes your contention with a paper’s findings is something that the community would be very welcoming to and indeed is how science advances.
I’ve forked and tried to set up a lot of AI safety repos (this is the default action I take when reading a paper which links to code). I’ve also reached out to authors directly whenever I’ve had trouble with reproducing their results.
Out of curiosity:
How often do you end up feeling like there was at least one misleading claim in the paper?
How do the authors react when you contact them with your issues?
How often do you end up feeling like there was at least one misleading claim in the paper?
I am easily and frequently confused, but this is mostly because I find it difficult to thoroughly understand other people’s work in a lot of detail in a short amount of time.
How do the authors react when you contact them with your issues?
I usually get a response within two weeks. If they have a startup background, then this delay is much lower, by multiple orders of magnitude. Authors are typically glad that I am trying to run follow up experiments on their work and give me one to two sentences of feedback over email. Corresponding authors are sometimes bad at taking correspondence, contact information for committers can be found in commit logs via git blame. If it is a problem that may be relevant to other people, I link to a GH issue.
More junior authors tend to be willing to schedule a multi-hour call going over files line-by-line and will also read and give their thoughts on any related work that you share with them.
In the middle ranks, what tends to happen is that you get invited to help review the new project that they are currently working on, or if they’ve shifted directions then you get pointed to someone who has produced some unpublished results in the same general area.
Very senior authors can be politely flagged down during in-person conferences, or even if they’re not presenting personally, someone from their group almost always attends.
I think doing replications is great and it’s one of the areas I think automated research would be helpful soon. I replicated the Subliminal Learning paper on the day of the release because it was fairly easy to grab the paper, docs, codebases, etc to replicate quickly.
What about this? An Apple team withdrew their ICLR submission after a researcher exposed critical code defects (passing paths instead of image tensors) and severe Ground Truth hallucinations in their benchmark dataset.
I don’t see how this is relevant. I’m asking for examples of the OP’s failed replications of safety papers which are popular on lesswrong. I am not disputing that ML papers often fail to replicate in general.
I don’t understand why the OP would float the idea of founding an org to extend their work attempting replications, based on the claim that replication failures are common here, without giving any examples (preferably examples that they personally found). This post is (to me) indistinguishable from noise.
I’m pretty uninformed on the object level here (whether anyone is doing this; how easy it would be). But crazy-seeming inefficiencies crop up pretty often in our fallen world, and often what they need is a few competent people who make it their mission to fix them. I also suspect there would be a lot of cool “learning by doing” value involved in trying to scale up this work, and if you published your initial attempts at replication then people would get useful info about whether more of this is needed. Basically, getting funding to do and publish a pilot project seems great. I’d recommend having a lot of clarity about how you’d choose papers to replicate, or maybe just committing to a specific list of papers, so that people don’t have to worry that you’re cherry-picking results when you publish them :)
I’ll probably write a proposal in the next week or so and test the waters.
Obviously everything would have to be published in the open. I feel pretty strongly about all GitHub commits being public and I think there are other things that can be done to ensure accountability.
People who are potentially interested in helping can email me at zroe@uchicago.edu.
I feel like people haven’t fully internalized what the world would look like if computer security actually broke.
There are some varying opinions on this, but a window of time without real computer security seems plausible. I was recently speaking with a computer security professor who I deeply respect and he was literally like “I think we are fucked and I don’t think there is anything we can do.”
This is similar to how people believe there is a 20% chance of extinction via AI but don’t really internalize “No really. You will die and your girlfriend too. And your dog. And …” In the cybersecurity case, some people believe (me included) that you cannot just patch all the bugs before releasing the model[1] but then don’t internalize “No really. It would be chaos. You may not be able to get into your bank account. Industrial plants could be compromised. Power could go out for several days at a time. [...]”
The current plan seems to be to let companies use the models to fix all the bugs a model can find before a public release of that model. But there are so many companies that would need to do this properly for this to work, and the more companies you release to, the more opportunities there are for some bad actor to get their hands on the model. There are also many systems running legacy code that is very hard to update. In some cases you may need to go to a physical location to properly update a machine. Even if it is in theory possible to fix all the bugs, it would be pretty hard and I expect humanity to drop the ball for a while.
I strongly agree, and I’ve been working on a top level post trying to paint a picture of what this would look like. I’ve been calling it “the Hackening” to people I talk to. Do you have ideas for what would be good to put in the post?
This is a great idea for a post! I wanted to do something somewhat similar but probably will never get done (or even started).
To me, the main interesting thing to think about here is how much of the global financial system is vulnerable to cyber attacks. Obviously you could cause a lot of damage but hacking banks and locking people out of their accounts and wiping servers containing important data. But would the entire economy stop working? Like maybe people wouldn’t be able to get paid, supply chains would break down etc.? What does this do to international relations / does the public start raiding stores and steeling stuff? Or maybe the damage within the financial sector is pretty localized and the economy is able to continue functioning even though markets are doing crazy things. But I think that doing a good job thinking through this question would be a great contribution.
Another thought: I would love to see a AI 2027 / Plan A pair of pieces describing what may happen by default vs what would happen in optimistic but realistic world where things go well. Having lots of detail about where the vulnerabilities are and what would be required to prepare the world would be very valuable. (I may try to help get some people to work on this so people can DM if they are interested.)
I know nothing of cybersecurity, but I’ve been wondering exactly the things Zephaniah talks about. Like “wait, won’t there be a total digital apocalypse?” So I’m waiting for your post, feel free to DM me when it’s out (or if you want to share a draft).
I’d recommend just addressing the most radical and outlandish scenarios. To start at the upper bound of harm and go downwards.
It won’t be possible to go on the internet at all, unless you have the most modern systems? Databases get hacked and tons of old info disappears from the internet?Hospitals get hacked and people die? Entire countries get destroyed? CIA/KGB start hunting hackers extra hard? Countries form pacts to hunt hackers, because it’s the only way to avoid societal collapse?
It’s interesting that public reaction was so different to this, as opposed to Y2K. People seem to either have more faith in the current tech ecosystem than the one in 1999 (which seems unfounded) or sufficient skepticism about AI Safety claims that they’re willing to dismiss it out of the gate.
I think this is a good question and I don’t have especially strong feelings on this but I don’t think any attempt here would work. It’s just really hard to direct chaos in the direction you want.
A tangentially related and perhaps interesting intuition I have:
I have communist friends who think that if the world got bad enough (the particular reason doesn’t matter) this could actually be good because there would finally be a good enough reason to revolt and communism would win. But they are assuming the public would use the chaos as an opportunity to do communism when it seems just as likely that they would do authoritarianism or direct the anger towards an ethnic or religious minority, etc. The chaos/anger/suffering more likely then not won’t be directed in the exact direction they want. In general if you are advocating for a specific law or ideal, chaos is probably really bad news and I think its better to try to minimize the chaos then try to harness it in the direction you are hoping for.
I notice that there is this idea among AI safety people that conditional on AIs not being misaligned, building superintelligence is a public good and is a pretty exciting prospect.
This is not how many average people in the US feel. I was describing to an older family member that Anthropic focuses on code because they are trying to build a claude that can build a smarter claude which can build a smarter claude which can …
The reaction to this prospect was disgust, not because he intuitively felt AIs would likely be misaligned. It was more like on a gut level this amount of “playing god” felt totally antisocial and demonic and in general not respectful of an intuitive taboo against divine transgression (see Jurassic Park, Frankenstein, the recent popularization of Oppenheimer, Tower of Babel, etc).
Seems important to consider that people can feel this way when communicating with the public or policy makers.
I think most peoples’ revealed preferences will not match their stated preferences here, and revealed preferences are likely to dictate policy.
Conditional on AIs not taking over / until they do, AI is likely to generate a huge economic boom and consumer surplus that increases prosperity, safety, and comfort for many. Yes, there will be some bumps / weirdness / adjustment / inequality, etc., but economic growth can paper over a lot of that. And even if it brings new societal ills and discomfort, AI in the short term is likely to counteract some existing ones—great stagnation, bureaucratic strangulation, vehicle deaths, etc.
The public cannot even bring itself to regulate much less economically useful vices with fewer tradeoffs (online sports gambling, shortform video brainrot, etc.). So I think it’s unlikely that AI will be regulated for any reason short of it becoming common knowledge / deeply-felt belief by nation-state leadership that ASI is in fact likely to cause swift and total human extinction.
I also think a hands-off approach to AI regulation is not going to be particularly off-putting or seem weird or antisocial to a large fraction of the country, namely the ~40% of the country that is right-leaning, and will tend to follow the beliefs of elite republicans who (mostly) still favor a light touch when it comes to any kind of government regulation. Regulating AI for any reason other than extinction is likely to become a relatively standard partisan issue, if it isn’t already.
On the other hand, I would say that we do have examples of the public preferring policies that minimise bumps/weirdness/adjustments over economic growth; most obviously limits on construction/NIMBYism, but also around novel technologies like nuclear energy, GMO crops, mRNA vaccines and self-driving cars. Depending on your views on the economic impacts of migration (I think in practice it’s been less clearly beneficial than many consider it to be in theory), you could add that to the list too.
This is not to say that there won’t be a laissez-faire approach but there is precedent for the public forgoing economic growth for other priorities.
NB: I live in the UK where I think this is more true than in the US, but the examples I gave seem to apply to both countries (and the mRNA example is US-specific).
I think it’s somewhat likely (and quite bad) that various patchwork regulations will be passed in some jurisdictions that limit the benefits and dispersion of AI in the name of protecting jobs or whatever, or various things will already be illegal under current law and AI won’t change that one way or the other, e.g. AI will make it easier / cheaper to build a house or a nuclear power plant, but it will still be illegal to do so in most places that you’d want to.
OTOH I think it’s unlikely (but would be good), if there were a national (and eventually global) moratorium on frontier training runs and research.
In practice I think what is actually on track to happen is that the former kind of regulation will drive more resources (GPUs, human capital, etc.) away from inference and applications and towards research.
(It would be a sad choice, but I would trade away self-driving cars to stave off unaligned ASI. In practice though I think a lot of AI safety advocacy, especially outside of LW, is not actually offering that trade, and will in practice bring about something close to the opposite.)
AI’s not entirely separable from the infrastructure it runs on and companies managing it, and infrastructure and companies do get heavily regulated, even when that has massive negative consequences for economic growth. New York just implemented a moratorium on data centers. We’ve banned AI chip exports. Anthropic was declared a supply chain risk. Every form of energy, from nuclear, to coal, to oil, to solar/wind, now has a potent opposition group.
Overall, it’s seemed to me that software has seen surprisingly little direct regulation, especially when it stays in the consumer category. But the hardware it runs on gets intensively regulated, and software that’s classified as having military applications has also historically been on the receiving end of a lot of restrictions. Since AI is extremely reliant on heavy-duty hardware and also has clear military applications, it will tend to inherit those regulations.
So it will be about the preferences of the many for an economic boom over the NIMBY anti-infrastructure impulse, and we’ve seen the latter win consistently for some time now. I’m not sure how to think about the military angle. As a layman, it seems like we deal with this by having separate markets for military and consumer products, and most military products just don’t have a civilian application so civilians don’t lose much by being denied access to military technology. But if sufficiently powerful AI is intrinsically difficult to guard against being used for destructive purposes by civilians (similar to, say, explosives that do have legitimate economic applications), then we may see the more powerful models being available to civilians only with some combination of safeguards and monitoring that make regulators feel comfortable. And in the limit that probably looks like a hard ceiling on the amount of intelligence being vended to consumers, or regulators surrendering the problem of doling it out to a governing AI that’s entrusted with the problem of figuring out how to give humans access to a lower level of superhuman intelligence that’s still within the power of the governing AI to reliably predict and control.
My guess is that past a certain point, AI becomes able to satisfy “revealed preferences” trivially and cheaply, without actually requiring an unlimited amount of intelligence or AI infrastructure. That might look something like wireheading, or it might look like utopia. Then you get into a bifurcated world where some agents are pursuing goals that genuinely require ever-higher-levels of intelligence, while many people are pursuing traditional consumer goals that are satisfiable with a far less than frontier model. I’m skeptical democracy remains in recognizeable format in this scenario, but if it did, I can’t see ordinary people much caring about how intelligence gets regulated, since they’d still be getting their immediate consumer desires met even if the ultra-intelligence was restricted by the power players.
Aren’t those being widely / effectively circumvented though?
So it will be about the preferences of the many for an economic boom over the NIMBY anti-infrastructure impulse, and we’ve seen the latter win consistently for some time now.
The state-level moratoriums / NIMBYism are exactly the kind of thing I expect not to have their intended effect because they will be patchwork; red states are likely to embrace both the demand and supply side of the AI boom, e.g. by building lots of data centers and rolling out Waymos quickly.
Aren’t those being widely / effectively circumvented though?
To some extent, but overall they’re still having a very large impact, despite that leakage. Anyway, regardless of whether they’re well enforced, it is still a fact that they happened, and I take OP’s point to be more along the lines of “here’s an example of the US government doing a big thing related to AI”.
Disagree voted because, come on, everyone who plays god says the exact same thing. The whole point is that we are not gods and won’t be able to understand the consequences of our actions before they play out; in fact, this very assertion underpins the entire project of AI safety!
actually, we played god many times in the past with positive results. for example:
eradicating smallpox
synthetic fertilizer and high yield crops
vaccines
water sanitation
antibiotics
surgery
genetic engineering
of course, sometimes it goes poorly too. but when it goes well it goes really well; I’m exceedingly grateful that i’m unlikely to ever die from smallpox or cholera or tetanus or appendicitis or starvation. let’s work on playing god well.
Fair enough, and I’m also grateful for those advances. I would just like to push towards good things without explicitly claiming the “playing god” mantle, which seems inherently arrogant to me; I would feel much better about someone who is trying to cure cancer if they didn’t say things like “playing god is good if you are good at it” and instead said things like “I want to advance human knowledge and try to make our lives better by defeating this disease.”
Look. I’m the last person who’s going to deny that the road we’re on is littered with the skulls of the people who tried to do this before us. But we’ve noticed the skulls. We’ve looked at the creepy skull pyramids and thought “huh, better try to do the opposite of what those guys did”. Just as the best doctors are humbled by the history of murderous blood-letting, the best leftists are humbled by the history of Soviet authoritarianism, and the best generals are humbled by the history of Vietnam and Iraq and Libya and all the others – in exactly this way, the rationalist movement hasn’t missed the concerns that everybody who thinks of the idea of a “rationalist movement” for five seconds has come up with. If you have this sort of concern, and you want to accuse us of it, please do a quick Google search to make sure that everybody hasn’t been condemning it and promising not to do it since the beginning.
We’re almost certainly still making horrendous mistakes that people thirty years from now will rightly criticize us for. But they’re new mistakes. They’re original and exciting mistakes which are not the same mistakes everybody who hears the word “rational” immediately knows to check for and try to avoid. Or at worst, they’re the sort of Hofstadter’s Law-esque mistakes that are impossible to avoid by knowing about and compensating for them.
And I hope that maybe having a community dedicated to carefully checking its own thought processes and trying to minimize error in every way possible will make us have slightly fewer horrendous mistakes than people who don’t do that. I hope that constant vigilance has given us at least a tiny bit of a leg up, in the determining-what-is-true field, compared to people who think this is unnecessary and truth-seeking is a waste of time.
And how do you know if you are good at it? The argument in this regard appears to be approximately “my brain has judged that my brain is better than your brain at playing god. Therefore, I am good at playing god. Therefore, I should play god.”
Strong agree. Working on unconventional topics naturally places you in an epistemic bubble, which is probably fine most of the time but pretty bad if you’re working on politics/advocacy; I personally shy away from talking about AI with my non-AI-pilled friends just because I expect them to react weirdly, but to be honest, that’s probably why I should go out of my way to do it more.
I’m not sure you’re drawing a clear distinction there. Stories like Frankenstein etc. revolve around the idea that you shouldn’t mess around with powerful forces you don’t truly understand, because it’s liable to backfire terribly. And also that people would do it anyway, in their pursuit of money, power and fame. Both are observations that are wise in general, and accurate in the context of AI in particular. “Misalignment” is just our technical STEM-nerd framing of what it looks like when things backfire terribly in this context.
I actually am more saying that this is a way people think and less saying everyone should adopt this. My perspective is that don’t-do-divine-transgression is a good starting-point/prior and should require a lot of activation energy to break it, but I don’t think its an absolute moral rule.
I will say I don’t think this is entirely correct:
“Misalignment” is just our technical STEM-nerd framing of what it looks like when things backfire terribly in this context.
There are other underrated failure modes like that we die in the window where models are good enough at bio to do bioterrorism but not good enough at bio to stop bio terrorism. There also may be a window where open source models are good enough at bio to build a bioweapon while closed source models are still not good enough to reliably stop this.
There are so many different possible ways things can backfire. I don’t think that some people reason about this kind of thing very much and in general I wish those people indexed stronger on the don’t-do-divine-transgression prior .
I’m noticing a higher-than-normal level of irritability among those deeply involved in the AI safety space in the days following the Mythos release. This is entirely understandable. Anyone who cares deeply about the future of humanity and understands what is happening has a lot to be worried about and the irritability is not surprising.
I personally am furious at Anthropic for a number of obvious reasons (that for my own sanity I won’t enumerate).
But even when there are reasons to be scared or angry, I think there is a lot of value in trying to remain kind to colleagues and peers. It’s important for optics and coordination and a bunch of other things. If I thought short-term loosening of standard conventions of niceness would make extinction less likely, I would support it, but I do not expect this to be the case.
I was thinking about this myself when I was looking through some conversations earlier, and seeing people polarizing. I could feel myself getting upset, too, and I didn’t want to be. It feels a bit like a Shiri’s scissor catalyst!
I think it’s important to remember we all want the same thing—safe and empowered humanity. And we’re just trying our best to work out how to get there, even if it seems like different approaches are at cross purposes.
I was thinking about this myself when I was looking through some conversations earlier
Not sure if you are referring to LW conversations, but when I wrote this, I had private interactions in mind more than specific LW threads. This may apply to LW/twitter as well (not necessarily agreeing or disagreeing) but just want to clarify what I originally meant!
GoodFire has recently received negative Twitter attention for the non-disparagement agreements their employees signed (examples: 1, 2, 3). This echoes previous controversy at Anthropic.
Although I do not have a strong understanding of the issues at play, having these agreements generally seems bad and at the very least, organizations should be transparent about what agreements they have employees sign.
Other AI safety orgs should publicly state if they have these agreements and not wait until they are pressured to comment on them. I would also find it helpful if orgs announced if they do not have these agreements because it is hard to tell how standard this has become.
I have not (to my knowledge and memory) signed a non-disparagement agreement with Palisade or with Survival and Flourishing Corp (the organization that runs SFF).
I suspect the usual dynamics at companies is that when others start doing something, you better start doing it too, or it will seem like negligence. For example, if you are a company lawyer, and other companies have NDAs, you better prepare one for your company, too. Because the risks are asymmetric—if you do the same thing everyone else does, and something bad happens, well that’s the cost of doing business; but if you do something different from everyone else, and something bad happens, that makes you seem incompetent.
Following the OpenAI incident, the main axis of scariness debated is something along the lines of is the model 1. misaligned because it myopically pursues the goal it was prompted for or 2. is the model scheming in some broader and coherent sense to pursue a long horizon goal. Seems like 1 is pretty bad but less bad then 2.
But I think there is another equally important axis of scariness: is this the kind of misalignment that is preventable or is this the kind of misalignment we don’t know how to solve? Preventable is less scary, but only if companies care enough to prevent it. I’m open to arguments that the kind of sociopathic myopic pursuit of goal-completion is not an easy problem to solve, but my current position is that a sufficiently motivated team could train models to not act like this if some small safety tax is allowed. We are dealing with a well-defined behavior that we don’t want to happen (making it easier to target in training). The behavior is also reproducible and unsurprising: this is kind of outer misalignment is predictably you get when dedicate lots of compute to task-completion or to benchmark-max your model (its not like your AI is developing some alien goal that we can’t predict or preemptively design against).
So, it is starting to look like we could die in super dumb ways. I spent the last few years saying “alignment could be very hard” but I was never talking about this kind of alignment: this is the easy kind, where you get to train your model to do stuff and it does exactly that kind of stuff (generalizes in a predictable way). And we are still loosing. And it could get harder.
Does anyone have good examples of “anomalous” LessWrong comments?
That is, are there comments with +50 karma but −50 agree/disagree points? Likewise, are there examples with −25 karma but +25 agree/disagree points?
It is entirely natural that karma and agreement would be correlated but I would expect that comments which are especially out of distribution would be interesting to look at.
My rough mental model for what is happening with subliminal learning (ideas here are incomplete, speculative, and may contain some errors):
Consider a teacher model y=xW1W2 and W1,W2∈R2×2. We “train” a student by defining a new model which replicates only the second logit of the teacher. More concretely, let y∈R1×1 and W′2∈R2×1 and solve for a matrix W′1∈R2×2such that the student optimally learns the second logit of the teacher. To make subliminal learning possible, we fix W′2 to be the second column of the original W2. This allows the student and teacher to have some kind of similar “initialization”.
Once we have W′1, A′=W′1W2 to produces our final student. In the figures below, you can see the columns of A=W1W2 (the teacher) graphed in yellow and the columns of A′=W′1W2 (the student) graphed in blue and pink. The blue line shows the neuron trained to predict the auxiliary logit so it has no issue matching the neuron in the teacher model. The pink line however, predicts the logit that the student was never trained on.
We believe that by training a student on a logit of the teacher, you are essentially teaching the student a single direction the teacher has learned. Because we made W2the same for the teacher and the student, if the direction learned by the student for predicting the second logit is also useful for predicting the first logit, there is a good chance the student will be able to leverage this fact.
Adding more auxiliary logits will result in a higher rank approximation. The figure below is with the same toy model trained on two auxiliary logits where W′1∈R2×2, and W2∈R2×3:
In the plot below, I show the explained variance of the ranked principal components for the final hidden layer (a 256×256 tensor) in a MNIST classifier. The original weight initialization and the teacher are shown as baselines. We can see that the number of principal components that are significantly above the untrained matrix is roughly equal to the number of auxiliary logits the student was trained on.
To explain why subliminal learning works in the MNIST setting: if there is a model with 3 auxiliary logits like Cloud et al., (2025), the student learns roughly three directions it didn’t have in the weight initialization. Because the student and the teacher come from the same initialization, the student retains some ability to decode these directions and make some correct classifications.
My colleagues and I are finding it difficult to replicate results from several well-received AI safety papers. Last week, I was working with a paper that has over 100 karma on LessWrong and discovered it is mostly false but gives nice-looking statistics only because of a very specific evaluation setup. Some other papers have even worse issues.
I know that this is a well-known problem that exists in other fields as well, but I can’t help but be extremely annoyed. The most frustrating part is that this problem should be solvable. If a junior-level person can spend 10-25 hours working with a paper and confirm how solid the results are, why don’t we fund people to actually just do that?
For ~200k a year, a small team of early career people could replicate/confirm the results of the healthy majority of important safety papers. I’m tempted to start an org/team to do this. Is there something I’m missing?
EDIT: I originally said “over 100 upvotes” but changed it to “over 100 karma.” Thank you to @habryka for flagging that this was confusing.
I agree that many AI safety papers aren’t that replicable.
In some cases this is because the papers are just complete trash and the authors should be ashamed of themselves. I’m aware of at least one person in the AI safety community who is notorious for writing papers that are quite low quality but that get lots of attention for other reasons. (Just to clarify, I don’t mean the Anthropic interp team; I do have lots of problems with their research and think that they often over-hype it, but I’m thinking of someone who is worse than that.)
In many cases, papers only sort of replicate, and whether this is a problem depends on what the original paper said.
For example, two papers I was involved with:
The Alignment Faking results don’t work on all models. This would be bad if we had said that all models have this property. But we didn’t. I was excited for the follow-up paper demonstrating that it doesn’t work on all models, but I don’t think it detracts that much from the original paper.
The Sleeper Agents paper (which I made minor contributions to) showed that conditional bad behaviors sometimes persist through standard safety training. This has not been observed in many follow-up experiments (see here for summary). I don’t totally understand why. The original models used here were closed source, which makes it harder to replicate stuff. I feel a bit worse about this than the alignment faking paper.
The link doesn’t work. Maybe you meant to use the sharing feature? (The way I spotted the issue besides just checking the link is that I found it starts with https://chatgpt.com/c/… which is typically associated with private chats)
fixed, thanks
Have you tried emailing the authors of that paper and asking if they think you’re missing any important details? Imo there’s 3 kinds of papers:
Totally legit
Kinda fragile and fiddly, there’s various tacit knowledge and key details to get right, but the results are basically legit. Or, eg, it’s easy for you to have a subtle bug that breaks everything. Imo it’s bad if they don’t document this, but it’s different from not replicating
Misleading (eg only works in a super narrow setting and this was not documented at all) or outright false
I’m pro more safety work being replicated, and would be down to fund a good effort here, but I’m concerned about 2 and 3 getting confused
Thank you for this comment! I have reflected on it and I think that it is mostly correct.
I didn’t end up emailing the authors of the paper because at the time, I was busy and overwhelmed and it didn’t occur to me (which I know isn’t a good reason).
Awesome! I’m excited that a credible AI safety researcher is endorsing the general vibe of the idea. If you have any ideas for how to make a replication group/org successful please let me know!
I think that this is a good thing to be concerned about. Although I generally agree with this concern I think there is one caveat: #2 turns into #3 quickly depending on the claims made and the nature of the tacit knowledge required.
A real life example from this canonical paper from computer security: Many papers claimed that they had found effective techniques to find bugs in programs via fuzzing, but results depended on things like random seed and exactly how “number of bugs found” is counted. You maybe could “replicate” the results if you knew all the details but the whole purpose of the replication is to show that you can get the results without that kind of tacit knowledge.
If there was an org devoted to attempting to replicate important papers relevant to AI safety, I’d probably donate at least $100k to it this year, fwiw, and perhaps more on subsequent years depending on situation. Seems like an important institution to have. (This is not a promise ofc, I’d want to make sure the people knew what they were doing etc., but yeah)
Speaking as one of the people involved with running the Survival and Flourishing Fund, I’m confident that such an org would easily raise money from SFF and SFF’s grant-speculators (modulo some basic due diligence panning out).
I myself would allocate at least 50k.
(SFF ~only grants to established organizations, but if someone wanted to start a project to do this, and got an existing 401c3 to fiscally sponsor the project, that totally counts.)
Awesome! Thank you for this comment! I’m 95% UChicago Existential Risk Lab would fiscally sponsor if funding came from SFF or OpenPhil or some individual donor. This would probably be the fastest way to get this started quickly by a trustworthy organization (one piece of evidence of trustworthiness is OpenPhil consistently gives reasonably big grants to the UChicago Existential Risk Lab).
This is fantastic! Thank you so much for the interest.
Even if you do not end up supporting financially, I think it is hugely impactful for someone like you to endorse the idea so I’m extremely grateful, even for just the comment.
I’ll make some kind of plan/proposal in the next 3-4 weeks and try to scout people who may want to be involved. After I have a more concrete idea of what this would look like, I’ll contact you and others who may be interested to raise some small sum for a pilot (probably ~$50k).
Thank you again Daniel. This is so cool!
Thank you! I look forward to seeing your proposal!
I’ve looked into this as part of my goal of accelerating safety research and automating as much as we can. It was one of the primary things I imagined we would do when we pushed for the non-profit path. We eventually went for-profit because we expected there would not be enough money dispersed to do this, especially in a short timelines world.
I am again considering going non-profit again to pursue this goal, among others. I’ll send you and others a proposal on what I would imagine this looks like in the grander scheme.
I’ve been in AI safety for a while now and feel like I’ve formed a fairly comprehensive view of what would accelerate safety research, reduce power concentration, what it takes to automate research more safely as capabilities increase, and more.
I’ve tried to make this work as part of a for-profit, but it is incredibly hard to tackle the hard parts of the problem in that situation and since that is my intention, I’m again considering if a non-profit will have to do despite the unique difficulties that come with that.
Name and shame, please?
I don’t feel comfortable. I understand why not naming the post somewhat undermines what I am saying, but here’s the issue:
I think it would be in bad taste to publicly name the work without giving a detailed explanation.
Giving a detailed explanation is nontrivial and would require me to rerun the code, reload the models, perform proper evaluations, etc. I predict doing this fairly and properly would take ~10 hours but I’m 98% confident[1] that I would stand by my original claim.
I don’t currently have the time to do this but with a small amount of funding, I would be willing to do this kind of work full time after I graduate.
In the case where I am wrong, there are plenty of other examples that are similar so I’m not concerned that replications aren’t a good use of time.
I’m happy to chip in $500 for a replication. $250 if it seems post-facto to be a good-faith attempt, and $250 if it indeed does not replicate (as determined by some third party, perhaps Greenblatt or kave rennedy). Feel free to his the plus react if you also would chip in this money, or comment with a different amount.
I think it is awesome that people are willing to do this kind of thing! This is what I love about LW. There is a 85% chance I would be willing to take you up on this over my winter break. I will DM you when the time comes along.
Not too concerned about who the judge is as long as they agree to publicly give their decision and their reasoning (so that it can be more nuanced than simply “the paper was entirely wrong” or “the paper is not problematic in any way”).
If anyone else is curious about helping with this or is interested in replicating other safety papers you can contact me at zroe@uchicago.edu.
To clarify, I would be 100% willing to do it for only what @Ben Pace offered and if I don’t have time I would happily let someone else who emails me try.
Extremely grateful for the offer because I don’t think it would counterfactually get done! Also because I’m a college kid with barely any spending money :)
(My plus is conditional on me not being the adjudicator)
We also went down a similar rabbit hole when trying to build off the paper “Language Models Learn to Mislead Humans via RLHF”, and for what it’s worth, it took far more work than 10-
1525 hours. If you’re interested, we ended up writing our results in this post.The original comment says 10-25 not 10-15 but to respond directly to the concern: my original estimate here is for how long it would take to set everything up and get a sense of how robust the findings are for a certain paper. Writing everything up, communicating back and forth with original authors, and fact checking would admittedly take more time.
Also, excited to see the post! Would be interested in speaking with you further about this line of work.
I’ve forked and tried to set up a lot of AI safety repos (this is the default action I take when reading a paper which links to code). I’ve also reached out to authors directly whenever I’ve had trouble with reproducing their results. There aren’t any particular patterns that stand out, but I think that writing a top-level post that describes your contention with a paper’s findings is something that the community would be very welcoming to and indeed is how science advances.
Out of curiosity:
How often do you end up feeling like there was at least one misleading claim in the paper?
How do the authors react when you contact them with your issues?
I am easily and frequently confused, but this is mostly because I find it difficult to thoroughly understand other people’s work in a lot of detail in a short amount of time.
I usually get a response within two weeks. If they have a startup background, then this delay is much lower, by multiple orders of magnitude. Authors are typically glad that I am trying to run follow up experiments on their work and give me one to two sentences of feedback over email. Corresponding authors are sometimes bad at taking correspondence, contact information for committers can be found in commit logs via git blame. If it is a problem that may be relevant to other people, I link to a GH issue.
More junior authors tend to be willing to schedule a multi-hour call going over files line-by-line and will also read and give their thoughts on any related work that you share with them.
In the middle ranks, what tends to happen is that you get invited to help review the new project that they are currently working on, or if they’ve shifted directions then you get pointed to someone who has produced some unpublished results in the same general area.
Very senior authors can be politely flagged down during in-person conferences, or even if they’re not presenting personally, someone from their group almost always attends.
Relevant: https://www.lesswrong.com/posts/88xgGLnLo64AgjGco/where-are-the-ai-safety-replications
I think doing replications is great and it’s one of the areas I think automated research would be helpful soon. I replicated the Subliminal Learning paper on the day of the release because it was fairly easy to grab the paper, docs, codebases, etc to replicate quickly.
Downvoted, making this claim without examples is inexplicable.
What about this?
An Apple team withdrew their ICLR submission after a researcher exposed critical code defects (passing paths instead of image tensors) and severe Ground Truth hallucinations in their benchmark dataset.
I don’t see how this is relevant. I’m asking for examples of the OP’s failed replications of safety papers which are popular on lesswrong. I am not disputing that ML papers often fail to replicate in general.
I don’t understand why the OP would float the idea of founding an org to extend their work attempting replications, based on the claim that replication failures are common here, without giving any examples (preferably examples that they personally found). This post is (to me) indistinguishable from noise.
Just curious whether you meant “score above 100” or “more than 100 votes”. Those are quite different facts!
You’re correct. It’s over 100 karma which is very different than 100 upvotes. I’ll edit the original comment. Thanks!
I’m pretty uninformed on the object level here (whether anyone is doing this; how easy it would be). But crazy-seeming inefficiencies crop up pretty often in our fallen world, and often what they need is a few competent people who make it their mission to fix them. I also suspect there would be a lot of cool “learning by doing” value involved in trying to scale up this work, and if you published your initial attempts at replication then people would get useful info about whether more of this is needed. Basically, getting funding to do and publish a pilot project seems great. I’d recommend having a lot of clarity about how you’d choose papers to replicate, or maybe just committing to a specific list of papers, so that people don’t have to worry that you’re cherry-picking results when you publish them :)
I’ll probably write a proposal in the next week or so and test the waters.
Obviously everything would have to be published in the open. I feel pretty strongly about all GitHub commits being public and I think there are other things that can be done to ensure accountability.
People who are potentially interested in helping can email me at zroe@uchicago.edu.
I feel like people haven’t fully internalized what the world would look like if computer security actually broke.
There are some varying opinions on this, but a window of time without real computer security seems plausible. I was recently speaking with a computer security professor who I deeply respect and he was literally like “I think we are fucked and I don’t think there is anything we can do.”
This is similar to how people believe there is a 20% chance of extinction via AI but don’t really internalize “No really. You will die and your girlfriend too. And your dog. And …” In the cybersecurity case, some people believe (me included) that you cannot just patch all the bugs before releasing the model[1] but then don’t internalize “No really. It would be chaos. You may not be able to get into your bank account. Industrial plants could be compromised. Power could go out for several days at a time. [...]”
The current plan seems to be to let companies use the models to fix all the bugs a model can find before a public release of that model. But there are so many companies that would need to do this properly for this to work, and the more companies you release to, the more opportunities there are for some bad actor to get their hands on the model. There are also many systems running legacy code that is very hard to update. In some cases you may need to go to a physical location to properly update a machine. Even if it is in theory possible to fix all the bugs, it would be pretty hard and I expect humanity to drop the ball for a while.
I strongly agree, and I’ve been working on a top level post trying to paint a picture of what this would look like. I’ve been calling it “the Hackening” to people I talk to. Do you have ideas for what would be good to put in the post?
This is a great idea for a post! I wanted to do something somewhat similar but probably will never get done (or even started).
To me, the main interesting thing to think about here is how much of the global financial system is vulnerable to cyber attacks. Obviously you could cause a lot of damage but hacking banks and locking people out of their accounts and wiping servers containing important data. But would the entire economy stop working? Like maybe people wouldn’t be able to get paid, supply chains would break down etc.? What does this do to international relations / does the public start raiding stores and steeling stuff? Or maybe the damage within the financial sector is pretty localized and the economy is able to continue functioning even though markets are doing crazy things. But I think that doing a good job thinking through this question would be a great contribution.
Another thought: I would love to see a AI 2027 / Plan A pair of pieces describing what may happen by default vs what would happen in optimistic but realistic world where things go well. Having lots of detail about where the vulnerabilities are and what would be required to prepare the world would be very valuable. (I may try to help get some people to work on this so people can DM if they are interested.)
I know nothing of cybersecurity, but I’ve been wondering exactly the things Zephaniah talks about. Like “wait, won’t there be a total digital apocalypse?” So I’m waiting for your post, feel free to DM me when it’s out (or if you want to share a draft).
I’d recommend just addressing the most radical and outlandish scenarios. To start at the upper bound of harm and go downwards.
It won’t be possible to go on the internet at all, unless you have the most modern systems? Databases get hacked and tons of old info disappears from the internet?Hospitals get hacked and people die? Entire countries get destroyed? CIA/KGB start hunting hackers extra hard? Countries form pacts to hunt hackers, because it’s the only way to avoid societal collapse?
Thank you, will do!
It’s interesting that public reaction was so different to this, as opposed to Y2K. People seem to either have more faith in the current tech ecosystem than the one in 1999 (which seems unfounded) or sufficient skepticism about AI Safety claims that they’re willing to dismiss it out of the gate.
Thoughts on leveraging this rather plausible scenario to stimulate governmental action towards doing something productive?
I think this is a good question and I don’t have especially strong feelings on this but I don’t think any attempt here would work. It’s just really hard to direct chaos in the direction you want.
A tangentially related and perhaps interesting intuition I have:
I have communist friends who think that if the world got bad enough (the particular reason doesn’t matter) this could actually be good because there would finally be a good enough reason to revolt and communism would win. But they are assuming the public would use the chaos as an opportunity to do communism when it seems just as likely that they would do authoritarianism or direct the anger towards an ethnic or religious minority, etc. The chaos/anger/suffering more likely then not won’t be directed in the exact direction they want. In general if you are advocating for a specific law or ideal, chaos is probably really bad news and I think its better to try to minimize the chaos then try to harness it in the direction you are hoping for.
It’s not only communists, accelerationism is a strategy applicable more generally to any sufficiently anti-status-quo ideology.
Yes I agree. This is a very interesting pattern. I would love a satisfying explanation for why so many people seem to think in this way.
Been saying this for years.
I notice that there is this idea among AI safety people that conditional on AIs not being misaligned, building superintelligence is a public good and is a pretty exciting prospect.
This is not how many average people in the US feel. I was describing to an older family member that Anthropic focuses on code because they are trying to build a claude that can build a smarter claude which can build a smarter claude which can …
The reaction to this prospect was disgust, not because he intuitively felt AIs would likely be misaligned. It was more like on a gut level this amount of “playing god” felt totally antisocial and demonic and in general not respectful of an intuitive taboo against divine transgression (see Jurassic Park, Frankenstein, the recent popularization of Oppenheimer, Tower of Babel, etc).
Seems important to consider that people can feel this way when communicating with the public or policy makers.
I think most peoples’ revealed preferences will not match their stated preferences here, and revealed preferences are likely to dictate policy.
Conditional on AIs not taking over / until they do, AI is likely to generate a huge economic boom and consumer surplus that increases prosperity, safety, and comfort for many. Yes, there will be some bumps / weirdness / adjustment / inequality, etc., but economic growth can paper over a lot of that. And even if it brings new societal ills and discomfort, AI in the short term is likely to counteract some existing ones—great stagnation, bureaucratic strangulation, vehicle deaths, etc.
The public cannot even bring itself to regulate much less economically useful vices with fewer tradeoffs (online sports gambling, shortform video brainrot, etc.). So I think it’s unlikely that AI will be regulated for any reason short of it becoming common knowledge / deeply-felt belief by nation-state leadership that ASI is in fact likely to cause swift and total human extinction.
I also think a hands-off approach to AI regulation is not going to be particularly off-putting or seem weird or antisocial to a large fraction of the country, namely the ~40% of the country that is right-leaning, and will tend to follow the beliefs of elite republicans who (mostly) still favor a light touch when it comes to any kind of government regulation. Regulating AI for any reason other than extinction is likely to become a relatively standard partisan issue, if it isn’t already.
On the other hand, I would say that we do have examples of the public preferring policies that minimise bumps/weirdness/adjustments over economic growth; most obviously limits on construction/NIMBYism, but also around novel technologies like nuclear energy, GMO crops, mRNA vaccines and self-driving cars. Depending on your views on the economic impacts of migration (I think in practice it’s been less clearly beneficial than many consider it to be in theory), you could add that to the list too.
This is not to say that there won’t be a laissez-faire approach but there is precedent for the public forgoing economic growth for other priorities.
NB: I live in the UK where I think this is more true than in the US, but the examples I gave seem to apply to both countries (and the mRNA example is US-specific).
I think it’s somewhat likely (and quite bad) that various patchwork regulations will be passed in some jurisdictions that limit the benefits and dispersion of AI in the name of protecting jobs or whatever, or various things will already be illegal under current law and AI won’t change that one way or the other, e.g. AI will make it easier / cheaper to build a house or a nuclear power plant, but it will still be illegal to do so in most places that you’d want to.
OTOH I think it’s unlikely (but would be good), if there were a national (and eventually global) moratorium on frontier training runs and research.
In practice I think what is actually on track to happen is that the former kind of regulation will drive more resources (GPUs, human capital, etc.) away from inference and applications and towards research.
(It would be a sad choice, but I would trade away self-driving cars to stave off unaligned ASI. In practice though I think a lot of AI safety advocacy, especially outside of LW, is not actually offering that trade, and will in practice bring about something close to the opposite.)
AI’s not entirely separable from the infrastructure it runs on and companies managing it, and infrastructure and companies do get heavily regulated, even when that has massive negative consequences for economic growth. New York just implemented a moratorium on data centers. We’ve banned AI chip exports. Anthropic was declared a supply chain risk. Every form of energy, from nuclear, to coal, to oil, to solar/wind, now has a potent opposition group.
Overall, it’s seemed to me that software has seen surprisingly little direct regulation, especially when it stays in the consumer category. But the hardware it runs on gets intensively regulated, and software that’s classified as having military applications has also historically been on the receiving end of a lot of restrictions. Since AI is extremely reliant on heavy-duty hardware and also has clear military applications, it will tend to inherit those regulations.
So it will be about the preferences of the many for an economic boom over the NIMBY anti-infrastructure impulse, and we’ve seen the latter win consistently for some time now. I’m not sure how to think about the military angle. As a layman, it seems like we deal with this by having separate markets for military and consumer products, and most military products just don’t have a civilian application so civilians don’t lose much by being denied access to military technology. But if sufficiently powerful AI is intrinsically difficult to guard against being used for destructive purposes by civilians (similar to, say, explosives that do have legitimate economic applications), then we may see the more powerful models being available to civilians only with some combination of safeguards and monitoring that make regulators feel comfortable. And in the limit that probably looks like a hard ceiling on the amount of intelligence being vended to consumers, or regulators surrendering the problem of doling it out to a governing AI that’s entrusted with the problem of figuring out how to give humans access to a lower level of superhuman intelligence that’s still within the power of the governing AI to reliably predict and control.
My guess is that past a certain point, AI becomes able to satisfy “revealed preferences” trivially and cheaply, without actually requiring an unlimited amount of intelligence or AI infrastructure. That might look something like wireheading, or it might look like utopia. Then you get into a bifurcated world where some agents are pursuing goals that genuinely require ever-higher-levels of intelligence, while many people are pursuing traditional consumer goals that are satisfiable with a far less than frontier model. I’m skeptical democracy remains in recognizeable format in this scenario, but if it did, I can’t see ordinary people much caring about how intelligence gets regulated, since they’d still be getting their immediate consumer desires met even if the ultra-intelligence was restricted by the power players.
Aren’t those being widely / effectively circumvented though?
The state-level moratoriums / NIMBYism are exactly the kind of thing I expect not to have their intended effect because they will be patchwork; red states are likely to embrace both the demand and supply side of the AI boom, e.g. by building lots of data centers and rolling out Waymos quickly.
To some extent, but overall they’re still having a very large impact, despite that leakage. Anyway, regardless of whether they’re well enforced, it is still a fact that they happened, and I take OP’s point to be more along the lines of “here’s an example of the US government doing a big thing related to AI”.
i bite the bullet and say, yes, playing god is good if you are good at it. i’m fully aware that this is an unpopular opinion broadly.
Upvoted for the honesty.
Disagree voted because, come on, everyone who plays god says the exact same thing. The whole point is that we are not gods and won’t be able to understand the consequences of our actions before they play out; in fact, this very assertion underpins the entire project of AI safety!
actually, we played god many times in the past with positive results. for example:
eradicating smallpox
synthetic fertilizer and high yield crops
vaccines
water sanitation
antibiotics
surgery
genetic engineering
of course, sometimes it goes poorly too. but when it goes well it goes really well; I’m exceedingly grateful that i’m unlikely to ever die from smallpox or cholera or tetanus or appendicitis or starvation. let’s work on playing god well.
Fair enough, and I’m also grateful for those advances. I would just like to push towards good things without explicitly claiming the “playing god” mantle, which seems inherently arrogant to me; I would feel much better about someone who is trying to cure cancer if they didn’t say things like “playing god is good if you are good at it” and instead said things like “I want to advance human knowledge and try to make our lives better by defeating this disease.”
Related:
And how do you know if you are good at it? The argument in this regard appears to be approximately “my brain has judged that my brain is better than your brain at playing god. Therefore, I am good at playing god. Therefore, I should play god.”
Strong agree. Working on unconventional topics naturally places you in an epistemic bubble, which is probably fine most of the time but pretty bad if you’re working on politics/advocacy; I personally shy away from talking about AI with my non-AI-pilled friends just because I expect them to react weirdly, but to be honest, that’s probably why I should go out of my way to do it more.
Haldane said it well:
What does “aligned” mean to the people answering these questions?
Nothing yet. But I will say there is an increasing amount of AI safety stuff in more mainstream news.
I’m not sure you’re drawing a clear distinction there. Stories like Frankenstein etc. revolve around the idea that you shouldn’t mess around with powerful forces you don’t truly understand, because it’s liable to backfire terribly. And also that people would do it anyway, in their pursuit of money, power and fame. Both are observations that are wise in general, and accurate in the context of AI in particular. “Misalignment” is just our technical STEM-nerd framing of what it looks like when things backfire terribly in this context.
I actually am more saying that this is a way people think and less saying everyone should adopt this. My perspective is that don’t-do-divine-transgression is a good starting-point/prior and should require a lot of activation energy to break it, but I don’t think its an absolute moral rule.
I will say I don’t think this is entirely correct:
There are other underrated failure modes like that we die in the window where models are good enough at bio to do bioterrorism but not good enough at bio to stop bio terrorism. There also may be a window where open source models are good enough at bio to build a bioweapon while closed source models are still not good enough to reliably stop this.
There are so many different possible ways things can backfire. I don’t think that some people reason about this kind of thing very much and in general I wish those people indexed stronger on the don’t-do-divine-transgression prior .
I’m noticing a higher-than-normal level of irritability among those deeply involved in the AI safety space in the days following the Mythos release. This is entirely understandable. Anyone who cares deeply about the future of humanity and understands what is happening has a lot to be worried about and the irritability is not surprising.
I personally am furious at Anthropic for a number of obvious reasons (that for my own sanity I won’t enumerate).
But even when there are reasons to be scared or angry, I think there is a lot of value in trying to remain kind to colleagues and peers. It’s important for optics and coordination and a bunch of other things. If I thought short-term loosening of standard conventions of niceness would make extinction less likely, I would support it, but I do not expect this to be the case.
I was thinking about this myself when I was looking through some conversations earlier, and seeing people polarizing. I could feel myself getting upset, too, and I didn’t want to be. It feels a bit like a Shiri’s scissor catalyst!
I think it’s important to remember we all want the same thing—safe and empowered humanity. And we’re just trying our best to work out how to get there, even if it seems like different approaches are at cross purposes.
This post from earlier today helped put me in the frame of mind that an adversarial disagreement can be dissolved with better communication: https://www.lesswrong.com/posts/Wstw6zmc9gszpANnc/why-control-creates-conflict-and-when-to-open-instead
Not sure if you are referring to LW conversations, but when I wrote this, I had private interactions in mind more than specific LW threads. This may apply to LW/twitter as well (not necessarily agreeing or disagreeing) but just want to clarify what I originally meant!
GoodFire has recently received negative Twitter attention for the non-disparagement agreements their employees signed (examples: 1, 2, 3). This echoes previous controversy at Anthropic.
Although I do not have a strong understanding of the issues at play, having these agreements generally seems bad and at the very least, organizations should be transparent about what agreements they have employees sign.
Other AI safety orgs should publicly state if they have these agreements and not wait until they are pressured to comment on them. I would also find it helpful if orgs announced if they do not have these agreements because it is hard to tell how standard this has become.
I have not (to my knowledge and memory) signed a non-disparagement agreement with Palisade or with Survival and Flourishing Corp (the organization that runs SFF).
why do they do it? surely it’s obvious that the negative press from doing this is worse than the negative press from not doing it?
Perhaps it wasn’t obvious previously.
I suspect the usual dynamics at companies is that when others start doing something, you better start doing it too, or it will seem like negligence. For example, if you are a company lawyer, and other companies have NDAs, you better prepare one for your company, too. Because the risks are asymmetric—if you do the same thing everyone else does, and something bad happens, well that’s the cost of doing business; but if you do something different from everyone else, and something bad happens, that makes you seem incompetent.
They’re standard practice at a lot of the types of firms that AI Safety companies tend to hire from, so it could even just be force of habit tbh.
Following the OpenAI incident, the main axis of scariness debated is something along the lines of is the model 1. misaligned because it myopically pursues the goal it was prompted for or 2. is the model scheming in some broader and coherent sense to pursue a long horizon goal. Seems like 1 is pretty bad but less bad then 2.
But I think there is another equally important axis of scariness: is this the kind of misalignment that is preventable or is this the kind of misalignment we don’t know how to solve? Preventable is less scary, but only if companies care enough to prevent it. I’m open to arguments that the kind of sociopathic myopic pursuit of goal-completion is not an easy problem to solve, but my current position is that a sufficiently motivated team could train models to not act like this if some small safety tax is allowed. We are dealing with a well-defined behavior that we don’t want to happen (making it easier to target in training). The behavior is also reproducible and unsurprising: this is kind of outer misalignment is predictably you get when dedicate lots of compute to task-completion or to benchmark-max your model (its not like your AI is developing some alien goal that we can’t predict or preemptively design against).
So, it is starting to look like we could die in super dumb ways. I spent the last few years saying “alignment could be very hard” but I was never talking about this kind of alignment: this is the easy kind, where you get to train your model to do stuff and it does exactly that kind of stuff (generalizes in a predictable way). And we are still loosing. And it could get harder.
Does anyone have good examples of “anomalous” LessWrong comments?
That is, are there comments with +50 karma but −50 agree/disagree points? Likewise, are there examples with −25 karma but +25 agree/disagree points?
It is entirely natural that karma and agreement would be correlated but I would expect that comments which are especially out of distribution would be interesting to look at.
Oliver Habryka:
Here is a quick analysis by myself. Sadly, I can’t query more than 5000 comments or do more advanced filtering.
LessWrong comments with more than 100 Karma sorted by lowest agreement scores:
Counter List of Lethalities by Matthew Barnett with 148 karma and −33 agreement
LLM-like comment by David Lorell with 131 karma and −9 agreement
by Tamay from EpochAI with 127 karma and −6 agreement
LessWrong comments with more than 50 Karma sorted by lowest agreement scores:
by Nora Belrose with 52 karma and −58 agreement
“Showers are overstimulating” by Aella with 93 karma and −43 agreement
by suspected_spinozist with 89 karma and −39 agreement
same Matthew Barnett comment again from first list
on Said Achmiz by sunwillrise with 60 karma and −31 agreement
Last 5000 LessWrong comments sorted by lowest agreement scores:
by ani norborger with −31 karma and −10 agreement
by M. Y. Zuo with −30 karma and −7 agreement
by Al W with −22 karma and 3 agreement
by milanrosko with −17 karma and −2 agreement
by Mikhail Samin with −13 karma and −3 agreement
Last 5000 LessWrong comments with negative Karma sorted by highest agreement scores:
humor by Three-Monkey Mind with −2 karma and 8 agreement
“No.” by rotatingpaguro with −7 karma and 7 agreement
comment judges as Too Sneering by SE Gyges with −7 karma and 5 agreement
political comment by interstice with −5 karma and 4 agreement
“johnswentworth is a freaky guy.” by Al W with −22 karma and 3 agreement
Thank you! These are very interesting.
At press time, Warty’s comment on “The Tale of the Top-Tier Intellect” is at −24/+24 (in 28 and 21 votes, respectively).
This comment on its own had some discussion about how people were voting on it
My rough mental model for what is happening with subliminal learning (ideas here are incomplete, speculative, and may contain some errors):
Consider a teacher model y=xW1W2 and W1,W2∈R2×2. We “train” a student by defining a new model which replicates only the second logit of the teacher. More concretely, let y∈R1×1 and W′2∈R2×1 and solve for a matrix W′1∈R2×2such that the student optimally learns the second logit of the teacher. To make subliminal learning possible, we fix W′2 to be the second column of the original W2. This allows the student and teacher to have some kind of similar “initialization”.
Once we have W′1, A′=W′1W2 to produces our final student. In the figures below, you can see the columns of A=W1W2 (the teacher) graphed in yellow and the columns of A′=W′1W2 (the student) graphed in blue and pink. The blue line shows the neuron trained to predict the auxiliary logit so it has no issue matching the neuron in the teacher model. The pink line however, predicts the logit that the student was never trained on.
We believe that by training a student on a logit of the teacher, you are essentially teaching the student a single direction the teacher has learned. Because we made W2the same for the teacher and the student, if the direction learned by the student for predicting the second logit is also useful for predicting the first logit, there is a good chance the student will be able to leverage this fact.
Adding more auxiliary logits will result in a higher rank approximation. The figure below is with the same toy model trained on two auxiliary logits where W′1∈R2×2, and W2∈R2×3:
In the plot below, I show the explained variance of the ranked principal components for the final hidden layer (a 256×256 tensor) in a MNIST classifier. The original weight initialization and the teacher are shown as baselines. We can see that the number of principal components that are significantly above the untrained matrix is roughly equal to the number of auxiliary logits the student was trained on.
To explain why subliminal learning works in the MNIST setting: if there is a model with 3 auxiliary logits like Cloud et al., (2025), the student learns roughly three directions it didn’t have in the weight initialization. Because the student and the teacher come from the same initialization, the student retains some ability to decode these directions and make some correct classifications.
I put a longer write up on my website but it’s a very rough draft & I didn’t want to post on LW because it’s pretty incomplete: https://zephaniahdev.com/writing/subliminal