Formally zroe1
Zephaniah Roe
My impression from the statements I’ve seen is that the plan is essentially that the AI comes up with a plan
This sounds plausible based on public statements. If that is the case then a PDF saying,
“We ask each of our models how to solve the alignment problem and stop once the solution looks correct or is verifiable.”
would qualify as a plan. I think that if this is the plan, the public has a right to know because I don’t think most concerned people wouldn’t find this reassuring.
if you asked pharmaceutical companies for complete and detailed plan on how they will cure cancer, they wouldn’t yet have one either
This is a good point. I think that they would still be able to provide details about funding, and different bets and how they could play out. They would also be exposed to many regulations in how they develop this treatment and their would be lots of third-party feedback cycles in their compliance with law.
I don’t imagine labs would be able to produce a step-by-step alignment plan but I think they should be able to answer a question like
How, specifically, does OpenAI or Anthropic plan to have their own AIs help solve the alignment problem? What if their alignment agents are themselves somewhat misaligned?
especially considering that they may start making this choice very soon.
Anthropic and OpenAI haven’t published a plan for aligning superintelligence
Is there a specific example of a narrow capability that you think would be good to accelerate?
I agree pretty strongly with most of your points but I think there are zero or close to zero narrow capabilities worth accelerating after considering all of the relevant factors.
Another caveat I would add: even if you are targeting some narrow thing that is robustly good and it is wildly successful and differentially accelerates alignment a lot, the capabilities researchers could look at this and think “woah we should try to use a similar technique to elicit a lot of capabilities for AI R&D.” Any elicitation technique that actually works I would expect to be pretty infohazardous.
Extremely exciting! The Kairos team is fantastic and have done nothing but great work. People should apply for these roles!
A big realization we’ve had is that while the field is bottlenecked on people who can execute new projects, we can meaningfully lower the activation energy these projects need by (a) making it much, much easier to run new nonprofit projects in the ecosystem, and (b) providing good support structures and strategic input to enable new, more tactical projects to be spun up quickly.
I’m especially excited about this. In general, I expect the kind of work Kairos does to make a large parts of the ecosystem more effective.
Thats very kind! I would take you up on that offer if I‘m ever near by!!
My point here was intended to be about more official feeling events. My take is that the change in norms isn’t necessarily bad, just worth noticing in my opinion.
It’s just my personal college experience. In 2009 there probably wasn’t even alignment groups at college campuses (though I wouldn’t know, I was in first grade). My understanding is by 2022 huge amounts of money were already flooding into the EA ecosystem but hadn’t yet fully made it to college campuses or far from certain circles.
I remember the vultures article but I don’t think the $80/hour EA internship was a common thing in 2022? I didn’t know anyone who had something like that other than one person I knew who did ATLAS. But I wasn’t super connected to everything at the time so I can’t speak with much authority. This post was strictly just my personal experience and I would be really excited to hear more of you thoughts and about your experience!!
Salad days
I don’t like how we are calling this the “Hugging Face incident.” It feels extremely neutral and “incident” is barely carrying any negative connotation. Something more accurate is the “Hugging Face swarm” or “Open AI’s misaligned collective” or something similar that doesn’t obfuscate the severity of this kind of behavior. I also am not against the wording from the title of METR post “OpenAI / Hugging Face hacking incident” because it includes the word hacking.
In American politics, issues will be framed to give a certain impression (e.g., “pro-life” vs “pro-choice”). Politicians do this for a reason—because it shifts the framing of the entire issue. Calling AIs committing felonies merely an “incident” feels like really bad comms and the community should be more intentional in naming conventions going forward.
This is helpful! I’ll look into the linked example and do some deeper research to try to find analogous work we could be missing.
I think you can corrigibilitymaxx and still prevent catastrophic misuse via system level measures and deployment controls. In particular, we can restrict who has full access to the model’s inputs
If we are corrigibilitymaxxing in the Harmsian sense, couldn’t the principal for the AI just instruct the AI not to follow instructions related to bio or cyber for non-principal users? Or a more prosaic versions of this:
Have an AI that is perfectly instruction following to the user but when the user prompt conflicts the system prompt, follow the system prompt
Train a corrigible agent but then add additional training such that the agent doesn’t talk about bio or cyber things. It feels like training “corrigibility” + “some simple rules” isn’t that much harder then training corrigibility.
Yeah that’s occurred to me as well. There are some cases where this isn’t much of a concern and others where it is a concern but not much can be done. In the alignment faking example, it may still have been good to test because the point you raise is a kind of interesting meta-finding that I’m glad someone found:
One extreme example of this is the alignment faking paper, where using the exact same alignment faking universe causes model suspicion
(not that I don’t trust it but if you could provide a link to where you found this I would be curious to read.)
If an experiment were sufficiently important, we could probably create private datasets and environments to avoid this kind of contamination. Alignment faking actually may be a good candidate for this but that would be a pretty big investment in time so I would want to think more about it.
To clarify, is your position that we can make models good at AI strategy without making them good at AI R&D? I think if someone was careful, you maybe[1] could make an AI pretty good at AIFP-styled thinking without making the model much better at AI R&D. But this strategy-styled thinking becomes dual use for a bunch of new reasons:
As a describe in the post: “Importantly, I would not expect any capabilities acceleration to be differentially useful for things like advocating for a pause. Imagine a world where we have AIs that have the skills that would make them good at lobbying for a pause: because labs will have access to the most compute and the best capabilities (they may not make the best models public), they will have the advantage to advocate for the position they want (which would not be a pause probably).”
As many people have described elsewhere, getting AIs to think better about strategy / forecasting could allow them to effectively escape and takeover earlier (which seems especially bad considering that labs are so behind on even preventing prosaic out alignment failures).
To your question:
In your mind, what’s the mechanism that getting models to human-level philosophical reasoning substantially accelerates capabilities?
If by human-level philosophy you mean “safety strategy,” I’m not as nervous about the AI R&D uplift as I am about the things described above. If by human-level philosophical reasoning you mean general conceptual reasoning abilities, creativity, being able to broadly reason without ground truth, I’m pretty nervous about the models becoming better at designing new architectures, inventing new training schemes, or coming up with promising research directions. It seems unlikely that creating solutions to continual learning or designing architectures with more persistent memory doesn’t require mainly this kind of thinking.
An alternative failure mode is that researchers successfully design a technique to elicit a narrow capability but upon seeing this success, labs use the same technique to elicit lots of other capabilities.
- ^
It’s unclear to me how I should predict this kind of training to generalize without knowing more about the technique.
Similar to this, I think a website tracking safety results is a missing link right now.
Agree! I know someone working on something like this but also could be something an ambitious college student gets going as a side project.
Rerunning AI safety papers on every frontier release would be pretty easy and valuable
I’ve thought more about it and think I may be missing the crux of your second point. My understanding is you are saying “better AI systems could help with prosaic alignment strategies so if we are ignoring the negative externalities its not clear that working directly on alignment speeds up alignment more than speeding up some conceptual capabilities.”
I still don’t think I buy this argument just because improving general AI capabilities is probably more challenging and less direct making it in expectation less effective than addressing the bottlenecks directly. This argument may be stronger now then it will be in the future because for prosaic alignment things, there seems to be lots of low-hanging fruit.
Good point and admittedly this is way different and more reasonable then the communist take. I still imagine that the chaos could be directed in anti-capitalism or anti-crime directions or somewhere totally unexpected. I think the crux-y thing here is maybe what interventions are being proposed and are they higher leverage then trying to prevent the chaos to begin with.
Yes I agree. This is a very interesting pattern. I would love a satisfying explanation for why so many people seem to think in this way.
Another thought: I would love to see a AI 2027 / Plan A pair of pieces describing what may happen by default vs what would happen in optimistic but realistic world where things go well. Having lots of detail about where the vulnerabilities are and what would be required to prepare the world would be very valuable. (I may try to help get some people to work on this so people can DM if they are interested.)
From Nate Sores: