That seems like a very nonstandard use of the term “Constellation”
maxnadeau
(This is a comment is full of disagreements and I hope you can imagine me saying it to you in the friendly, inquisitive tone I’d have if we were chatting in person, rather than in a cranky and angry one)
A bunch of your claims here are confusing to confusing to me, are difficult for me to parse, or seem wrong to me. I think you have a pretty different way of dividing up the worldviews/perspectives of the AI safety community from me, and I don’t feel like I understand what you’re pointing at when you make reference to certain clusters and fault lines.
I suspect what’s going on here is that you’re using the term “Constellation” to refer almost entirely to Redwood. Personally, I have been heavily influenced by Paul and Ajeya’s threat modeling work (e.g. WFLL, AOAFS, WSCTEP...) and the conversations I’ve had with them, most of which took place within Constellation. And of course, they themselves have been known to frequent Constellation over the years. So I’m perplexed by the way your post contrasts them with Constellation and says their views were “mostly just missing from Constellation dialogue”. And I don’t think I’m unique in how much I’ve absorbed their views/thoughts; you may have been more influenced by your employer’s threat modeling work than Paul/Ajeya’s, but (as one noisy proxy) look how much less karma Redwood’s threat modeling posts have!
You write that “the discourse has focused somewhat too much on these more contingent stories for alignment risk and not enough on the basic “Goodharting” argument… the Constellation view looks somewhat too focused on non-central inductive-bias questions around scheming”. I realize that Redwood has done a lot of work predicated on the schemer threat model over the last few years, but I am not convinced that there is an important fault line between people more focused on Goodharting and people more focused on inductive biases, such that MIRI falls on the former side. After all, the idea that inductive biases towards optimizers may produce misalignment even in the absence of Goodhart (and/or that Goodhart has a multiplier effect on such biases) originates in MIRI’s writing and was first thoroughly explicated in MIRI’s paper, RFLO. Meanwhile, much work has been done in Constellation and by Constellation members on the topics of scalable oversight, ELK, and other attempts to mitigate Goodharting/reward-hacking. If anything, I think you have it backward which cluster is focused on which sources/varieties of misalignment.
You write that “it was importantly not necessary for risk that the AI only succeeded in training as a strategy for gaining beyond-episode power”. This, to me, looks like a strawman—I don’t think anyone had this view or acted as if they did. Perhaps you think various people’s work overindexed on scheming threat models, but that’s a different claim—everyone I know always had the view that other forms of misalignment (e.g. Hackistan!) were plausible too.
You write that Constellation underappreciated the issue of “Problem → Patch → Make it smarter → Whole New Problem → New Patch”. Following Nate’s post, I am happy to give MIRI credit for originating/popularizing this idea, and I agree it is core to what’s scary about alignment. But from my perspective it is suffused thoroughly and ubiquitously throughout the thoughts and writings of people in Constellation, especially those that were most influential (within and without Constellation). See for instance Holden’s writing, along with the posts of Paul’s and Ajeya’s mentioned earlier.
You write that “We rarely talked about the alignment problem in depth at Constellation, and should have fostered more of a culture of discussing the basic arguments… Constellation pedagogy foregrounds particular types of misalignment like “schemers” and “reward-seekers” too much, rather than the basic problem of Goodharting on outcomes”. This has not been my experience! Now, I expect some readers to object that they have spent plenty of time in both Constellation and in other hubs of AI safety discussion and observed much more discussion of the alignment problem, “the basic arguments”, and “the basic problem of Goodharting” in those other spaces. I’m sure that’s true. But there’s a big difference between the claims “People in Constellation don’t talk about alignment enough (implicitly: and they talk too much about other aspects of the AI safety situation)” and “People in Constellation rarely talk about the alignment problem in depth”.
You write “We (mainly Redwood) didn’t focus enough on superintelligence, and put too much emphasis on early schemers”. I think the jury is still out on this one, and nothing about recent events seems like a decisive update in this direction.
To round out this very disagreeable comment (again, which I intend to be friendly!), here is a major point of agreement: I agree that kludges of motivations and narrow/sphexish fitness-producing habits/instincts have been underrated historically. I appreciated this observation from Nate Soares, for instance, and think it’s a good corrective to me/others putting too much credence on simpler-to-describe motivational complexes. But I think the main reason these things were underemphasized was that people like to talk about superintelligences that have undergone processes towards coherence, which conflicts with some of the other updates you recommend in this post.
Here’s a brief attempt to ground/concretize people’s understanding of how much funding CG provides for “more ambitious” AI safety research: I can quickly think of around $40m worth of grants we’ve made over the last year (a lot in the last few months, so not all disbursed/posted yet) that meet the following criteria (and probably there are more I didn’t think of):
Technical research
Done by people trying to reduce misalignment risk and who are very familiar with the LW literature
Who are using mathematical methods to strive for higher-assurance safety
A few preemptive clarifications:
I don’t think this settles the question of “to what extent are funders setting incentives that pull against doing more ambitious research?”, but I think it provides some info/”upper bound” about how strong those incentives could possibly be. $40m/year is a lot more than $0—it’s more than the SFF spends, and more than we spent on our two “flagship” RFPs in 2023/2024.
I’m not saying it’s “enough”—CG would like to spend more money on more ambitious work, and I am trying to make it happen.
I think my criteria above are a reasonable operationalization of “more ambitious” work, but I don’t think they’re the only one.
While the Chain-of-Thought (CoT) for OpenAI models reasonably accurately reflects the model’s cognition, the CoT for Anthropic models does so to a substantially lesser extent. This may be due to “spillover” effects where reinforcement on outputs transfers to the CoT because Anthropic’s CoT is less distinct from the output
Could primary main cause instead be (as Tim Hua notes) that Anthropic has trained past models to have reward-model-pleasing chains of thought?
“When something quite bad can occur in a small number of actions (potentially even in a single trajectory) or in a short period of time, it’s important to sometimes block suspicious looking actions until they are reviewed.”
Why? I think there is not perfect overlap between the categories “single actions that are catastrophic, i.e. high stakes failures” and “actions that require monitoring and blocking in the moment”. One example where they come apart is “producing a batch of data that would cause the next generation of models to be misaligned”—this might be a high-stakes failure, but you have plenty of time to catch it before you actually train that model. Another example would be “writing code that is catastrophically insecure but only gets run rarely”.
What are your objections to Alex/Buck’s definition, for example? I think they define it pretty well. Of course, their definition admits some edge cases and ambiguities, but so do lots of concepts—still seems like useful concept to me.
I’m glad you’re betting on your own taste/expertise instead of donating on behalf of the community—that seems sensible to me.
(I am also a grantmaker at Coefficient/OP)
The arguments/evidence in this post seem true and underrated to me, and I think more people should come work with us.
In particular, I also have updated upward on how impactful the job is over the last year. It does really seem to me like each grantmaker enables a ton of good projects. Here’s an attempt to make more concrete how much is enabled by additional grantmakers: If Jake hadn’t joined OP, I think we would our interp/theory grants would have been fewer in number and less impactful, because I don’t know those areas nearly as well as Jake does. Jake’s superior knowledge improves our grantmaking in these areas in multiple ways:
Better sourcing: Jake’s involvement meant that the proposals in these areas that were even available to us to evaluate were much better. His contributions to the interp/theory sections of the RFP meant the incoming proposals were higher-quality than if I had attempted to write them, and he had good suggestions/steers for grant applicants that I couldn’t have offered. He was also able to proactively ideate and realize projects that I wouldn’t have thought of or wouldn’t have had time for.
More grants, in more varied subareas: because Jake knows those areas better, he can evaluate proposals faster and is more comfortable arguing for/defending these grants than I am. This allows us to make more, and more varied, grants in those areas.
[The obvious one] Jake has better discernment among proposals in these areas than I do, which straightforwardly increases the impact of our grantmaking.
I think there are probably more buckets of similar scale/impact grantmaking to interp and theory that we’re currently neglecting. We need to hire more people to open up these new vistas of TAIS grantmaking, each of which will contain not just mediocre/marginal grants, but also some real gems! I think this dynamic is often underappreciated; additional grantmakers take ownership for new areas, rather than just helping us make better choices on the margin.
I also think that Jake obviously had way more impact on theory/interp than if he had done direct work. He funded dozens of projects by capable researchers, many of whom wouldn’t have worked on AI safety otherwise. I think most TAIS researchers aren’t taking this nearly seriously enough, and I think the case for grantmaking roles looks very strong in light of this.
I fervently agree that the degree of inaction here is embarrassing and indefensible.
Here’s one proposed explanation, though definitely not a justification, of why the Biden admin didn’t do more. To be clear, I’m not saying they’re the only ones who should have done/do more:
The shrinking anti-pandemic agenda
After coming up with a $65 billion moonshot plan, Biden asked for about half of that as part of his initial Build Back Better proposal. But as the entirety of BBB shrank in an effort to secure the support of Joe Manchin and Kyrsten Sinema, the pandemic prevention shrank to about $2.7 billion, of which roughly half is to modernize the CDC’s labs.
And it’s far from clear that even this relatively small amount will pass.
The extreme shrinkage of the pandemic prevention agenda in part reflects a partisan calculation. To Democrats who agree that this should be a priority, it doesn’t feel like it’s a distinctively progressive priority that should squeeze out ideas like free preschool or Medicaid expansion, which everyone understands Republicans oppose. They feel like this bill is supposed to be dessert, and pandemic prevention is vegetables.
And the good news is that it’s true — pandemic prevention is not a super partisan topic, and there are prospects for bipartisan cooperation.
The problem is that once you get into the regular appropriations process, the logic of base rates starts to dominate everything. To secure a 30% increase in pandemic preparedness funding would be a big step for appropriators since obviously most programs can’t score increases nearly that large. But we are currently spending peanuts on a problem that has both massive economic consequences and carries genuine existential risk. We don’t need a large increase in pandemic preparedness funding; we need to go from “not seriously investing in preventing pandemics” to “genuinely trying to prevent pandemics” with a gargantuan investment in funds.
What’s particularly galling is that even as the need exceeds the demand of standard appropriations, it’s still relatively modest compared to the $725 billion defense budget for fiscal year 2020. And due to base rate issues, the Biden administration’s 2021 requested increase — though modest in percentage terms — still amounts to $12 billion for one year in a world where asking for a $7 billion per year increase in defending ourselves against pandemics is considered outrageous.
Thanks for writing this! Just booked a time in your calendly to discuss at more length.
For more discussion of the hard cases of exploration hacking, readers should see the comments of this post.
Typo: should be “Gell-Mann”
I figured out the encoding, but I expressed the algorithm for computing the decoding in different language from you. My algorithm produces equivalent outputs but is substantially uglier. I wanted to leave a note here in case anyone else had the same solution.
Alt phrasing of the solution:
Each expression (i.e. a well-formed string of brackets) has a “degree”, which is defined as the the number of well-formed chunks that the encoding can be broken up into. Some examples: [], [[]], and [-[][]] have degree one, [][], -[][[]], and [][[][]] have degree two, etc.
Here’s a special case: the empty string maps to 0, i.e. decode(“”) = 0
When an encoding has degree one, you take off the outer brackets and do 2^decode(the enclosed expression), defined recursively. So decode([]) = 2^decode(“”) = 2^0 = 1, decode([[]]) = 2^decode([]) = 2, etc.
Negation works as normal. So decode([-[]]) = 2^decode(-[]) = 2^(-decode([])) = 2^(-1) = 1⁄2
So now all we have to deal with is expressions with degree >1.
When an expression has degree >1, you compute its decoding as the product of decoding of the first subexpression and inc(decode(everything after the first subexpression)). I will define the “inc” function shortly.
So decode([][[]]) = decode([]) * inc(decode([[]]) = 1 * inc(2)
decode([[[]]][][[]]) = decode([[[]]]) * inc(decode([][[]])) = 4 * inc(decode([]) * inc(decode([[]]))) = 4 * inc(1 * inc(2))
What is inc()? inc() is a function that computes a prime factorization of a number and the increments (from one prime to the next) all the prime bases. So inc(10) = inc(2 * 5) = 3 * 7 = 21, and inc(36) = inc(2^2 * 3^2) = 3^2 * 5^2 = 225. But inc() doesn’t just take in integers, it can take in any number representable as a product of primes raised to powers. So inc(2^(1/2) * 3^(-1)) = 3^(1/2) * 5^(-1) = sqrt(3)/5. I asked the language models whether there’s a standard name for the set of numbers definable in this way, and they didn’t have ideas.
On point 6, “Humanity can survive an unaligned superintelligence”: In this section, I initially took you to be making a somewhat narrow point about humanity’s safety if we develop aligned superintelligence and humanity + the aligned superintelligence has enough resources to out-innovate and out-prepare a misaligned superintelligence. But I can’t tell if you think this conditional will be true, i.e. whether you think the existential risk to humanity from AI is low due to this argument. I infer from this tweet of yours that AI “kill[ing] us all” is not among your biggest fears about AI, which suggests to me that you expect the conditional to be true—am I interpreting you correctly?
To make a clarifying point (which will perhaps benefit other readers): you’re using the term “scheming” in a different sense from how Joe’s report or Ryan’s writing uses the term, right?
I assume your usage is in keeping with your paper here, which is definitely different from those other two writers’ usages. In particular, you use the term “scheming” to refer to a much broader set of failure modes. In fact, I think you’re using the term synonymously with Joe’s “alignment-faking”—is that right?
People interested in working on these sorts of problems should consider applying to Open Phil’s request for proposals: https://www.openphilanthropy.org/request-for-proposals-technical-ai-safety-research/
This section of our RFP has some other related work you might want to include, e.g. Orgad et al.
I don’t think I had that misunderstanding exactly. I am disputing your characterization of the intellectual environment that “many others” experience from being at Constellation lunches, talks, and on slack (but perhaps this is a good characterization of what you experienced, or of being a junior Redwood employee, I’m not sure).
I don’t agree that the lunch/talks/slack have the biases you refer to (i.e. inductive biases arguments for misalignment over Goodhart, the belief that terminal training-gaming poses no risk, ignoring misalignment emerging/spreading during deployment, ignoring the limitations of band-aid solutions to alignment). I don’t see much evidence that there has been tons of discussion of inductive biases arguments for misalignment in the slack/talks/lunch conversations.
I now agree with your position that there is rare discussion of the alignment problem in lunch/talks/slack. I was largely thinking about my early years in Constellation, in which I talked about that subjects very frequently.