Member of Technical Staff at OpenAI.
Previously: MIRI → interp with Adrià and Jason → Head of linear regression at METR
I have signed no contracts or agreements whose existence I cannot mention.
Member of Technical Staff at OpenAI.
Previously: MIRI → interp with Adrià and Jason → Head of linear regression at METR
I have signed no contracts or agreements whose existence I cannot mention.
Out of curiosity, do you have any idea what an alignment solution would look like though?
Something like this was mentioned long ago by Paul Christiano in the original ELK report, as a possible outer alignment target. (See the indirect normativity section on page 62). It solves the issue where we need to delegate to some state to move utopia forward, but need information from other branches too.
LessWrong really needs to allow agents to read LW comments. I paste or screenshot a lesswrong shortform or comment into Codex like twice a day.
Hmm this isn’t quite what I was looking for, I wanted answers that were more concrete about why we think the AI is aligned. How will we know we’ve found the universal ethics? Will the fact that we have the “theory of everything” be load-bearing to alignment? If not, how do we know the AI generalized well? What is the model spec and how do we know the AI satisfies it?
Note I also don’t claim that the solution to alignment will be alignment evals or anything that looks like them at all. Also my median conditional on survival is that we muddle through with just enough engineering understanding and are only able to write eliezer’s textbook from the future after the fact.
I just think that if we solve alignment and alignment turns out to be hard, we’ll be much better at measuring alignment than we are currently. For example with confessions, adversarially constructed evals, production monitoring, interp stuff, and like a dozen other plausible things. But even with regular alignment evals, Anthropic was able to turn a reconstruction of the HF incident into a large advance on the SoTA of alignment evals.
This is if we solve alignment. Eval awareness is an unsolved problem, so the alignment evals may be far more sophisticated than just putting the AI in some artificial situation and checking if it seems to have aligned behavior.
An AI risk skeptic asked this on twitter and I replied:
Probably like the solution to airline safety, which is thousands of pages of design best practices plus a robust safety culture. These principles cannot be “written in blood” thru trial and error bc we’d go extinct first, so they need to result from a better understanding of ML + careful experiments. We probably can’t prove we’ve solved it, just get to a high level of confidence. But if alignment is hard, we will be able to generate much harder alignment evals that all current models fail.
What would a solution to alignment actually, concretely look like? Could we prove we have solved alignment without actually deploying an ASI?
I don’t think hyperstition through pretraining data is worth worrying about. There are several more important factors in whether AIs organize into unauthorized collectives than what language we use for them. Better to use accurate language that maintains relevant humans’ understanding of the world.
comms and comms-adjacent people are being confusing so I will directionally move towards leo’s policy instead
As of last month METR neither has the management capacity nor the engineers to have a team of 60 people, so probably more like six? Not sure what it would be with management assistance.
There is no single correct y axis for “AI capabilities”. It’s pretty clear that time horizon-like measures increase exponentially, while the ECI is constructed such that constant rates of progress (according to a particular definition inspired by Item Response Theory) mean linear increases.
Linearly increasing ECI empirically turns out to mean exponentially decreasing hazard rates, which means exponentially increasing time horizons https://www.tobyord.com/writing/half-life
The fact that time horizon increases exponentially replicated across domains.
In math, models have gone from being able to solve AMC12 problems (which experts can solve in a couple of minutes) through IMO problems (hours), to now solving open problems (years+)
From what you’re saying, the theory of general intelligence seems to be the bottleneck rather than having a superintelligence. GPT-6 seems to have enough of the pieces of general intelligence that if we had theory and could reverse engineer it, we could refine the theory.
My read is that this is an descriptive adjective, not a restrictive one. That is, the authors think all instances of unsanctioned actions on the internet [edit: at least since they deployed the monitoring system] are “unexpected” and are just saying it for emphasis.
How would we get from reverse-engineered gpt3 to an aligned superintelligence? If misalignment is pervasive enough to get >50% p(doom), it will be an emergent property intertwined with capabilities, and so we probably can’t just prune the misalignment circuits.
Suppose we had a superintelligence architecture scaled down to gpt-6 astra capability level. Or an actual superintelligence that is prevented by output safeguards from being immediately dangerous to us. What could we demonstrate about it that you would find really impressive progress towards alignment?
Within scalable oversight, what exactly would a really impressive concrete empirical result be?
Argh, forgot about that. Agreed. But cartoonishly evil is a stronger statement so it will be weaker evidence
Let’s say doom = permanent disempowerment or worse, for the purpose of this question.
It’s sad we can’t pay them because the UKAISI-Redwood pay gap (order of $300k) is tiny compared to the cost of donations that it takes to match a safety researcher’s impact (order of $30M), or even the Redwood-Anthropic pay gap (order of $3M). I feel like at the very least, former lab employees and others without financial concerns should preferentially work for UKAISI and CAISI, as long as they don’t have unresolvable COIs from their equity.