In the same way, the persona model of alignment may turn out ‘mostly right’, but still not give you enough precision to do what you want to do.
Adrià Garriga-alonso
We haven’t seriously tried to have physics-like theories of intelligence! Only a few people (various academics, MIRI) have done something close to trying. Compare to how much effort has been put into optimizing deep learning.
Maybe physics was also confusing and weird until we understood it.
I also think the strategic focus on “capabilities bad!” is in tension with the praise for scientific and conceptual progress
I think you’re missing the point here. It’s not that Richard thinks capabilities are bad (on vibes he does, but it’s not the point). Or that capabilities are bad (in my assessment, they’ve been good so far and things are going OK).
Rather, it’s that the explicit goal of the alignment community was to differentially advance alignment over capabilities, but instead they ended up advancing capabilities much more effectively than anybody else, while not advancing alignment much. This is a failure to follow the stated goal of colossal proportions, we in fact optimized the opposite of the goal. Why did this happen, and how can we avoid failing on goals this hard in the future?
This is the paragraph that most supports my comment:
To be clear, I’m not taking a strong stance in this sequence on whether AI will go well or badly—that seems up for grabs. My concern is that the alignment community had a plan to make good outcomes more likely (differentially advancing alignment over capabilities) but has mostly pushed the world in the opposite direction, while some parts of it gained a lot of power by doing so
Since… of course, if this were true, then alignment research (including agent foundations) would accelerate capabilities, right? Understanding how things work helps to make things work better? Obviously?
Yes, but it’s better than trying to make things work better without understanding them. (That said, if neural networks didn’t work this well, people would still entertain fanciful theories about Bayesian learning or parameter restriction, so maybe the effect of practice on theory has been good!)
would be more likely brought in by the things you put in the “prestige” bucket (like Superintelligence) than the “fanfic” bucket.
So many educated people are starstruck by HPMOR, or were starstruck before becoming really educated. Disagree with you that Superintelligence is more effective at recruiting (it’s too boring, facts alone don’t steer like emotions do).
and first (to my knowledge) made the link to a limiting infinite-width Gaussian process (which later evolved into Neural Tangent Kernel work.)
This is true. Not only that, but the use of Gaussian processes in machine learning comes from Radford Neal’s thesis and that they’re the limiting behavior of wide NNs in the first place.
YUGE! welcome back Paul!
Cool post, thank you! I’m very surprised that this natural selection model of SGD has only 1 mutant offspring. I would have sent many mutations in all directions and then made the selection distribution really sharp based on fitness. Which is effectively zeroth-order optimization, which approximates gradient descent on a smoothed function.
---
> where mutation (xt+1−xt)∼Pm.I think this is a typo and ought to be (x’_t—x_t) ~ Pm
So the distribution over xt+1 given xt is
P(xt+1|xt)=αsPm(xt+1−xt)Ps(f(xt+1)−f(x))
This is another typo, and the expression on the very right should be Ps(f(xt+1)−f(x_t))
funny to come back to this debate, when gradient hacking was demonstrated in late 2024, just one year later.
This is how frequentist statistics survived for so long, by the way.
This is correct for popularized-for-working-scientists frequentist statistics. I would just construct an estimator of the value of the action and then act based on that, without any posterior probabilities.
Or, if I just want to cache a fact, construct an estimator of whether it is 1 (true) or 0 (false) with consistency, low variance and low bias.
Nitpicking aside, this was a great post!! Thank you for writing it.
This post is heavy on math but light on explaining what it’s even trying to do. For example, why does this incorporate cryptographic signing? A very unorthodox choice, that I cannot see any purpose for.
What does QACI even stand for? It’s not in this post or the summary of Orthogonal. Is this esoteric on purpose?
it was shown by Lee et al.
Also contemporaneously Alexander G. de G. Matthews et al.! And, while less famous, that paper was better in one way: it took the limit of the width of all layers simultaneously, instead of one by one. That is, Lee et al was a statement about:
lim(width->infty) [ b_2 + W_2 nonlinearity( lim(width → infty) [W_1x + b_1])]
whereas Matthews et al was a statement about:
lim(width->infty)[ b_2 + W_2 nonlinearity(W_1x + b_1)]
which is more complicated
And then, if you are drawn in, next week it will be something a little further from the rules, and next year something further still, but all in the jolliest, friendliest spirit
This is incredibly similar to one interesting passage from Tolstoy’s The Death of Ivan Illyich, another famous moralizing story:
but all this [affairs, immoral things] was done with such a tone of good breeding that no hard names could be applied to it. It all came under the heading of the French saying: Il faut que jeunesse se passe. It was all done with clean hands, in clean linen, with French phrases, and above all among people of the best society and consequently with the approval of people of rank.
You can probably take out a loan to pay Nectome and MAiD, then pay back the loan with life insurance a couple of weeks later once you’re already preserved.
How much superposition is there?
We can monitor that and mitigate it when we get there, using the previous generation of AIs.
It goes from 1.05 to 1.2.
Thank you! I have fixed the link now.
Yeah, true. It’s gone so well for so long that I forgot. I didn’t spend a lot of time thinking about this list.
No, I think the blue-team will keep having the latest and best LLMs and be able to stop such attempts from randos. These AGIs won’t be so much magically superintelligent that they can take all the unethical actions needed to take over the world, without other AGIs stopping them.
On Chrome on a Mac you can just C-f in the PDF, it just OCRs automatically. I didn’t have this problem.
Really? What do you find so confusing about it?
It’s the kind of thing you do if you want to affect the world through power. It’s quite uncertain how far we are from understanding NNs well (anywhere from 1-10 years seems reasonable, even with our current AI helpers). And in the meantime someone else could invest in scaling NNs and get there, or you won’t be able to sustain enough expectation of future revenue for your investment multiples, or if you’re OpenAI then Anthropic will do it first and you lose power. (Though people are spooked enough by recent incidents that these two labs might start pacing the frontier now, so maybe your intuitions are more correct.)