Thanks to the “alignment engineering,” we have at least some understanding of what’s going on and partially know why. Alignment is a moving target, and engineers mostly register events, sometimes explaining them in retrospect; that’s why this cannot be called science in principle. The AGI companies invested everything in this “moving target”, putting everything on the line, so we should not expect any other results from them. The “misalignment science” can suffer from the same problems.
In my opinion, the solution isn’t to run after the target, but to rebuild the target itself. The transformer is an extremely simple architecture with just a few inductive biases. We can’t expect aligned behavior from this technology. This was a theoretical idea, but now it’s backed by experience, and experience builds up.
Alignment science is just computer science. If this isn’t possible, then alignment isn’t possible in principle.
Thanks to the “alignment engineering,” we have at least some understanding of what’s going on and partially know why. Alignment is a moving target, and engineers mostly register events, sometimes explaining them in retrospect; that’s why this cannot be called science in principle. The AGI companies invested everything in this “moving target”, putting everything on the line, so we should not expect any other results from them. The “misalignment science” can suffer from the same problems.
In my opinion, the solution isn’t to run after the target, but to rebuild the target itself. The transformer is an extremely simple architecture with just a few inductive biases. We can’t expect aligned behavior from this technology. This was a theoretical idea, but now it’s backed by experience, and experience builds up.
Alignment science is just computer science. If this isn’t possible, then alignment isn’t possible in principle.