To put it further, I think there’s no way we crack “theoretically pure alignment” in time, and we only survive by handling matters of degree, not matters of kind. We play the cat and mouse game, and we win.
For people who advocate for pure alignment, do you think humanity is capable of solving pure alignment in the next 50 years?
the cat and mouse game is unwinnable if we are vastly dumber than the models. if pure alignment is impossible we need to shut down progress instead of whining that geopolitics is hard to our graves.
This is not necessarily true if we have previous models help us align future models, though it’s not necessarily untrue either. I wrote about this recently here.
To put it further, I think there’s no way we crack “theoretically pure alignment” in time, and we only survive by handling matters of degree, not matters of kind. We play the cat and mouse game, and we win.
For people who advocate for pure alignment, do you think humanity is capable of solving pure alignment in the next 50 years?
the cat and mouse game is unwinnable if we are vastly dumber than the models. if pure alignment is impossible we need to shut down progress instead of whining that geopolitics is hard to our graves.
This is not necessarily true if we have previous models help us align future models, though it’s not necessarily untrue either. I wrote about this recently here.