I expect he’s probably overstating the extent to which MIRI’s strategy is actually well-summarized as trying to implement RSI themselves
This may be a reasonable prior but I think I’ve offered enough evidence to overcome it. (See also Outside View(s) and MIRI’s FAI Endgame if you need more.)
And even if I believed him on that point, I don’t know how he thinks that should affect any of the specific points I made in my post—he’s leaving all of that interpretative labor up to his readers
In the OP you wrote “The early rationalist community was a beacon of intellectual clarity.” and I’m pointing out that part of this supposed clarity was a plan to build FAI with a small team and potentially just 1 philosopher. Here’s a longer explanation if it’s still not clear:
Ngo attributes the current dysfunction of the alignment field to external pressures, bad incentives, and cognitive errors like “jumping down the slippery slope” to chase power and conform to mainstream machine learning.
Wei Dai shifts the diagnosis inward. He argues that the community’s failures are not just the result of recent compromises, but rather a reflection of a more fundamental, ongoing issue: that humans in general, and MIRI leadership in particular, have historically been very bad at strategy and philosophy.
I guess the upvotes come from people really liking Wei’s arguments about the difficulty of philosophy?
My guess is that it’s more a combination of people seeing the relevance of the comment to the OP, and liking solid criticisms of Eliezer and MIRI. (E.g. Holden’s post had huge karma for that time.)
It would feel less like that if he had an alternative he endorsed instead of doing philosophy (since we’re so bad at it), or a historically-grounded argument for why he thinks philosophy is so important, or even a clear characterization of what he thinks philosophy is. (This post is the closest I can find, but it’s from a while ago, and is still pretty handwavey.)
Did you see The Long (Self-)Correction which seems relevant here? From the conclusion: “My main hope for a Long Self-Correction eventually succeeding rests on the fact that humans have seemingly, mysteriously, made progress on these issues over a very long period of time, so if we preserve the environment in which we can seemingly do this, and not give anyone or anything the power to permanently derail such progress, then maybe we can continue to snowball The Correction until we reach a point when we can rightly justify reshaping the universe according to our volition.”
And yes, Some Thoughts on Metaphilosophy is still my best take on what philosophy is, and I fully admit that it’s very handwavy. Lack of further progress is of course disappointing, but I’ve always said that trying to solve metaphilosophy would be a long shot (in the short/medium run).
And yes, Some Thoughts on Metaphilosophy is still my best take on what philosophy is
OT, but what do you think of my take that philosophy is about mesa-optimizers extracting patterns (relevant to their functioning) from their underlying statistical model into a legible format?
This may be a reasonable prior but I think I’ve offered enough evidence to overcome it. (See also Outside View(s) and MIRI’s FAI Endgame if you need more.)
In the OP you wrote “The early rationalist community was a beacon of intellectual clarity.” and I’m pointing out that part of this supposed clarity was a plan to build FAI with a small team and potentially just 1 philosopher. Here’s a longer explanation if it’s still not clear:
Ngo attributes the current dysfunction of the alignment field to external pressures, bad incentives, and cognitive errors like “jumping down the slippery slope” to chase power and conform to mainstream machine learning.
Wei Dai shifts the diagnosis inward. He argues that the community’s failures are not just the result of recent compromises, but rather a reflection of a more fundamental, ongoing issue: that humans in general, and MIRI leadership in particular, have historically been very bad at strategy and philosophy.
My guess is that it’s more a combination of people seeing the relevance of the comment to the OP, and liking solid criticisms of Eliezer and MIRI. (E.g. Holden’s post had huge karma for that time.)
Did you see The Long (Self-)Correction which seems relevant here? From the conclusion: “My main hope for a Long Self-Correction eventually succeeding rests on the fact that humans have seemingly, mysteriously, made progress on these issues over a very long period of time, so if we preserve the environment in which we can seemingly do this, and not give anyone or anything the power to permanently derail such progress, then maybe we can continue to snowball The Correction until we reach a point when we can rightly justify reshaping the universe according to our volition.”
And yes, Some Thoughts on Metaphilosophy is still my best take on what philosophy is, and I fully admit that it’s very handwavy. Lack of further progress is of course disappointing, but I’ve always said that trying to solve metaphilosophy would be a long shot (in the short/medium run).
OT, but what do you think of my take that philosophy is about mesa-optimizers extracting patterns (relevant to their functioning) from their underlying statistical model into a legible format?