Yeah, the “Bitter Lesson” refers to a special case of this classic mistake, as do the other essays I linked. Some of those essays were quite well known in their day, at least to various groups of practitioners.
You could do it up in the classic checklist meme format:
Your brilliant AI plan will fail because:
[ ] You assume that you can somehow make the inner workings of intelligence mostly legible.
The people who learn this unpleasant lesson the fastest are AI researchers who process inputs that are obviously arrays of numbers. For example, sound and images are giant arrays of numbers, so speech recognition researchers have known what’s up for decades. But researchers who worked with either natural language or (worse) simplified toy planning systems often thought that they could handwave away the arrays of numbers and find a nice, clear, logical “core” that captured the essence of intelligence.
I want to be clear: Lots of terrifyingly smart people made this mistake, including some of the smartest scientists who ever lived. Many of them made this mistake for a decade or more before wising up or giving up.
But if you slap a camera and a Raspberry Pi onto a Roomba chassis, and wire up a simple gripper arm, then you can speed-run the same brutal lessons in a year, max. You’ll learn that the world is an array of numbers, and that the best “understanding” you can obtain about the world in front of your robot is a probability distribution over “apple”, “Coke can”, “bunch of cherries” or “some unknown reddish object”, each with a number attached. The transformation that sits between the array and the probability distribution always includes at least one big matrix that’s doing illegible things.
Neural networks are just bunches of matrices with even more illegible (non-linear) complications. Biological neurons take the matrix structure and bury it under more than a billion years of biochemistry and incredible complications we’re only starting to discover.
Like I said, this is a natural mistake, and smarter people than most of us here have made this mistake, sometimes for a decade or more.
I want to be clear: Lots of terrifyingly smart people made this mistake, including some of the smartest scientists who ever lived. Many of them made this mistake for a decade or more before wising up or giving up.
Imagine this. Imagine a future world where gradient-driven optimization never achieves aligned AI. But there is success of a different kind. At great cost, ASI arrives. Humanity ends. In his few remaining days, a scholar with the pen name of Rete reflects back on the 80s approach (i.e. using deterministic rules and explicit knowledge) with the words: “The technology wasn’t there yet; it didn’t work commercially. But they were onto something—at the very least, their approach was probably compatible with provably safe intelligence. Under other circumstances, perhaps it would have played a more influential role in promoting human thriving.”
Yeah, the “Bitter Lesson” refers to a special case of this classic mistake, as do the other essays I linked. Some of those essays were quite well known in their day, at least to various groups of practitioners.
You could do it up in the classic checklist meme format:
The people who learn this unpleasant lesson the fastest are AI researchers who process inputs that are obviously arrays of numbers. For example, sound and images are giant arrays of numbers, so speech recognition researchers have known what’s up for decades. But researchers who worked with either natural language or (worse) simplified toy planning systems often thought that they could handwave away the arrays of numbers and find a nice, clear, logical “core” that captured the essence of intelligence.
I want to be clear: Lots of terrifyingly smart people made this mistake, including some of the smartest scientists who ever lived. Many of them made this mistake for a decade or more before wising up or giving up.
But if you slap a camera and a Raspberry Pi onto a Roomba chassis, and wire up a simple gripper arm, then you can speed-run the same brutal lessons in a year, max. You’ll learn that the world is an array of numbers, and that the best “understanding” you can obtain about the world in front of your robot is a probability distribution over “apple”, “Coke can”, “bunch of cherries” or “some unknown reddish object”, each with a number attached. The transformation that sits between the array and the probability distribution always includes at least one big matrix that’s doing illegible things.
Neural networks are just bunches of matrices with even more illegible (non-linear) complications. Biological neurons take the matrix structure and bury it under more than a billion years of biochemistry and incredible complications we’re only starting to discover.
Like I said, this is a natural mistake, and smarter people than most of us here have made this mistake, sometimes for a decade or more.
Imagine this. Imagine a future world where gradient-driven optimization never achieves aligned AI. But there is success of a different kind. At great cost, ASI arrives. Humanity ends. In his few remaining days, a scholar with the pen name of Rete reflects back on the 80s approach (i.e. using deterministic rules and explicit knowledge) with the words: “The technology wasn’t there yet; it didn’t work commercially. But they were onto something—at the very least, their approach was probably compatible with provably safe intelligence. Under other circumstances, perhaps it would have played a more influential role in promoting human thriving.”