I think that’s the fundamental question. Does LLMs’ ability to autonomously perform basic hyperparameter search get them far enough that they can perform architecture optimization? Does that get them far enough that they can pursue new paradigms for language modeling[1]? If it takes 1 intelligence to go from 1 to 2, but 2.5 intelligence to go from 2 to 3, then 2 is where you stop.
The practical answer is that, from our perspective, “as smart as a human engineer across all relevant domains” gets us to “the best AI humans will ever be able to create” quite a bit quicker than we’d otherwise get there, and without the need for any further input from human engineers.
“as smart as a human engineer across all relevant domains”
I dispute that LLMs are like this; I think they and their training have a bunch of performance capability and not much ability to generate those de novo.
gets us to “the best AI humans will ever be able to create” quite a bit quicker than we’d otherwise get there,
Maybe; to some extent I’d expect this to hit various walls, though not sure; Amdahl’s law; and IDK how people get very confident of this.
I think that’s the fundamental question. Does LLMs’ ability to autonomously perform basic hyperparameter search get them far enough that they can perform architecture optimization? Does that get them far enough that they can pursue new paradigms for language modeling[1]? If it takes 1 intelligence to go from 1 to 2, but 2.5 intelligence to go from 2 to 3, then 2 is where you stop.
The practical answer is that, from our perspective, “as smart as a human engineer across all relevant domains” gets us to “the best AI humans will ever be able to create” quite a bit quicker than we’d otherwise get there, and without the need for any further input from human engineers.
I’m not suggesting that this is the exact trajectory.
I dispute that LLMs are like this; I think they and their training have a bunch of performance capability and not much ability to generate those de novo.
Maybe; to some extent I’d expect this to hit various walls, though not sure; Amdahl’s law; and IDK how people get very confident of this.