I disagree because I’m yet to see any of those “promising new architectures” outperform even something like GPT-2 345M, weight for weight, at similar tasks. Or show similar performance with a radical reduction in dataset size. Or anything of the sort.
I don’t doubt that a better architecture than LLM is possible. But if we’re talking AGI, then we need an actual general architecture. Not a benchmark-specific AI that destroys a specific benchmark, but a more general purpose AI that happens to do reasonably well at a variety of benchmarks it wasn’t purposefully trained for.
I disagree because I’m yet to see any of those “promising new architectures” outperform even something like GPT-2 345M, weight for weight, at similar tasks. Or show similar performance with a radical reduction in dataset size. Or anything of the sort.
I don’t doubt that a better architecture than LLM is possible. But if we’re talking AGI, then we need an actual general architecture. Not a benchmark-specific AI that destroys a specific benchmark, but a more general purpose AI that happens to do reasonably well at a variety of benchmarks it wasn’t purposefully trained for.
We aren’t exactly swimming in that kind of thing.