I think your point about the “something missing in the deep understanding” part is true—the math examples are a great one.
But we seem to have gotten to our current level of AI without any deep understanding of intelligence, mostly scaling data, some trivial algorithms and brute force engineering, and tacking on some RL to it.
I don’t know if you’d consider the transformer architecture a deep insight, or something that a current Fable/Astra level model would be able to discover independently? (in a hypothetical world where it didn’t know about it or any other related developments).
I don’t know if you’d consider the transformer architecture a deep insight, or something that a current Fable/Astra level model would be able to discover independently?
Depends on the starting point. Is it RNNs, or pre-deep-learning machine learning, or...?
I think your point about the “something missing in the deep understanding” part is true—the math examples are a great one.
But we seem to have gotten to our current level of AI without any deep understanding of intelligence, mostly scaling data, some trivial algorithms and brute force engineering, and tacking on some RL to it.
I don’t know if you’d consider the transformer architecture a deep insight, or something that a current Fable/Astra level model would be able to discover independently? (in a hypothetical world where it didn’t know about it or any other related developments).
Depends on the starting point. Is it RNNs, or pre-deep-learning machine learning, or...?
I do think LLMs are ultimately capable of e. g. RNN → Transformer jumps, or say Vanilla Transformer → Kimi K3.