Mammalian/human memory systems feel seamless to us, but neurological/psychological study of them shows they’re actually quite a complicated cludge of working memory, short-term, medium-term, and long-term episodic memory, procedural memory, traumatic memory, and so forth: there are probably something of the order of half-a-dozen-to-a-dozen subsystems working together in human memory/learning.
For LLMs, the capabilities people have been working hard on this for the last half-decade or so, and so far as widely-used memory/learning mechanisms we have: initial model training, retrieval-augmented generation and memory summarization, and in-context learning plus context compaction. That’s not nothing, and each of them is improving, but the combination still has a distinct weakness in the sort of on-them-job learning that a lot of human workers do, often even during a specific extended project.
At this rate of advance, it seems unlikely to me that we’re going to devise a complicated architecture as capable as the human one in the next couple of years. So this makes me a little septical about some of the “full AGI in the next 2–3 years” claims: on those timescales I’m expecting “sort of AGI but with online learning less good than humans”.
Mammalian/human memory systems feel seamless to us, but neurological/psychological study of them shows they’re actually quite a complicated cludge of working memory, short-term, medium-term, and long-term episodic memory, procedural memory, traumatic memory, and so forth: there are probably something of the order of half-a-dozen-to-a-dozen subsystems working together in human memory/learning.
For LLMs, the capabilities people have been working hard on this for the last half-decade or so, and so far as widely-used memory/learning mechanisms we have: initial model training, retrieval-augmented generation and memory summarization, and in-context learning plus context compaction. That’s not nothing, and each of them is improving, but the combination still has a distinct weakness in the sort of on-them-job learning that a lot of human workers do, often even during a specific extended project.
At this rate of advance, it seems unlikely to me that we’re going to devise a complicated architecture as capable as the human one in the next couple of years. So this makes me a little septical about some of the “full AGI in the next 2–3 years” claims: on those timescales I’m expecting “sort of AGI but with online learning less good than humans”.