It’s great to collect these as canaries, though in my framing the underlying failure is a lack of general intellegence. The road from “1600 on the SAT” (or whatever benchmark) to “drop-in remote worker” is paved with thousands of illegible out-of-distribution tasks, thus no such LLM workers exist today.
Any canary that is too forecasting-culture-load-bearing (e.g. ARC-AGI I) gets subjected to lab optimization pressure (e.g. special RLVR envs, or even just researcher attention) and thus ceases to be part of the positive manifold, i.e. it becomes a spike in the the spiky capabilities of LLMs.
With “positive manifold” here I mean the psychometrical concept: A broad set of cognitive abilities is positively correlated, you can expect a person with an “1600 on the SAT” to be a “drop-in remote worker” by default.
I don’t think this is entirely true. ARC-AGI has been going for years, is under a huge amount of optimization pressure, and is still one of the domains where LLM gets beaten by a 12 year old.
If it turns out we can optimize any trait we want with a year of attention, that’s pretty important information in and of itself. It means ASI will probably be able to fix any remaining gaps in the jagged frontier once it hits a certain skill in self-improvement.
I don’t think this is entirely true. ARC-AGI has been going for years, is under a huge amount of optimization pressure, and is still one of the domains where LLM gets beaten by a 12 year old.
I meant something like: “This ARC-AGI-1 benchmark isn’t beaten by any frontier model”. Then, it catches OpenAI’s attention, and o1 beats it Currently, ARC-AGI seems to be in the third version, and I think it’s a testimony to the jagged frontier that a relatively small modification (i.e. v1 to v2 to v3) throws the LLM’s off balance.[1]
If it turns out we can optimize any trait we want with a year of attention, that’s pretty important information in and of itself. It means ASI will probably be able to fix any remaining gaps in the jagged frontier once it hits a certain skill in self-improvement.
The problem is that the jagged frontier, in my view, is more like a bunch of tiny islands of legible benchmark-tasks in the vast ocean of tasks that human general intelligence is able to solve, the tasks that you would need to have an end-to-end remote worker. For example, to the best of my knowledge, there isn’t even a single job that today’s LLM’s can do fully autonomously.
That is also the reason why evolution didn’t converge on “massive modularity”, i.e. the brain isn’t an enourmous bag of tricks that we deploy specific to each new task, it’s more like a single trick (general intelligence) that does all the other tricks, to use the words of Yudkowsky.
Therefore we can’t have ASI “close any remaining gaps”, because we would need a thing that doesn’t have such gaps in the first place to have AGI.
Epistemic status disclaimer: I’ve only briefly checked the history of ARC-AGI, maybe someone with more subject-matter knowledge has a different framing here.
ARC-AGI-v1 was fairly simple static visual problems. ARC-AGI-v3 involves playing a simple set of video games. So it’s more like the former skills were a prerequisite for latter challenges to even make sense.
I do feel like we’ve seen a lot of broadening of intelligence, though: they can write poetry, cold read humans, operate a desktop computer using a mouse, solve long-standing math problems, etc. all without too much explicit training.
They don’t learn to solve “ARC-AGI-v1” puzzles, they learn to solve “the entire category that problem represents” and I’d personally lean towards there only being a few dozen such categories.
If we saw narrower learning, where “writing” didn’t include poetry, or “art” couldn’t mimic specific styles, I’d agree with a thousand islands. Although I won’t be surprised if we do see a few islands outside the major continents of learning.
It’s great to collect these as canaries, though in my framing the underlying failure is a lack of general intellegence. The road from “1600 on the SAT” (or whatever benchmark) to “drop-in remote worker” is paved with thousands of illegible out-of-distribution tasks, thus no such LLM workers exist today.
Any canary that is too forecasting-culture-load-bearing (e.g. ARC-AGI I) gets subjected to lab optimization pressure (e.g. special RLVR envs, or even just researcher attention) and thus ceases to be part of the positive manifold, i.e. it becomes a spike in the the spiky capabilities of LLMs.
With “positive manifold” here I mean the psychometrical concept: A broad set of cognitive abilities is positively correlated, you can expect a person with an “1600 on the SAT” to be a “drop-in remote worker” by default.
I don’t think this is entirely true. ARC-AGI has been going for years, is under a huge amount of optimization pressure, and is still one of the domains where LLM gets beaten by a 12 year old.
If it turns out we can optimize any trait we want with a year of attention, that’s pretty important information in and of itself. It means ASI will probably be able to fix any remaining gaps in the jagged frontier once it hits a certain skill in self-improvement.
I meant something like: “This ARC-AGI-1 benchmark isn’t beaten by any frontier model”. Then, it catches OpenAI’s attention, and o1 beats it Currently, ARC-AGI seems to be in the third version, and I think it’s a testimony to the jagged frontier that a relatively small modification (i.e. v1 to v2 to v3) throws the LLM’s off balance.[1]
The problem is that the jagged frontier, in my view, is more like a bunch of tiny islands of legible benchmark-tasks in the vast ocean of tasks that human general intelligence is able to solve, the tasks that you would need to have an end-to-end remote worker. For example, to the best of my knowledge, there isn’t even a single job that today’s LLM’s can do fully autonomously.
That is also the reason why evolution didn’t converge on “massive modularity”, i.e. the brain isn’t an enourmous bag of tricks that we deploy specific to each new task, it’s more like a single trick (general intelligence) that does all the other tricks, to use the words of Yudkowsky.
Therefore we can’t have ASI “close any remaining gaps”, because we would need a thing that doesn’t have such gaps in the first place to have AGI.
Epistemic status disclaimer: I’ve only briefly checked the history of ARC-AGI, maybe someone with more subject-matter knowledge has a different framing here.
ARC-AGI-v1 was fairly simple static visual problems. ARC-AGI-v3 involves playing a simple set of video games. So it’s more like the former skills were a prerequisite for latter challenges to even make sense.
I do feel like we’ve seen a lot of broadening of intelligence, though: they can write poetry, cold read humans, operate a desktop computer using a mouse, solve long-standing math problems, etc. all without too much explicit training.
They don’t learn to solve “ARC-AGI-v1” puzzles, they learn to solve “the entire category that problem represents” and I’d personally lean towards there only being a few dozen such categories.
If we saw narrower learning, where “writing” didn’t include poetry, or “art” couldn’t mimic specific styles, I’d agree with a thousand islands. Although I won’t be surprised if we do see a few islands outside the major continents of learning.