Would you say that LLMs are already good at the ‘invention’ part?
it is hard to measure how “good” they are without a truly formidable experimentation program. however, models have been observed inventing languages to solve a hard problem (and others in its class), which is already astounding to me, and is literally the definition of intelligence used by some computer scientists.
for example see section 6.4.5 in the Claude Sonnet 5 System Card …
the behavior is somewhat forced by limited prompt context and hours long runs, but is still super interesting.
> they have a decent grasp of already existing languages and can apply them in appropriate circumstances that humans have overlooked
”overlooked” makes it sound like everything was right before the expert and they still failed to notice it. I think that there is a little more going on in mathematics in that “proof space” is so vast that an expert could search it forever and still never encounter the particular region with the key that unlocks the solution. while the patterns of mathematics and the LLM’s brute force speed makes up for some of this, it is still only a few of orders of magnitude, so that’s where the mysterious “research taste” comes in, which somehow predicts productive areas to look with only a preliminary survey of the problem (both in humans and LLM-based autoresearchers!).
it is hard to measure how “good” they are without a truly formidable experimentation program. however, models have been observed inventing languages to solve a hard problem (and others in its class), which is already astounding to me, and is literally the definition of intelligence used by some computer scientists.
for example see section 6.4.5 in the Claude Sonnet 5 System Card …
https://www.anthropic.com/claude-sonnet-5-system-card
the behavior is somewhat forced by limited prompt context and hours long runs, but is still super interesting.
> they have a decent grasp of already existing languages and can apply them in appropriate circumstances that humans have overlooked
”overlooked” makes it sound like everything was right before the expert and they still failed to notice it. I think that there is a little more going on in mathematics in that “proof space” is so vast that an expert could search it forever and still never encounter the particular region with the key that unlocks the solution. while the patterns of mathematics and the LLM’s brute force speed makes up for some of this, it is still only a few of orders of magnitude, so that’s where the mysterious “research taste” comes in, which somehow predicts productive areas to look with only a preliminary survey of the problem (both in humans and LLM-based autoresearchers!).