In a recent post, I mentioned FrontierMath Open Problems benchmark as one of the few I take seriously for measuring intelligence (as opposed to e.g. coding benchmarks). My mostly-a-joke LaughBench benchmark is also starting to show minor results. This is the first time I’m sensing that the AI models I use have some kind of spark of intelligence inside them.
@Zvi somehow missed the first Solid Result from EpochAI’s FrontierMath Open Problems benchmark.
In a recent post, I mentioned FrontierMath Open Problems benchmark as one of the few I take seriously for measuring intelligence (as opposed to e.g. coding benchmarks). My mostly-a-joke LaughBench benchmark is also starting to show minor results. This is the first time I’m sensing that the AI models I use have some kind of spark of intelligence inside them.