LaughBench
Introducing LaughBench.
For a long time, I’ve considered the ability for AI models to tell novel, funny jokes that actually make people laugh to be a robust indicator of real general intelligence (as opposed to, say, coding tasks). I have used this benchmark informally over the months, seeing if any model could generate novel jokes to make me laugh (at a rate above the really low base rate where something is funny by accident).
So far, the answer has always been no. As you can see from the chart in the repo, none of the models succeeded in producing laughter:
The frontier models (GPT-5.6 Sol and Fable) are able to occasionally make me smile. This indicates there’s a spark of intelligence present that wasn’t present to the same degree in prior models.
There are a few reasons I think laughter-producing comedy is a good benchmark for intelligence (as mentioned in the repo):
Professional comedians, whose careers depend almost entirely on generating laughter, are observed to have well-above-average IQs.
Humor probably evolved as a reward signal for detecting faulty assumptions in your predictive model of the world. To make someone laugh, you have to model their mind, anticipate a false inference they would make, and avoid making the same false inference yourself. This is IQ-hard.
Other benchmarks for intelligence, such as those based on math or programming tasks, are aided by general intelligence, but you can also go a long way with an incredible memory for past problems, and lower-level computation/recombination based on these memories. That is, math and programming tasks used in existing benchmarks are not often “IQ-hard”.
Other IQ-hard tasks (I believe) included things like writing good stories, doing AI research that leads to real capability gains, coming up with competitive business ideas, and so on. But these are hard to measure. It’s easy to convince yourself your business idea is better than it actually is (and besides, an excellent execution would permit a mediocre initial idea). But laughter is fairly objective. You laugh or you don’t. You can’t squint and say, “yeah, I think this is a good poem” like people did back in the ChatGPT 3.5 days with terrible AI-generated poetry.
It seems important to have benchmarks aimed at measuring “real intelligence” of AI models, whatever that means, since I suspect many dangerous activities will become possible once real intelligence hits a certain point (creating bioweapons, manipulating/deceiving human beings, creating long-term plans that aren’t doomed to failure, advancing physics, robotics, biology, or AI research, etc.) Most of what we want in the good AI future also requires this ethereal “real intelligence”: curing all diseases and restructuring society for the common good are also IQ-hard activities.
Some benchmarks already have the flavor of this. When I looked into it, I thought FrontierMath Open Problems was at least a partial fit. The ARC-AGI stuff too. But ARC-AGI seems more like something that is easier if you have more IQ, but sufficient spatial reasoning and other memorizable skills can make it easier and generalize across games. (Similar complaint about FrontierMath.)
In the past I have mentioned the possibility of benchmarking AI performance on novel board games (with branching complexity at least equal to chess etc.). But generating arbitrary board games with no significant commonalities to existing board games seems pretty hard. You could take the ARC-AGI approach of hand crafting them, but doing it correctly is IQ-hard for the author.
Or you could just ask the AI to tell you a novel joke.
The ironic thing about humor as an IQ test is that IMO, it’s one that LessWrong and the Rationalist community is not doing so great on… It did occur to me one day when I realized what an outlier Scott Alexander is in this regard. As far as my personal funny bone is considered, anyway.
I wonder if Anthropic got this partially covered. Page 212 of Mythos Preview’s card had Anthropic describe how “Claude Mythos Preview comes up with decent and seemingly novel ones [puns—S.K.], often relating to its preferred technical and philosophical topics” (alas, Fable 5′s jokes aren’t described in Fable’s card, and neither are Opus 5′s jokes).
P.S. What jokes by GPT-5.6 Sol or Fable 5 made you smile?
The distinction between “decent” jokes and ones that actually make a person laugh seems like an important one to me.
I unfortunately don’t remember what the jokes were.
My experience is that LLMs are better at being funny when they are doing ordinary writing and mixing in a witty line, than when they are explicitly prompted to tell a joke. This has made me laugh out loud once and grin a decent amount, though it’s the kind of thing that’s awkward to benchmark given that explicit prompting makes them “try too hard” and get less funny.
True for people as well.
Maybe when the labs make an rlvr environment for this will this benchmark finally be saturated