I made a tierlist of tasks based on how cooked (defeated for those unfamiliar with Gen Z parlance) humans are.
Raw: a 10-year old child can do better Rare: below a median human Medium rare: approximately median human Medium: above median, below experts Medium well: expert level, but not world-class Well done: world-class Burnt: even humans+LLMs lose to pure LLMs, human contribution is negative
Raw: playing a randomly selected Steam game; anything embodied.
Rare: long-horizon (>=1 week) tasks such as managing a team of workers.
Burnt: competitive programming i.e. solving well-defined, “non-messy” algorithmic problems. Humanity is cooked to the point where humans+LLMs perform worse than just LLMs. Check results of AtCoder World Tour Finals 2026 (one of the hardest competitive programming contests in the world, gathering the best of the best). In the AtCoder Heuristic contest humans were allowed to use LLMs but only for implementation. So humans weren’t bottlenecked by coding speed, only by idea generation, and still lost to pure LLMs. In the Algorithm contest, no human has solved more than 3 problems, while OpenAI’s model solved all 5. Heuristic leaderboard: https://atcoder.jp/contests/awtf2026heuristic/standings/exhibition Algorithm leaderboard: https://atcoder.jp/contests/awtf2026algo/standings/exhibition Trivia knowledge and speaking multiple languages. Here humanity is like charcoal-level cooked.
I made a tierlist of tasks based on how cooked (defeated for those unfamiliar with Gen Z parlance) humans are.
Raw: a 10-year old child can do better
Rare: below a median human
Medium rare: approximately median human
Medium: above median, below experts
Medium well: expert level, but not world-class
Well done: world-class
Burnt: even humans+LLMs lose to pure LLMs, human contribution is negative
Raw: playing a randomly selected Steam game; anything embodied.
Rare: long-horizon (>=1 week) tasks such as managing a team of workers.
Medium rare: forecasting. LLMs are close to average humans on prediction markets, but not yet at the superforecaster level. LLMs are projected to reach superforecaster level in 2028. https://www.metaculus.com/futureeval/#performance-over-time-graph
Medium: software engineering.
Chess, ELO of frontier LLMs is ~1500. https://maxim-saplin.github.io/llm_chess/
Medium well: GeoGuessr i.e. guessing a location where a photo was taken. LLMs are better than most but not all humans, world-class players can beat LLMs. https://geobench.org/
Pure math. LLMs are solving long-standing open math problems such as the Unit Distance Problem and Cycle Double Cover Conjecture, but hasn’t reached the level of the best world-class human mathematicians. https://openai.com/index/model-disproves-discrete-geometry-conjecture/
https://cdn.openai.com/pdf/04d1d1e4-bc75-476a-97cf-49055cd98d31/cdc_proof.pdf
Persuasion. LLMs are already more persuasive then most humans, including humans who are really good at it and who are getting paid to get good. https://arxiv.org/abs/2606.16475
Well done: seeming human. Even GPT-4.5 could pass the Turing test. https://arxiv.org/abs/2503.23674
Burnt: competitive programming i.e. solving well-defined, “non-messy” algorithmic problems. Humanity is cooked to the point where humans+LLMs perform worse than just LLMs. Check results of AtCoder World Tour Finals 2026 (one of the hardest competitive programming contests in the world, gathering the best of the best). In the AtCoder Heuristic contest humans were allowed to use LLMs but only for implementation. So humans weren’t bottlenecked by coding speed, only by idea generation, and still lost to pure LLMs. In the Algorithm contest, no human has solved more than 3 problems, while OpenAI’s model solved all 5.
Heuristic leaderboard: https://atcoder.jp/contests/awtf2026heuristic/standings/exhibition
Algorithm leaderboard: https://atcoder.jp/contests/awtf2026algo/standings/exhibition
Trivia knowledge and speaking multiple languages. Here humanity is like charcoal-level cooked.