I work primarily on AI Alignment. Scroll down to my pinned Shortform for an idea of my current work and who I’d like to collaborate with.
Website: https://jacquesthibodeau.com
Twitter: https://twitter.com/JacquesThibs
GitHub: https://github.com/JayThibs
LinkedIn: https://www.linkedin.com/in/jacques-thibodeau/
fwiw, I’ve previously had much more weight on short timelines before[1] (I was already working on automated alignment in 2022 partly for this reason and have kept at it since then!), but the distribution has certainly expanded.[2]
Largely, I’ve found that many people arguing for short timelines have collapsed into lazy thinking, relying far too heavily on “the models are getting better” and “those who predicted longer timelines have continued to be proven wrong” (I’m not saying this doesn’t count as bits to update on). I don’t say this (and haven’t said this) to disparage the short-timeline view (I still place considerable weight on it), but I want to encourage clearer writing and more convincing arguments (not just for myself). Ultimately, I find many arguments lacking in the mechanistic description of where current capabilities come from and how that relates to RSI and “True AGI.”
Personally, my core reason for fixating on this is how much the details matter for understanding progress on superalignment, and how much hand-waving can make much of the AI safety field focus on entirely the wrong things! It may be true that timelines are short (again, I also can imagine quick paths to AGI), but I’d like more work on disentangling such predictions (which I intend to work on).
Either way, I think some folks are falling prey to the situation described in this post, “The goalposts are shrouded, not moving”.
When there’s a capability advance, there’s a tendency in some people to say folks are ‘moving goalposts’ in response to people saying, “Ok, but the model is not really doing x.” I’ve called people out on the goalpost moving too (and still think it is important to say in some contexts)!
Though, I feel like it’s more complicated than that, and worth investigating why.
Usually it’s because someone like Gary Marcus says something, but it’s the kind of thing that’s happened in the x-risk community too.
For example, it seems to me that many of the things that LLMs do today would have been predicted as ‘superintelligence’ or at least ‘AGI’ by the MIRI crew 6-7 years ago. However, this was likely because they guessed that AIs that can code and do impressive-seeming things like this would ALSO be good at xyz. Instead, on the journey to superintelligence, we ended up in this valley where AIs can make progress on the Riemann hypothesis yet can’t reliably do other basic tasks.
They had an underlying assumption that just wasn’t really articulated, and now it is labelled as ‘moving the goalpost’.
I think people would have a lot more clarity on AI progress and where things are going if they took a step back before having the knee-jerk “this person is dumb and moving goalposts” reaction. Sometimes goalpost-moving is a defence mechanism, but other times it’s about having a nuanced opinion (not generalized AI skepticism).
For example, someone could think, “ok, I still believe that the end state is that AIs will be incredible goal seekers, but now I need to make sense of why it could do x without capability y, which I had wrongly assumed was necessary.”
On the other side, I suspect that others are also developing underlying assumptions they may not have considered strongly. That is, they might assume, “If the AI is capable of making progress on the Riemann hypothesis, then it means that math is basically solved.”
In that case, they are sweeping under the rug things like AI being able to come up with a completely new mathematical paradigm that matters and does not leverage lots of previous work. Yet, we could be in a world where the model is superhuman at Riemann hypothesis-like tasks and counterexamples, but just keeps being pisspoor at coming up with new paradigms for much longer than they are implicitly assuming.
Some related thoughts from a tweet I previously wrote:
I think it depends on what we mean by “researchers”. In the past it was typically assumed that a researcher AI equates a “human-level researcher”, but the cognitive shape of LLM agents is different.
The difference here can lead to a completely different pace of progress. You could be 10000x-ing your ability to make plots, 1000x-ing your ability to solve specific kinds of problems, and 0.5 to 10x-ing your ability to make progress on novel, necessary approaches. We are now at the stage where we need to be specific about the types of problems these 10M researchers would solve.
People were wrong when they said “frontier math requires novel thinking”, but that doesn’t mean certain types of frontier math don’t require ‘novel thinking’. They just didn’t properly specify that there are indeed famous open problems LLMs will make progress on, largely due to interpolation and search (which allow LLMs to make progress, whereas humans are limited in how many fields they can know and how much persistence they have).
In other words, LLMs get access to a bunch of new low-hanging fruit because of their specific capability profile, but this doesn’t necessarily translate into the types of novelty that may be the main limiter (or at least slow progress).
Note that this is not to say that LLMs lack the capability of producing “novel” outputs, just that we now need to be more precise about the taxonomy of novelty and how much each branch is required to go faster and deeper down the tech tree.
The popularity and assumed difficulty (judged by humans) doesn’t matter *that* much. What matters is which cognitive moves were required for the agent to arrive at an answer to those problems and what that allows us to predict about future progress.
This is not a claim that “timelines are long”, but that the path down automated AI R&D may be much more complicated than doing a bunch of inference-time compute in the current paradigm. And what provides the unlock gets us fundamentally closer to the actual problems in alignment, in ways that make many current alignment techniques potentially ineffective.
And by short I was thinking >50% “AGI” by 2026-2027. To the point I started working on AI safety startups in 2024 with that thesis in mind (though partly as a hedge since I felt the field and philanthropy were dropping the ball on this happening).
I’ve communicated some of my reasoning here (and in the comments [1] [2] [3]) and here.