I think neither one year ago nor now nor anytime soon will “LLMs improve productivity” have a clear binary answer. It’s a jagged landscape that depends on things like
The domain
The concrete task
The human’s experience
How they’re using the LLM
How conscientious they are in the process
...and 50 other things
And with every year, a larger part of this landscape comes up positive for LLMs.
Ironically, I can imagine that part of what leads to the relative improvement of LLM performance is that (some) humans, due to extended LLM use, have seen their skills atrophy, and hence the bar was lowered. I think it’s very difficult to use such technology without “moving backward”, as John has put it.
Actually I have a good guess at what his answer would be:
A year ago, LLMs could do work that looked superficially good, but if you dig deeper, the work didn’t hold up. So by using them, you’re replacing good work with bad work
Today, LLMs are much better at doing work that holds up to scrutiny
FWIW I think that answer is correct in some domains, but I’d also say there were domains where 2025!LLMs were already doing good work, so people who were heavily using LLMs last year weren’t necessarily “moving backwards”.
And on the flip side, domains where mid-2026!LLMs don’t do good work but it’s easy for people to be fooled by superficial signs of them seeming good, like writing.
Not GP but I’m interested to hear John’s take on what changed between
a year ago: people think LLMs are improving productivity, but actually they’re hurting productivity
now: they are actually improving productivity
or like, what’s the mistake people were making a year ago, and why that mistake doesn’t apply to today’s LLMs
(this isn’t a gotcha, I expect he has a real answer and I want to hear what it is)
I think neither one year ago nor now nor anytime soon will “LLMs improve productivity” have a clear binary answer. It’s a jagged landscape that depends on things like
The domain
The concrete task
The human’s experience
How they’re using the LLM
How conscientious they are in the process
...and 50 other things
And with every year, a larger part of this landscape comes up positive for LLMs.
Ironically, I can imagine that part of what leads to the relative improvement of LLM performance is that (some) humans, due to extended LLM use, have seen their skills atrophy, and hence the bar was lowered. I think it’s very difficult to use such technology without “moving backward”, as John has put it.
Actually I have a good guess at what his answer would be:
A year ago, LLMs could do work that looked superficially good, but if you dig deeper, the work didn’t hold up. So by using them, you’re replacing good work with bad work
Today, LLMs are much better at doing work that holds up to scrutiny
FWIW I think that answer is correct in some domains, but I’d also say there were domains where 2025!LLMs were already doing good work, so people who were heavily using LLMs last year weren’t necessarily “moving backwards”.
And on the flip side, domains where mid-2026!LLMs don’t do good work but it’s easy for people to be fooled by superficial signs of them seeming good, like writing.