I briefly discussed with Dwarkesh why I’m skeptical AI progress is heavily driven by scaling up spending on human experts labeling/making data. My main argument is that spending on researchers and experiment compute seem much higher. But I didn’t say very much in the podcast.
More precisely, my view is that if spending on having human experts label/make individual data points were fixed at ~$100 million / year (per company), then AI progress would be <25% slower.
We didn’t have time to get into everything in this podcast (and some content about this recorded at an earlier point was cut), so I’ll spell out my view in a bit more detail here:
Spending directly on data (rather than on R&D about data) isn’t growing that fast and isn’t that high (relative to spending on researchers and especially compared to spending on experiment compute).
It’s important to make a distinction between spending directly on making data and science about better processes for making data. E.g., better data mixes like FineWeb count as R&D (and the person doing the R&D needs almost no understanding of individual sequences).
AI automation seems differentially good at accelerating data generation, such that I think improvements in RL environments have mostly been driven by improved AI (and R&D into making better environments) rather than spending on humans making/labeling data, and I expect this to continue. Like data stuff seems particularly amenable to acceleration from weaker AIs.
Structurally, most of what human data labeling does (though not all!) depends on having generally decent judgment rather than on having more expertise than the AIs being trained.
Transfer without domain-specific labels looks decent in practice. E.g., it doesn’t seem like Anthropic is hiring a ton of mathematicians and cyber experts to do data labeling, and the AIs are still good and getting better at these domains. Maybe this depends on having labels in some domain, but so long as AIs can label in domains that transfer well enough and/or can make RL envs that don’t require much labeling, that would be fine.
I briefly discussed with Dwarkesh why I’m skeptical AI progress is heavily driven by scaling up spending on human experts labeling/making data. My main argument is that spending on researchers and experiment compute seem much higher. But I didn’t say very much in the podcast.
More precisely, my view is that if spending on having human experts label/make individual data points were fixed at ~$100 million / year (per company), then AI progress would be <25% slower.
We didn’t have time to get into everything in this podcast (and some content about this recorded at an earlier point was cut), so I’ll spell out my view in a bit more detail here:
Spending directly on data (rather than on R&D about data)
isn’t growing that fast andisn’t that high (relative to spending on researchers and especially compared to spending on experiment compute).It’s important to make a distinction between spending directly on making data and science about better processes for making data. E.g., better data mixes like FineWeb count as R&D (and the person doing the R&D needs almost no understanding of individual sequences).
AI automation seems differentially good at accelerating data generation, such that I think improvements in RL environments have mostly been driven by improved AI (and R&D into making better environments) rather than spending on humans making/labeling data, and I expect this to continue. Like data stuff seems particularly amenable to acceleration from weaker AIs.
Structurally, most of what human data labeling does (though not all!) depends on having generally decent judgment rather than on having more expertise than the AIs being trained.
Transfer without domain-specific labels looks decent in practice. E.g., it doesn’t seem like Anthropic is hiring a ton of mathematicians and cyber experts to do data labeling, and the AIs are still good and getting better at these domains. Maybe this depends on having labels in some domain, but so long as AIs can label in domains that transfer well enough and/or can make RL envs that don’t require much labeling, that would be fine.