Daniel Paleka

Karma: 404

Daniel Paleka 13 Nov 2025 6:50 UTC
1 point
0
on: Daniel Paleka’s Shortform
AIs being bad at AI research says nothing about acceleration
Lots of people are trying to make AI good at AI research. How are they doing?
One way to measure this is to assume AIs are gradually doing more and more complex tasks independently. Eventually it would grow to doing whole research projects. Something like this, but for software engineering instead of research, is captured in the METR “time horizons” benchmark.
I think extending this line of thinking to forecasting progress in AI research is wrong. Instead, a better way to accelerate AI research for the time being is combining AI and people to do research together, in a way that uses the complementary strengths of each; with the goal of the researcher’s feedback loops shortening.
What is difficult to automate in AI research?
If you decompose a big AI research project into tasks, there’s lots of “dark matter” that does not neatly fit into any category. Some examples are given in Large-Scale Projects Stress Deep Cognitive Skills, which is a much better post than mine.
But I think that the most central argument is: the research process involves taste, coming up with ideas, and various such intangibles that we don’t really know how to train for.
The labs are trying to make superhuman AI researchers. We do not yet know how to do it, which means at least some of our ideas are lacking. To improve our ideas, we need either:
1. (the proper way) conceptual advances in machine learning;
2. (the way it’s actually going to get done) reinforcement learning on the idea->code->experiment->result process, to figure out which ideas are good.
Measuring which ideas are good is difficult; it requires sparse empirical outcomes that happen long after the idea is formulated. How can we accelerate this process?
I want to make two claims:
1. There are large gains from accelerating AI researchers.
2. Much more importantly, those gains are achievable without inventing new things in machine learning.
The careful reader might ask, ok, this sounds fine in the abstract, but I don’t understand what exactly the lab is doing then, if not “automate AI research as a whole”? How is this different from making autonomous AI researchers directly?
Here is a list of tasks that would be extremely valuable if we wanted to make the research feedback loops faster.
- implementing instructions of varying level of detail into code efficiently;
- extrapolating user intent and implementing the correct thing;
- relatedly: learning user intent from working with a researcher over time;
- given code, running experiments autonomously, fixing minor deployment issues;
- checking for bugs and suspicious logic in the code;
- observing and pinpointing anomalies in the data;
- monitoring experiments, reporting updates, and raising alarm when something is off; and so on.
I believe all of these tasks possess properties that make them attractive to attack directly.
1. They consume a significant amount of time of a researcher (or add communication overhead if we add people focused on research engineering);
2. There are clear ways to generate many datapoints + labels / reward functions for each of these tasks;
3. Alternatively, these are done by people typing into a keyboard, so labs can do imitation learning by collecting all the human actions from their own researchers.
This seems easier than automating the full research process. If labs have the goal of speeding up the lab’s ability to do AI research as opposed to other goals, they are probably doing these things; and measuring the ability of AIs to do research autonomously is not going to give a good grasp on how quickly the lab is accelerating.

Daniel Paleka 8 Nov 2025 2:20 UTC
101 points
30
on: Daniel Paleka’s Shortform
Slow takeoff for AI R&D, fast takeoff for everything else
Why is AI progress so much more apparent in coding than everywhere else?
Among people who have “AGI timelines”, most do not set their timelines based on data, but rather update them based on their own day-to-day experiences and social signals.
As of 2025, my guess is that individual perception of AI progress correlates with how closely someone’s daily activities resemble how an AI researcher spends their time. The reason why users of coding agents feel a higher rate of automation in their bones, whereas people in most other occupations don’t, is because automating engineering has been the focus of the industry for a while now. Despite the expectations for 2025 to be the year of the AI agent, it turns out the industry is small and cannot have too many priorities, hence basically the only competent agents we got in 2025 so far are coding agents.
Everyone serious about winning the AI race is trying to automate one job: AI R&D.
To a first approximation, there is no point yet in automating anything else, except to raise capital (human or investment), or to earn money. Until you are hitting diminishing returns on your rate of acceleration, unrelated capabilities are not a priority. This means that a lot of pressure is being applied to AI research tasks at all times; and that all delays in automation of AI R&D are, in a sense, real in a way that’s not necessarily the case for tasks unrelated to AI R&D. It would be odd if there were easy gains to be made in accelerating the work of AI researchers on frontier models in addition to what is already being done across the industry.
I don’t know whether automating AI research is going to be smooth all the way there or not; my understanding is that slow vs fast takeoff hinges significantly on how bottlenecked we become by non-R&D factors over time. Nonetheless, the above suggests a baseline expectation: AI research automation will advance more steadily compared to automation of other intellectual work.
For other tasks, especially for less immediately lucrative ones, it will make more sense to automate them quickly after we’re done with automating AI research. Hence, a teacher’s or a fiction writer’s experience of automation will be somewhat more abrupt than a researcher’s. In particular, I anticipate there will be a period of a year or two in which publicly available models are severely underelicited in tasks unrelated to AI R&D, as top talent is increasingly incentivized to work on capabilities that compound in R&D value.
This “differential automation” view naturally separates the history of AI capabilities into three phases:
- intentional, pre-scaling: we train the AI on a specific dataset, or even write the code for a specific task
- unintentional (2019-2023; scaling on broad data on the internet; AI improves across the board)
- intentional again: the AIs improve being finetuned on carefully sourced data, and RL environments
There will likely be another phase, after say a GPT-3 moment for RL, where RL is going to generalize somewhat further, and we will get gains on tasks that we do not directly train for; but I think the sheer amount of “unintentional” increase of capabilities across the board is less likely, because the remaining capabilities are inherently more specialized and unrelated to each other than they were in the pretraining scaling phase.

A/B testing could lead LLMs to retain users instead of helping them

Daniel Paleka4 Nov 2025 19:30 UTC

28 points

0 comments4 min readLW link

(newsletter.danielpaleka.com)

Daniel Paleka’s Shortform

Daniel Paleka31 Mar 2025 10:10 UTC

4 points

3 comments1 min readLW link

Daniel Paleka 31 Mar 2025 10:10 UTC
3 points
0
on: Daniel Paleka’s Shortform
GPT-4o’s drawings of itself as a person are remarkably consistent: it’s more or less always a similar-looking white male in his late 20s with brown hair, often sporting facial hair and glasses, unless you specify otherwise. All the men it generates might as well be brothers. I reproduced this on two ChatGPT accounts with clean memory.
On the contrary, its drawings of itself when it does not depict itself as a person are far more diverse: a wide range of robot designs and abstract humanoids, often featuring OpenAI logo as a head or on the word “GPT” on the chest.

Daniel Paleka 22 Mar 2025 19:25 UTC
3 points
0
in reply to: nostalgebraist’s comment on: What are the strongest arguments for very short timelines?
I think the labs might well be rational in focusing on this sort of “handheld automation”, just to enable their researchers to code experiments faster and in smaller teams.
My mental model of AI R&D is that it can be bottlenecked roughly by three things: compute, engineering time, and the “dark matter” of taste and feedback loops on messy research results. I can certainly imagine a model of lab productivity where the best way to accelerate is improving handheld automation for the entirety of 2025. Say, the core paradigm is fixed; but inside that paradigm, the research team has more promising ideas than they have time to implement and try out on smaller-scale experiments; and they really do not want to hire more people.
If you consider the AI lab as a fundamental unit that wants to increase its velocity, and works on things that make models faster, it’s plausible they can be aware how bad the model performance is on research taste, and still not be making a mistake by ignoring your “dark matter” right now. They will work on it when they are faster.

You should delay engineering-heavy research in light of R&D automation

Daniel Paleka7 Jan 2025 2:11 UTC

42 points

3 comments5 min readLW link

(newsletter.danielpaleka.com)

Daniel Paleka 18 Aug 2024 5:47 UTC
1 point
0
in reply to: KhromeM’s comment on: Using an LLM perplexity filter to detect weight exfiltration
N = #params, D = #data
Training compute = const .* N * D
Forward pass cost (R bits) = c * N, and assume R = Ω(1) on average
Now, thinking purely information-theoretically:
Model stealing compute = C * fp16 * N / R ~ const. * c * N^2
If compute-optimal training and α = β in Chinchilla scaling law:
Model stealing compute ~ Training compute
For significantly overtrained models:
Model stealing << Training compute
Typically:
Total inference compute ~ Training compute
=> Model stealing << Total inference compute
Caveats:
- Prior on weights reduces stealing compute, same if you only want to recover some information about the model (e.g. to create an equally capable one)
- Of course, if the model is producing much fewer than 1 token per forward pass, then model stealing compute is very large

Daniel Paleka 21 Jan 2024 0:53 UTC
2 points
1
in reply to: Adrià Garriga-alonso’s comment on: Does literacy remove your ability to be a bard as good as Homer?
The one you linked doesn’t really rhyme. The meter is quite consistently decasyllabic, though.
I find it interesting that the collection has a fairly large number of songs about World War II. Seems that the “oral songwriters composing war epics” meme lived until the very end of the tradition.

Daniel Paleka 14 Jan 2024 10:12 UTC
2 points
0
on: Takeaways from the NeurIPS 2023 Trojan Detection Competition

With Greedy Coordinate Gradient (GCG) optimization, when trying to force argmax-generated completions, using an improved objective function dramatically increased our optimizer’s performance.

Do you have some data / plots here?

Daniel Paleka 15 Nov 2023 19:41 UTC
2 points
1
in reply to: lberglund’s comment on: Paper: LLMs trained on “A is B” fail to learn “B is A”
Oh so you have prompt_loss_weight=1, got it. I’ll cross out my original comment. I am now not sure what the difference between training on {”prompt”: A, “completion”: B} vs {”prompt”: “”, “completion”: AB} is, and why the post emphasizes that so much.

Daniel Paleka 15 Nov 2023 19:32 UTC
1 point
0
in reply to: ryan_greenblatt’s comment on: Paper: LLMs trained on “A is B” fail to learn “B is A”
The key adjustment in this post is that they train on the entire sequence
Yeah, but my understanding of the post is that it wasn’t enough; it only worked out when A was Tom Cruise, not Uriah Hawthorne. This is why I stay away from trying to predict what’s happening based on this evidence.
Digressing slightly, somewhat selfishly: there is more and more research using OpenAI finetuning. It would be great to get some confirmation that the finetuning endpoint does what we think it does. Unlike with the model versions, there are no guarantees on the finetuning endpoint being stable over time; they could introduce a p(A | B) term when finetuning on {”prompt”: A, “completion”: B} at any time if it improved performance, and experiments like this would then go to waste.

Daniel Paleka 15 Nov 2023 18:37 UTC
7 points
0
in reply to: gwern’s comment on: Paper: LLMs trained on “A is B” fail to learn “B is A”
So there’s a post that claims p(A | B) is sometimes learned from p(B | A) if you make the following two adjustments to the finetuning experiments in the paper:
~~(1) you finetune not on p(B | A), but p(A) + p(B | A) instead~~ finetune on p(AB) in the completion instead of finetuning on p(A) in the prompt + p(B | A) in the completion, as in Berglund et al.
(2) A is a well-known name (“Tom Cruise”), but B is still a made-up thing

~~The post is not written clearly, but this is what I take from it. Not sure how model internals explain this.~~
~~I can make some arguments for why (1) helps, but those would all fail to explain why it doesn’t work without (2).~~
Caveat: The experiments in the post are only on A=”Tom Cruise” and gpt-3.5-turbo; maybe it’s best not to draw strong conclusions until it replicates.

Daniel Paleka 5 Oct 2023 14:00 UTC
3 points
0
in reply to: 1a3orn’s comment on: What evidence is there of LLM’s containing world models?
I made an illegal move while playing over the board (5+3 blitz) yesterday and lost the game. Maybe my model of chess (even when seeing the current board state) is indeed questionable, but well, it apparently happens to grandmasters in blitz too.

Daniel Paleka 24 Aug 2023 9:39 UTC
1 point
0
on: Reducing sycophancy and improving honesty via activation steering
Do the modified activations “stay in the residual stream” for the next token forward pass?
Is there any difference if they do or don’t?
If I understand the method correctly, in Steering GPT-2-XL by adding an activation vector they always added the steering vectors on the same (token, layer) coordinates, hence in their setting this distinction doesn’t matter. However, if the added vector is on (last_token, layer), then there seems to be a difference.

Daniel Paleka 22 Aug 2023 8:49 UTC
1 point
0
in reply to: james387’s comment on: Evaluating Superhuman Models with Consistency Checks
Thank you for the discussion in the DMs!

Wrt superhuman doubts: The models we tested are superhuman. https://www.melonimarco.it/en/2021/03/08/stockfish-and-lc0-test-at-different-number-of-nodes/ gave a rough human ELO estimate of 3000 for a 2021 version of Leela with just 100 nodes, 3300 for 1000 nodes. There is a bot on Lichess that plays single-node (no search at all) and seems to be in top 0.1% of players.
I asked some Leela contributors; they say that it’s likely new versions of Leela are superhuman at even 20 nodes; and that our tests of 100-1600 nodes are almost certainly quite superhuman. We also tested Stockfish NNUE with 80k nodes and Stockfish classical with 4e6 nodes, with similar consistency results.
Table 5 in Appendix B.3 (“Comparison of the number of failures our method finds in increasingly stronger models”): this is all on positions from Master-level games. The only synthetically generated positions are for the Board transformation check, as no-pawn positions with lots of pieces are rare in human games.
We cannot comment on different setups not reproducing our results exactly; pairs of positions do not necessarily transfer between versions, but iirc preliminary exploration implied that the results wouldn’t be qualitatively different. Maybe we’ll do a proper experiment to confirm.
There’s an important question to ask here: how much does scaling search help consistency? Scaling Scaling Laws with Board Games [Jones, 2021] is the standard reference, but I don’t see how to convert their predictions to estimates here. We found one halving of in-distribution inconsistency ratio with two doublings of search nodes on the Recommended move check. Not sure if anyone will be working on any version of this soon (FAR AI maybe?). I’d be more interested in doing a paper on this if I could wrap my head around how to scale “search” in LLMs, with a similar effect as what increasing the number of search nodes does on MCTS trained models.

Daniel Paleka 9 Aug 2023 5:50 UTC
4 points
1
on: Model Organisms of Misalignment: The Case for a New Pillar of Alignment Research
It would be helpful to write down where the Scientific Case and the Global Coordination Case objectives might be in conflict. The “Each subcomponent” section addresses some of the differences, but not the incentives. I do acknowledge that first steps look very similar right now, but the objectives might diverge at some point. It naively seems that demonstrating things that are scary might be easier and is not the same thing as creating examples which usefully inform alignment of superhuman models.

Evaluating Superhuman Models with Consistency Checks

Daniel Paleka and Lukas Fluri

1 Aug 2023 7:51 UTC

21 points

2 comments9 min readLW link

(arxiv.org)

Daniel Paleka 25 May 2023 21:57 UTC
1 point
0
in reply to: Richard_Ngo’s comment on: Coercion is an adaptation to scarcity; trust is an adaptation to abundance
So I’ve read an overview ^[1] which says Chagnon observed a pre-Malthusian group of people, which was kept from exponentially increasing not by scarcity of resources, but by sheer competitive violence; a totalitarian society that lives in abundance.

There seems to be an important scarcity factor shaping their society, but not of the kind where we could say that “we only very recently left the era in which scarcity was the dominant feature of people’s lives.”

Although, reading again, this doesn’t disprove violence in general arising due to scarcity, and then misgeneralizing in abundant environments… And again, “violence” is not the same as “coercion”.
1. ^
  Unnecessarily political, but seems to accurately represent Chagnon’s observations, based on other reporting and a quick skim of Chagnon’s work on Google Books.

Daniel Paleka 25 May 2023 18:58 UTC
1 point
0
on: Coercion is an adaptation to scarcity; trust is an adaptation to abundance
I don’t think “coercion is an evolutionary adaptation to scarcity, and we’ve only recently managed to get rid of the scarcity” is clearly true. It intuitively makes sense, but Napoleon Chagnon’s research seems to be one piece of evidence against the theory.

Daniel Paleka

AIs being bad at AI research says nothing about acceleration

Slow takeoff for AI R&D, fast takeoff for everything else

A/​B test­ing could lead LLMs to re­tain users in­stead of helping them

Daniel Paleka’s Shortform

You should de­lay en­g­ineer­ing-heavy re­search in light of R&D automation

Eval­u­at­ing Su­per­hu­man Models with Con­sis­tency Checks

A/B testing could lead LLMs to retain users instead of helping them

You should delay engineering-heavy research in light of R&D automation

Evaluating Superhuman Models with Consistency Checks