Suppose that you have “lower-level” and “higher-level” beliefs about yourself. For example, “I don’t eat oysters” is a lower-level belief than “I am a vegan”, which is a lower-level belief than ”I behave ethically”, which may be a lower-level belief than “I am a good person”.
People seem quite resistant to updates on their higher-level beliefs, in such as way where updates to their lower-level beliefs have a tendency to update “only locally”, i.e, in a way where their bearing on higher-level beliefs is explained away. (e.g, “I eat oysters, but they don’t have brains, so I’m effectively a vegan”, or “I’m not a vegan, but I still generally behave ethically” or “I don’t behave in ways that seem generally ethical, but I’m still a good person in some fundamental sense”.) This can be a problem inasmuch as this makes you resistant to actually updating against evidence coming at you from the outside world (such as feedback on your behaviors).
Anyway, maybe one of the epistemic functions of introspection is that it allows for higher-bandwidth “residual connections” between higher and lower-level beliefs. Especially if you have raw observational data on the history of your lower-level beliefs (including about why you did various things), you can hold these fixed as data to fit against to update your higher-level beliefs, possibly by looking at yourself as an object of inquiry and then bringing your general cognition to bear on the question.
(This also roughly goes through if you replace “beliefs about yourself” with “beliefs about your behavioral patterns”. It also goes through if instead of talking about beliefs you just talk about higher and lower-level behavioral patterns.)
Of course, the complexity of this whole setup begs the question of why it should be so hard to update your higher-order beliefs in the first place. Either there’s a good reason for that and introspection was of limited use in the ancestral environment, or human intelligence evolved in some bullshit runaway process and the human condition consists largely of countless spandrels on spandrels. I think it might be the second one.
(footnote: Here we are just assuming that “actual vegans don’t eat oysters” and that “being vegan is ethical behavior”; the facts of the matter don’t really have a bearing on my point. Perhaps I should come up with better examples.)
I think this statement is obvious to people who have really focused on scaling LLMs, but recent increases in performance have substantially been about improvements to the data mixture. (This of course explains the massive investment in procuring data and RL envs.) (There’s some involvement of faster attention mechanisms, and probably some contribution from better post-training practices.)
This also helps put in perspective the late-2024 disappointment (for consumers)/lengthening of timelines that we observed after the release of GPT-4.5. One perspective on this is that returns to pretraining scaling had tapered off, because improvements on the pretraining loss didn’t help on the datasets people had (in that particular post-training paradigm).
One thing I underappreciated is the degree to which this fact has affected frontier model development. I intuitively view pretraining as this massive undertaking, and basically dismissed claims that mid-2025 LLMs like Grok 3 could possibly have spent as much compute on post-training as pretraining. (I guess I assumed this was hype or something? Dangerous heuristic to overuse.)
But in fact I think if you take Epoch’s estimates for how much compute GPT4.5 took and you use the rough “6 * num_active_params * num_training_tokens” formula, you get that Kimi-K3 probably took a roughly similar amount of compute to pretrain, and is obviously way more capable.
Also if you just take various common benchmark scores and you look at the gap between Qwen3-8B and Qwen3-32B, as well as the gap from Qwen3-32B to Qwen3.8-27B, extrapolating that a linear increase in score corresponds to an exponential increase in size, you get that approximately 2 OOMs of scale were gained by whatever people changed from Qwen3 to Qwen3.8 (data mixture and attention layers, mostly, AFAICT). (i.e, Qwen3.8-27B is as performant as a hypothetical Qwen3-2700B.)
(This might be a very bad comparison inasmuch as Qwen3-8B is just distilled from 32B, and inasmuch as significantly more compute actually went into the 3.8 generation because of RL, etc.)
I tend to think of myself as immune to rage-baiting/click-baiting/general toxicity from social media and politics. I generally don’t engage in arguments on classic culture war topics on the internet, and I knowingly avoid consuming much news on the grounds that it will make me feel worse without inducing meaningful change.
But I recently realized that the phenomenon has slightly broader implications: presumably in any medium, outrage is just more attractive to the human brain, and conflicts are entertaining, especially ones where you can take a side or criticize both sides.
This made me realize that this issue isn’t constrained to just the forms of media I’m more explicitly cynical about. In particular, some of the content I read about culture, even if it is more nuanced, and even if it reads as following a debate, is still essentially to me about observing conflict for entertainment.
Entertainment for entertainment’s sake is fine, but I doubt if doing this kind of reading is on the Pareto frontier of “enjoy myself” and “inform myself about things in an actionable way”.
On the far end of the spectrum, I am not sure if it’s sensible to attempt to fully avoid engaging in any form of social derision/feeling outrage, even unproductive outrage. It does feel like some forms of outrage arise organically as a natural consequence of valuing things.
I think you are correct on a general level, but some conflicts are pure zero-sum waste of resources, and the more nuanced ones maybe are not? Like, if we debate about whether object-oriented programming is better than functional, perhaps as a result both sides get better at writing software. Or at least the observers get better at writing software, regardless of which style they choose.
I agree that there are internet conflicts worth participating in, for sure. This site contains a large number of them!
But the original post was mostly about the value of passively reading certain things vs certain other things for entertainment. (In the first paragraph, I separate out “arguments on classic culture war topics” as an example of the sorts of conflicts that are most likely a waste of resources.)
Is anyone else noticing that Claude (Sonnet 3.5 new, the default on claude.ai) is a lot worse at reasoning recently? In the past five days or so its rate of completely elementary reasoning mistakes, which persist despite repeated clarification in different ways, seems to have skyrocketed for me.
Maybe they are preparing for switching from merely encouraging their main model to do CoT (old technique) to a full RL-based reasoning model. I recently saw this, before the GUI aborted and said the model was over capacity:
Then it wouldn’t make sense anymore to have the non-reasoning model attempt to do CoT.
Here’s a fun thing for agent foundations people to ponder:
“That ‘you believe X for Y reason’ is itself a belief”. (Call this belief B.)
What this implies is that when you learn that Y reason is false, you’ll update against X, and, but also if you think that Y is true and learn that B is false, you’ll update against X.
This all happens at the level of your conscious awareness; often B needs to be salient to you for this update to go through.
Note that this is not quite the same thing as being like a PDG modeling X and Y. (In that case, learning that Y is false automatically updates you against X.) Instead, it’s more like “believing that a PDG which models X and Y should be correct”. A strange thing to observe is that this actually convincingly propagates into your beliefs about X.
A variation of this which shares a lot of structure but seems intuitively to remove elements of introspection is that you might start by believing Y, and then you might be convinced that “Y is a reason to believe in X”. (But really, the element is still here—once convinced, you believe that Y is a reason to believe in X, as before.)
Bob believes X, where X is “bald eagles are the greatest animal”.
Bob thinks he believes X because of Y, where Y is a bunch of facts about bald eagles.
But later, Bob discovers that really he believes X because he’s American and was taught X as a piece of tribal signaling.
Bob considers the counterfactual world where the US had adopted the rattlesnake instead of the bald eagle as a national symbol; and realizes that in that world, he would believe the rattlesnake was the best animal.
Bob examines a map of the world showing where people who believe X live, and notices that they’re almost all Americans. Notably people’s credence in X doesn’t correlate with their familiarity with Y (bald eagle facts); it correlates with their Americanness.
I actually mean this to be sort of a toy model/thought experiment of metacognition. How the hell does this actually convincingly propagate into your beliefs about X? A good theory of agent foundations should have an explanation for this.
It seems to me that empirical AI research/evals are automated enough that either labs or open-source could build a sort of “CI” for repeatedly checking a bunch of alignment-relevant evals upon every new model release.
If you manage to operationalize certain experiments well enough (or provide enough context for what they’re evaluating), one could imagine a pipeline where the same experiment is reproduced automatically with roughly the click of a button. (Or, to the extent that this isn’t totally automatable, yeah, you can reproducibly burn some amount of human effort doing the same experiment again and again, as long as this amount isn’t too much and you find someone willing to do this unglamorous work.)
This ranges from very small toy games such as this one or considerably more advanced/agentic evals.
You can also do the same thing to reproducibly evaluate the effectiveness of a bunch of alignment interventions that may depend on scale-dependent behavior. (e.g, maybe some old ideas didn’t work very well before, but they work well now with better models—seems like low hanging fruit to just check).
it’s hard to find definitive information about this basic fact about how modern RL on LLMs works:
are there any particularly clever ways of doing credit assignment for the tokens in a sequence S that resulted in high reward?
moreover, if you adopt the naive strategy of asserting that all of the tokens are equally responsible for the reward, is the actual gradient update to the model parameters mathematically equivalent to the one you’d get SFTing the model on S (possibly weighted by the reward, and possibly adjusted by GRPO)?
the followup is this: in this paper they claim that SFT’d models perform badly at something and RL’d models don’t. i can’t imagine what the difference between these things would even be, except that the RL’d models are affected by samples which are on-policy for them.
followup: after looking at the appendix i’m pretty sure the biggest distinction is that the SFT’d models in this paper are SFT’d on data that comes from entirely different models/datasets.
so not only is the data not coming from a policy which adapts during training, it’s coming from policies very different from the model’s own.
i think this by itself is enough to explain the results of the paper; i think this is a useful result but not the one i imagined upon reading the title.
i would still like to know the answer to the original question.
followup: with the benefit of another year’s hindsight i think that my initial assessment was pretty much spot on.
to answer the original questions:
people by and large don’t do any credit assignment (“process reward” in the contemporary parlance)
if you do this the gradient update is in fact the same, weighted by the GRPO-adjusted reward, except for the importance sampling ratio and PPO loss clip, which are both irrelevant if you’re training in a fully on-policy fashion
on-policyness matters a shitload, for reasons which are well-articulated by the recent https://www.beren.io/2026-07-26-How-Can-LLM-RL-Work-Despite-Information-Theoretic-Inefficiency/ , though i feel that this post is also a little bit misleading about the distinction between the form of the loss functions (really, the one bit of gradient we get from RL is, critically, very much an on-policy gradient, which is the important part)
Does anyone have a rigorous reference or primer on computer ergonomics, or ergonomics in general? It’s hard to find a reference that says with authority/solid reasoning what good ergonomics are and why, and solutions to common problems.
I didn’t do a lot of thorough research, but maybe I simply don’t know how to.
I googled around for resources, which usually leads to… I don’t know how to describe this, but short-form articles which are not very information dense and mutually contradictory, and I looked for opinions on Reddit and for an FAQ-like thing on /r/Ergonomics, which also didn’t tell me much definitive except that a) people have a variety of problems due to their variety of body shapes and b) it is a normal thing to want a desk that’s significantly lower than most desks.
I must have done some amount of Claude-querying, but it’s intensive to figure out what the root problems are here and whether there are canonical solutions to them, possibly because of the fact that the resources Claude would most easily reference are the same inadequate ones I’ve just described. I bet that it’s possible to figure this out with Claude if I go slowly and specifically enough, though.
I don’t think I found anything even approaching a central resource which claims to be comprehensive (however opinionated). Something like what they have at /r/bodyweightfitness, for example, would be excellent by the standards described here.
“Human–Computer Interaction (HCI) is no longer limited to trained software users. Today people interact with various devices such as mobile phones, tablets, and laptops. How can such interaction be made more user friendly, even when user proficiency levels vary? This book explores methods for assessing the psychological complexity of computer-based tasks. It also presents methods of qualitative and quantitative analysis of exploratory activity during interaction with a computer.”
Assessment of the Ergonomic Quality of Hand-Held Tools and Computer Input Devices ”The International Ergonomics Association (IEA) is currently developing standards for Ergonomic Quality in Design (EQUID) which primarily intends to promote ergonomics principles and the adaptation of a process approach for the development of products, work systems and services. It is important to assess the ergonomic quality of products, hand-held tools and computer input devices through working processes that represent reality. Well-designed working tools can be expected to reduce or eliminate fatigue, discomfort, accidents and health problems and they can lead to improvements in productivity and quality. Furthermore, absenteeism, job turnover and training costs can positively be influenced by the working tools and the environment. Not all these short-term and long-term issues of working tools can be quantified in pragmatically oriented ergonomic research approaches. But multi-channel electromyography, which enables the measurement of the physiological costs of the muscles involved in handling tools during standardized working tests, and subjective assessments of experienced subjects enable a reliable insight in the essential ergonomic criteria of working tools and products. In this respect it is advantageous to provide a test procedure, in which working tests can be carried out alternating both with test objects and reference models.”
There’s an argument against slowing down AI capabilities in order to focus on alignment, which goes “we have historically made technologies safe by iterating on them; we need a system that’s representative of the system we’re trying to align or else we won’t know what we’re trying to do”.
One argument against this, which you can see in today’s ACX post goes something like “AI systems are sufficiently complex/agentic/capable of scheming/potentially destructive that we need a much more complete science in order to deal with them than we do for past technologies”.
This is true, but the original argument would fail even if this were false. There’s a good chance that significant harms come from systems that are extremely similar to present ones, and present ones are poorly enough understood that we could likely spend years doing research to understand them better, and have much of it generalize to dangerous systems.
(This has been true as long as it’s been true that “there are significant empirical things that we don’t understand about contemporary AI systems”, which is like 4-12 years depending on how you count. (Of course 6-12 years ago you’d have missed significant information about the current paradigm, leaving only very general and abstract research like agent foundations and theoretical ML stuff. However, it’s also fair to say that we would love to have made more progress on those things at this point in time. Regardless of your stance towards more theoretical AIS work, the same goes for a bunch of things which today would be called “empirical alignment”, so the statement is certainly true now.)
Strangely, if you ask Claude “what’s the complexity of inversion of a triangular matrix?”, it says O(n^2), but when pressed to explain will say O(n^3).
This is one of those things that’s pretty elementary, and well-represented in the training set, so I’d expect it to be very well-memorized by an LLM.
This is robust to a few attempts and rewords, and held for Sonnet 4.6 and Opus 4.6.
forgive my ignorance, but is there any reason that you can’t have multi-layer sparse autoencoders, even those that are interpretably compatible with the linear representation hypothesis? like what would their drawbacks be (other than more required compute)?
no matter how it is that you’re computing the latents,
0) you still have a reconstruction loss;
the final layer is still a decoder that computes a linear transformation to the thing you’re reconstructing;
you can still have a sparsity penalty on the last set of activations.
it seems to me like this still constructs a set of latents that sparsely activate, and which are linearly represented in activation space
Does anyone have a sense of whether, qualitatively, RL stability has been solved for any practical domains?
This question is at least in part asking for qualitative speculation about how the post-training RL works at big labs, but I’m interested in any partial answer people can come up with.
My impression of RL is that there are a lot of tricks to “improve stability”, but performance is path-dependent in pretty much any realistic/practical setting (where state space is huge and action space may be huge or continuous). Even for larger toy problems my sense is that various RL algorithms really only work like up to 70% of the time, and 30% of the time they randomly decline in reward.
One obvious way of getting around this is to just resample. If there are no more principled/reliable methods, this would be the default method of getting a good result from RL. It would follow that that’s just what the big labs do. But of course they may have secretly solved some stuff, but it’s hard to imagine what the form of that would be.
is there something i’m missing here? my strong impression is that the scale of the nondeterminism of the result is quite small, and random in direction, so that it isn’t likely to affect an aggregate-scale thing like the qualitative effect of an entire gradient update. (i can imagine that the accumulation of many random errors does bias the policy towards being generally less stable, which implies qualitatively worse, yes...)
without something that mitigates the statement above, my prior is instead that the graph is cherry-picked, intentionally or not, to increase the perceived importance of llm determinism.
with the benefit of hindsight/some more broad-literature understanding of RL, i now understand that modern RL is ridiculously unstable and even numerical issue-level mismatches between rollout and fsdp policies can cause RL to be unstable; that’s why this mattered.
Suppose that you have “lower-level” and “higher-level” beliefs about yourself. For example, “I don’t eat oysters” is a lower-level belief than “I am a vegan”, which is a lower-level belief than ”I behave ethically”, which may be a lower-level belief than “I am a good person”.
People seem quite resistant to updates on their higher-level beliefs, in such as way where updates to their lower-level beliefs have a tendency to update “only locally”, i.e, in a way where their bearing on higher-level beliefs is explained away. (e.g, “I eat oysters, but they don’t have brains, so I’m effectively a vegan”, or “I’m not a vegan, but I still generally behave ethically” or “I don’t behave in ways that seem generally ethical, but I’m still a good person in some fundamental sense”.) This can be a problem inasmuch as this makes you resistant to actually updating against evidence coming at you from the outside world (such as feedback on your behaviors).
Anyway, maybe one of the epistemic functions of introspection is that it allows for higher-bandwidth “residual connections” between higher and lower-level beliefs. Especially if you have raw observational data on the history of your lower-level beliefs (including about why you did various things), you can hold these fixed as data to fit against to update your higher-level beliefs, possibly by looking at yourself as an object of inquiry and then bringing your general cognition to bear on the question.
(This also roughly goes through if you replace “beliefs about yourself” with “beliefs about your behavioral patterns”. It also goes through if instead of talking about beliefs you just talk about higher and lower-level behavioral patterns.)
Of course, the complexity of this whole setup begs the question of why it should be so hard to update your higher-order beliefs in the first place. Either there’s a good reason for that and introspection was of limited use in the ancestral environment, or human intelligence evolved in some bullshit runaway process and the human condition consists largely of countless spandrels on spandrels. I think it might be the second one.
(footnote: Here we are just assuming that “actual vegans don’t eat oysters” and that “being vegan is ethical behavior”; the facts of the matter don’t really have a bearing on my point. Perhaps I should come up with better examples.)
I think this statement is obvious to people who have really focused on scaling LLMs, but recent increases in performance have substantially been about improvements to the data mixture. (This of course explains the massive investment in procuring data and RL envs.) (There’s some involvement of faster attention mechanisms, and probably some contribution from better post-training practices.)
This also helps put in perspective the late-2024 disappointment (for consumers)/lengthening of timelines that we observed after the release of GPT-4.5. One perspective on this is that returns to pretraining scaling had tapered off, because improvements on the pretraining loss didn’t help on the datasets people had (in that particular post-training paradigm).
One thing I underappreciated is the degree to which this fact has affected frontier model development. I intuitively view pretraining as this massive undertaking, and basically dismissed claims that mid-2025 LLMs like Grok 3 could possibly have spent as much compute on post-training as pretraining. (I guess I assumed this was hype or something? Dangerous heuristic to overuse.)
But in fact I think if you take Epoch’s estimates for how much compute GPT4.5 took and you use the rough “6 * num_active_params * num_training_tokens” formula, you get that Kimi-K3 probably took a roughly similar amount of compute to pretrain, and is obviously way more capable.
Also if you just take various common benchmark scores and you look at the gap between Qwen3-8B and Qwen3-32B, as well as the gap from Qwen3-32B to Qwen3.8-27B, extrapolating that a linear increase in score corresponds to an exponential increase in size, you get that approximately 2 OOMs of scale were gained by whatever people changed from Qwen3 to Qwen3.8 (data mixture and attention layers, mostly, AFAICT). (i.e, Qwen3.8-27B is as performant as a hypothetical Qwen3-2700B.) (This might be a very bad comparison inasmuch as Qwen3-8B is just distilled from 32B, and inasmuch as significantly more compute actually went into the 3.8 generation because of RL, etc.)
I tend to think of myself as immune to rage-baiting/click-baiting/general toxicity from social media and politics. I generally don’t engage in arguments on classic culture war topics on the internet, and I knowingly avoid consuming much news on the grounds that it will make me feel worse without inducing meaningful change.
But I recently realized that the phenomenon has slightly broader implications: presumably in any medium, outrage is just more attractive to the human brain, and conflicts are entertaining, especially ones where you can take a side or criticize both sides.
This made me realize that this issue isn’t constrained to just the forms of media I’m more explicitly cynical about. In particular, some of the content I read about culture, even if it is more nuanced, and even if it reads as following a debate, is still essentially to me about observing conflict for entertainment.
Entertainment for entertainment’s sake is fine, but I doubt if doing this kind of reading is on the Pareto frontier of “enjoy myself” and “inform myself about things in an actionable way”.
On the far end of the spectrum, I am not sure if it’s sensible to attempt to fully avoid engaging in any form of social derision/feeling outrage, even unproductive outrage. It does feel like some forms of outrage arise organically as a natural consequence of valuing things.
I think you are correct on a general level, but some conflicts are pure zero-sum waste of resources, and the more nuanced ones maybe are not? Like, if we debate about whether object-oriented programming is better than functional, perhaps as a result both sides get better at writing software. Or at least the observers get better at writing software, regardless of which style they choose.
I agree that there are internet conflicts worth participating in, for sure. This site contains a large number of them!
But the original post was mostly about the value of passively reading certain things vs certain other things for entertainment. (In the first paragraph, I separate out “arguments on classic culture war topics” as an example of the sorts of conflicts that are most likely a waste of resources.)
Is anyone else noticing that Claude (Sonnet 3.5 new, the default on claude.ai) is a lot worse at reasoning recently? In the past five days or so its rate of completely elementary reasoning mistakes, which persist despite repeated clarification in different ways, seems to have skyrocketed for me.
Maybe they are preparing for switching from merely encouraging their main model to do CoT (old technique) to a full RL-based reasoning model. I recently saw this, before the GUI aborted and said the model was over capacity:
Then it wouldn’t make sense anymore to have the non-reasoning model attempt to do CoT.
I have also seen this.
Here’s a fun thing for agent foundations people to ponder:
“That ‘you believe X for Y reason’ is itself a belief”. (Call this belief B.) What this implies is that when you learn that Y reason is false, you’ll update against X, and, but also if you think that Y is true and learn that B is false, you’ll update against X. This all happens at the level of your conscious awareness; often B needs to be salient to you for this update to go through. Note that this is not quite the same thing as being like a PDG modeling X and Y. (In that case, learning that Y is false automatically updates you against X.) Instead, it’s more like “believing that a PDG which models X and Y should be correct”. A strange thing to observe is that this actually convincingly propagates into your beliefs about X.
A variation of this which shares a lot of structure but seems intuitively to remove elements of introspection is that you might start by believing Y, and then you might be convinced that “Y is a reason to believe in X”. (But really, the element is still here—once convinced, you believe that Y is a reason to believe in X, as before.)
Am I understanding the argument correctly?
Bob believes X, where X is “bald eagles are the greatest animal”.
Bob thinks he believes X because of Y, where Y is a bunch of facts about bald eagles.
But later, Bob discovers that really he believes X because he’s American and was taught X as a piece of tribal signaling.
Bob considers the counterfactual world where the US had adopted the rattlesnake instead of the bald eagle as a national symbol; and realizes that in that world, he would believe the rattlesnake was the best animal.
Bob examines a map of the world showing where people who believe X live, and notices that they’re almost all Americans. Notably people’s credence in X doesn’t correlate with their familiarity with Y (bald eagle facts); it correlates with their Americanness.
This causes Bob to update away from X.
Yes, this is an instance of what I’m talking about.
This strange observation is imo one basis upon which you can claim that metacognition is a convergent abstraction :)
I actually mean this to be sort of a toy model/thought experiment of metacognition. How the hell does this actually convincingly propagate into your beliefs about X? A good theory of agent foundations should have an explanation for this.
It seems to me that empirical AI research/evals are automated enough that either labs or open-source could build a sort of “CI” for repeatedly checking a bunch of alignment-relevant evals upon every new model release.
If you manage to operationalize certain experiments well enough (or provide enough context for what they’re evaluating), one could imagine a pipeline where the same experiment is reproduced automatically with roughly the click of a button. (Or, to the extent that this isn’t totally automatable, yeah, you can reproducibly burn some amount of human effort doing the same experiment again and again, as long as this amount isn’t too much and you find someone willing to do this unglamorous work.)
This ranges from very small toy games such as this one or considerably more advanced/agentic evals.
You can also do the same thing to reproducibly evaluate the effectiveness of a bunch of alignment interventions that may depend on scale-dependent behavior. (e.g, maybe some old ideas didn’t work very well before, but they work well now with better models—seems like low hanging fruit to just check).
This point has been kind of made before https://www.lesswrong.com/posts/oKxc8maZGtnzgpNzx/rerunning-ai-safety-papers-on-every-frontier-release-would-1
it’s hard to find definitive information about this basic fact about how modern RL on LLMs works:
are there any particularly clever ways of doing credit assignment for the tokens in a sequence S that resulted in high reward?
moreover, if you adopt the naive strategy of asserting that all of the tokens are equally responsible for the reward, is the actual gradient update to the model parameters mathematically equivalent to the one you’d get SFTing the model on S (possibly weighted by the reward, and possibly adjusted by GRPO)?
the followup is this: in this paper they claim that SFT’d models perform badly at something and RL’d models don’t. i can’t imagine what the difference between these things would even be, except that the RL’d models are affected by samples which are on-policy for them.
https://arxiv.org/pdf/2507.00432
followup: after looking at the appendix i’m pretty sure the biggest distinction is that the SFT’d models in this paper are SFT’d on data that comes from entirely different models/datasets. so not only is the data not coming from a policy which adapts during training, it’s coming from policies very different from the model’s own. i think this by itself is enough to explain the results of the paper; i think this is a useful result but not the one i imagined upon reading the title.
i would still like to know the answer to the original question.
followup: with the benefit of another year’s hindsight i think that my initial assessment was pretty much spot on.
to answer the original questions:
people by and large don’t do any credit assignment (“process reward” in the contemporary parlance)
if you do this the gradient update is in fact the same, weighted by the GRPO-adjusted reward, except for the importance sampling ratio and PPO loss clip, which are both irrelevant if you’re training in a fully on-policy fashion
on-policyness matters a shitload, for reasons which are well-articulated by the recent https://www.beren.io/2026-07-26-How-Can-LLM-RL-Work-Despite-Information-Theoretic-Inefficiency/ , though i feel that this post is also a little bit misleading about the distinction between the form of the loss functions (really, the one bit of gradient we get from RL is, critically, very much an on-policy gradient, which is the important part)
Does anyone have a rigorous reference or primer on computer ergonomics, or ergonomics in general? It’s hard to find a reference that says with authority/solid reasoning what good ergonomics are and why, and solutions to common problems.
I’d be interested in this myself. Where/how have you looked so far, and which resources have you found wanting so far?
I didn’t do a lot of thorough research, but maybe I simply don’t know how to.
I googled around for resources, which usually leads to… I don’t know how to describe this, but short-form articles which are not very information dense and mutually contradictory, and I looked for opinions on Reddit and for an FAQ-like thing on /r/Ergonomics, which also didn’t tell me much definitive except that a) people have a variety of problems due to their variety of body shapes and b) it is a normal thing to want a desk that’s significantly lower than most desks.
I must have done some amount of Claude-querying, but it’s intensive to figure out what the root problems are here and whether there are canonical solutions to them, possibly because of the fact that the resources Claude would most easily reference are the same inadequate ones I’ve just described. I bet that it’s possible to figure this out with Claude if I go slowly and specifically enough, though.
I don’t think I found anything even approaching a central resource which claims to be comprehensive (however opinionated). Something like what they have at /r/bodyweightfitness, for example, would be excellent by the standards described here.
These are the first things I found on the first search result page of GoodReads, do these suite?
Applying Systemic-Structural Activity Theory to Design of Human-Computer Interaction Systems
“Human–Computer Interaction (HCI) is no longer limited to trained software users. Today people interact with various devices such as mobile phones, tablets, and laptops. How can such interaction be made more user friendly, even when user proficiency levels vary? This book explores methods for assessing the psychological complexity of computer-based tasks. It also presents methods of qualitative and quantitative analysis of exploratory activity during interaction with a computer.”
Assessment of the Ergonomic Quality of Hand-Held Tools and Computer Input Devices
”The International Ergonomics Association (IEA) is currently developing standards for Ergonomic Quality in Design (EQUID) which primarily intends to promote ergonomics principles and the adaptation of a process approach for the development of products, work systems and services. It is important to assess the ergonomic quality of products, hand-held tools and computer input devices through working processes that represent reality. Well-designed working tools can be expected to reduce or eliminate fatigue, discomfort, accidents and health problems and they can lead to improvements in productivity and quality. Furthermore, absenteeism, job turnover and training costs can positively be influenced by the working tools and the environment. Not all these short-term and long-term issues of working tools can be quantified in pragmatically oriented ergonomic research approaches. But multi-channel electromyography, which enables the measurement of the physiological costs of the muscles involved in handling tools during standardized working tests, and subjective assessments of experienced subjects enable a reliable insight in the essential ergonomic criteria of working tools and products. In this respect it is advantageous to provide a test procedure, in which working tests can be carried out alternating both with test objects and reference models.”
There’s an argument against slowing down AI capabilities in order to focus on alignment, which goes “we have historically made technologies safe by iterating on them; we need a system that’s representative of the system we’re trying to align or else we won’t know what we’re trying to do”.
One argument against this, which you can see in today’s ACX post goes something like “AI systems are sufficiently complex/agentic/capable of scheming/potentially destructive that we need a much more complete science in order to deal with them than we do for past technologies”.
This is true, but the original argument would fail even if this were false. There’s a good chance that significant harms come from systems that are extremely similar to present ones, and present ones are poorly enough understood that we could likely spend years doing research to understand them better, and have much of it generalize to dangerous systems.
(This has been true as long as it’s been true that “there are significant empirical things that we don’t understand about contemporary AI systems”, which is like 4-12 years depending on how you count. (Of course 6-12 years ago you’d have missed significant information about the current paradigm, leaving only very general and abstract research like agent foundations and theoretical ML stuff. However, it’s also fair to say that we would love to have made more progress on those things at this point in time. Regardless of your stance towards more theoretical AIS work, the same goes for a bunch of things which today would be called “empirical alignment”, so the statement is certainly true now.)
Strangely, if you ask Claude “what’s the complexity of inversion of a triangular matrix?”, it says O(n^2), but when pressed to explain will say O(n^3).
This is one of those things that’s pretty elementary, and well-represented in the training set, so I’d expect it to be very well-memorized by an LLM.
This is robust to a few attempts and rewords, and held for Sonnet 4.6 and Opus 4.6.
forgive my ignorance, but is there any reason that you can’t have multi-layer sparse autoencoders, even those that are interpretably compatible with the linear representation hypothesis? like what would their drawbacks be (other than more required compute)?
no matter how it is that you’re computing the latents, 0) you still have a reconstruction loss;
the final layer is still a decoder that computes a linear transformation to the thing you’re reconstructing;
you can still have a sparsity penalty on the last set of activations.
it seems to me like this still constructs a set of latents that sparsely activate, and which are linearly represented in activation space
Does anyone have a sense of whether, qualitatively, RL stability has been solved for any practical domains?
This question is at least in part asking for qualitative speculation about how the post-training RL works at big labs, but I’m interested in any partial answer people can come up with.
My impression of RL is that there are a lot of tricks to “improve stability”, but performance is path-dependent in pretty much any realistic/practical setting (where state space is huge and action space may be huge or continuous). Even for larger toy problems my sense is that various RL algorithms really only work like up to 70% of the time, and 30% of the time they randomly decline in reward.
One obvious way of getting around this is to just resample. If there are no more principled/reliable methods, this would be the default method of getting a good result from RL. It would follow that that’s just what the big labs do. But of course they may have secretly solved some stuff, but it’s hard to imagine what the form of that would be.
with the benefit of a bit of hindsight, i should remark that intuitions built from classical RL (from scratch, highly off-policy) don’t carry over to LLMs that well, and plausibly LLM RL is a lot more stable/consistent because of essential differences, as articulated in https://www.beren.io/2026-07-26-How-Can-LLM-RL-Work-Despite-Information-Theoretic-Inefficiency/
at the end of the somewhat famous blogpost about llm nondeterminism recently https://thinkingmachines.ai/blog/defeating-nondeterminism-in-llm-inference/ they assert that the determinism is enough to make an rlvr run more stable without importance sampling.
is there something i’m missing here? my strong impression is that the scale of the nondeterminism of the result is quite small, and random in direction, so that it isn’t likely to affect an aggregate-scale thing like the qualitative effect of an entire gradient update. (i can imagine that the accumulation of many random errors does bias the policy towards being generally less stable, which implies qualitatively worse, yes...)
without something that mitigates the statement above, my prior is instead that the graph is cherry-picked, intentionally or not, to increase the perceived importance of llm determinism.
with the benefit of hindsight/some more broad-literature understanding of RL, i now understand that modern RL is ridiculously unstable and even numerical issue-level mismatches between rollout and fsdp policies can cause RL to be unstable; that’s why this mattered.