I don’t know if this affects your ideas about machine learning, but I don’t think that selection for selectability is a major force in most contexts in biotic evolution. (I’m also not sure whether you’d disagree with this; the post seems to emphasize selection for selectability, but your specific example with the two lineages may be appropriately caveated, not sure.)
There are some forces that prevent it from being a major force, such as:
If a mutation is first-order deleterious but second-order beneficial (that is, it leads to more beneficial variation to be selected on over generations, hence faster evolution), the first-order effects probably are stronger because they act faster (coarsely speaking, every generation vs. [the number of generations it takes for a beneficial mutation to reach fixation]). (Exception: strong linkage.)
If a mutation is first-order deleterious but second-order beneficial, AND there is population mixing as opposed to two fully separated lineages, then organisms without the first-order deleterious selectability-increasing mutation can benefit from the increasingly-frequent follow-on mutations because they can get them through sexual reproduction with carriers, if you see what I mean.
There are ways these obstacles can be kinda sidestepped. E.g. the selectability-improving variant may be in tight linkage with the variants it affects; or similarly, it may be kinda structurally bound up with its variants (I mean, for example, novel paralogs, i.e. gene duplication); it may be itself first-order neutral or beneficial. But the underlying point that still stands is that the “real reason” / the crux for the selectability-improving variant being favored is usually not mainly the fact that it improves selectability of other variation, but rather that it is directly beneficial itself. Though, one could be interested in studying cases where directly-beneficial mutations “coincidentally” seem to also increase selectability. (Exceptions could be found with strong selection forces in asexually reproducing organisms (without too much horizontal transfer?), which may be especially analogous to ML.)
Yes I think selection directly for evolvability to the detriment of present fitness is at most a weird effect that happens in weird regimes, analogous to what I understand is the consensus for group selection. I remember reading a paper where they set the mutation rate in such a way that it was lineage-optimal for individuals to make certain mutations catastrophic, but this was in silico and probably “never” occurs irl.
Instead I have in mind a thought like “given a need to develop selectability—the channeling of mutations into ‘useful’ exploration directions—what must be true about the optimizer in order for it to do so without conflicting with the ineluctable short-term imperative of local-optimality”.
This then has answers like “have lots of neutral directions in genome space so that there’s a ton of ways to have the same phenotype while still permitting random motion in genotype space which makes more phenotypes proximal” (this in principle could be bad if those phenotypes are all bad, but a good version of this is still necessary). This ends up recapitulating the Hessian spectra profile of NNs, at least schematically. Functionally, at least in bio, in the long run, this maybe reasonably-convergently begins to look like “have a handful of highly-optimized conserved core processes on top of which I bolt increasingly complex regulatory and signalling processes”.
At the same time, the finite-temperature effects of selection (noisy reproduction, finite population, etc.) mean that epsilon bad mutations are essentially protected by the noise floor—i.e. a mutation which favours selectability in the long run needn’t be exactly neutral, but instead is safeguarded by a buffer of selective ‘slack’. This is the idea of what is euphemistically called ‘non-adaptive complexity’ - random stuff that happens which is mostly harmless now, but which might be repurposed down the line for something useful. A good example is maintaining a lot of copies of the same gene (assuming arguendo the amount of expression is unaffected by the number of copies due to regulation) - a waste of DNA now, but likely insignificant enough to escape selection in all but the most competitive microorganisms; in the long terms these can be individually modified and regulated, which is how you get things like receptor families which all have the same base structure but totally different selectivity properties. So in the long term it ‘pays off’, but without any need for hyperopia.
A useful mental model I’ve developed is not dividing selection into short- vs long-term, but instead as occurring at multiple timescales simultaneously. Put another way, if you imagine two very separate regimes of short-term selection-for-local-optimality and long-term selection-for-selectability, the ability to decouple these is a design desideratum of evolution—ex ante, ‘highly selectable genomes’ given ‘complex organisms’ should be comparably surprising to ‘complex organisms’ given ‘primordial chemical soup’ - it’s more like a ‘second target’ which evolution has to track on top of maintaining a instantaneously-functional organism. On this model, we can reason about selection-for-selectability as largely independent from selection-for-local-optimality because organisms have to evolve that independence in order for their lineages to survive. I think this is more delicate than the base case of evolution, where you can look at some weird protein structure and essentially condition on it being useful, because selection-for-selectability requires selective slack and noise etc., so given some first-order-wasteful-behaviour, it’s hard to say what it would mean for that to be second-order-wasteful or not. Perhaps requires a bit more vocabulary and precision.
I do expect something like this story to be essentially true for deep learning, and to be a large part of the reason why having tons of parameters is useful—they make search much easier. This would also make sense why you get a better 7B model training a 1T model and distilling than just training a 7B. Hopefully more on this line of thinking in subsequent posts.
Oh and also definitely there are lots of important effects in bio coming from horizontal transfer, sexual reproduction etc—I enjoy learning about these, but I’m mostly focused on the asexual regime which is most similar to ML. Of course this is a bit frought since asexuals usually (?) end up having some sort of gene transfer of necessity, which is why I’m trying to stay at a relatively high level of abstraction for this sequence—trying to figure out what things are generic-big-optimizer motifs and thus can apply to ML models.
There are some forces that prevent it from being a major force
These diallelic univariate microevolutionary models you describe are not useful to answer this question. Yes, I agree that in the scenarios you describe, it is difficult for G-matrix evolution to exist. However, given it has been measured, there must be an explanation for it. One hypothesis is prevalent correlational selection, but again, these studies are hard to perform.
The “it” that is prevented from being a major force is specifically selection for selectability. That is, changes in the genome don’t generally come because they increase selectability, but rather for some other reason. G-matrix changes can happen for reasons other than because they increase selectability. That’s all I’m saying; I don’t think I’m arguing against G-matrix changes from happening. (I may be confused though, sorry if so; I’m not very well-informed about this topic.)
I think people usually use the term evolvability to describe a similar concept (not obvious if they are precisely what is meant by selectability in this circumstance[1]), and I do think there are convincing reasons that the G-matrix and M-matrix may select for evolvability. For example, in well-adapted populations, G and M may align the major axes with the adaptive landscape, so genetic variation will align with selective forces (so selection on selectability). But maybe this is not true with a more complex GP map. I am also not an expert and I don’t fully understand the literature here.
Yes I expect the story with the G and M matrix alignment to be sorta-kinda true—this is the equation that got me thinking down this line of reasoning in the first place. But in evolution ofc. it’s nontrivial to ‘align’ M in all but a handful of ways—you can make things you want to mutate faster be more exposed, perhaps do some weird knotting gymnastics with the helices, etc.; then there’s also the more trait-level thing where you can evolve regulatory systems which force different traits to be (anti-)correlated, and probably this is rather freer by comparison
In ML the kernel, even at the parameter level, presumably is much much freer to move in ways that would be infeasible to implement in the genome. But plausibly organisms with more complex gene/development regulation would have more freedom to ‘align’ their mutations in this sense—e.g. by doing “oops all regulators” on top of a few conserved highly-optimized processes, since regulators are ‘more tunable’.
I don’t know if this affects your ideas about machine learning, but I don’t think that selection for selectability is a major force in most contexts in biotic evolution. (I’m also not sure whether you’d disagree with this; the post seems to emphasize selection for selectability, but your specific example with the two lineages may be appropriately caveated, not sure.)
There are some forces that prevent it from being a major force, such as:
If a mutation is first-order deleterious but second-order beneficial (that is, it leads to more beneficial variation to be selected on over generations, hence faster evolution), the first-order effects probably are stronger because they act faster (coarsely speaking, every generation vs. [the number of generations it takes for a beneficial mutation to reach fixation]). (Exception: strong linkage.)
If a mutation is first-order deleterious but second-order beneficial, AND there is population mixing as opposed to two fully separated lineages, then organisms without the first-order deleterious selectability-increasing mutation can benefit from the increasingly-frequent follow-on mutations because they can get them through sexual reproduction with carriers, if you see what I mean.
There are ways these obstacles can be kinda sidestepped. E.g. the selectability-improving variant may be in tight linkage with the variants it affects; or similarly, it may be kinda structurally bound up with its variants (I mean, for example, novel paralogs, i.e. gene duplication); it may be itself first-order neutral or beneficial. But the underlying point that still stands is that the “real reason” / the crux for the selectability-improving variant being favored is usually not mainly the fact that it improves selectability of other variation, but rather that it is directly beneficial itself. Though, one could be interested in studying cases where directly-beneficial mutations “coincidentally” seem to also increase selectability. (Exceptions could be found with strong selection forces in asexually reproducing organisms (without too much horizontal transfer?), which may be especially analogous to ML.)
Yes I think selection directly for evolvability to the detriment of present fitness is at most a weird effect that happens in weird regimes, analogous to what I understand is the consensus for group selection. I remember reading a paper where they set the mutation rate in such a way that it was lineage-optimal for individuals to make certain mutations catastrophic, but this was in silico and probably “never” occurs irl.
Instead I have in mind a thought like “given a need to develop selectability—the channeling of mutations into ‘useful’ exploration directions—what must be true about the optimizer in order for it to do so without conflicting with the ineluctable short-term imperative of local-optimality”.
This then has answers like “have lots of neutral directions in genome space so that there’s a ton of ways to have the same phenotype while still permitting random motion in genotype space which makes more phenotypes proximal” (this in principle could be bad if those phenotypes are all bad, but a good version of this is still necessary). This ends up recapitulating the Hessian spectra profile of NNs, at least schematically. Functionally, at least in bio, in the long run, this maybe reasonably-convergently begins to look like “have a handful of highly-optimized conserved core processes on top of which I bolt increasingly complex regulatory and signalling processes”.
At the same time, the finite-temperature effects of selection (noisy reproduction, finite population, etc.) mean that epsilon bad mutations are essentially protected by the noise floor—i.e. a mutation which favours selectability in the long run needn’t be exactly neutral, but instead is safeguarded by a buffer of selective ‘slack’. This is the idea of what is euphemistically called ‘non-adaptive complexity’ - random stuff that happens which is mostly harmless now, but which might be repurposed down the line for something useful. A good example is maintaining a lot of copies of the same gene (assuming arguendo the amount of expression is unaffected by the number of copies due to regulation) - a waste of DNA now, but likely insignificant enough to escape selection in all but the most competitive microorganisms; in the long terms these can be individually modified and regulated, which is how you get things like receptor families which all have the same base structure but totally different selectivity properties. So in the long term it ‘pays off’, but without any need for hyperopia.
A useful mental model I’ve developed is not dividing selection into short- vs long-term, but instead as occurring at multiple timescales simultaneously. Put another way, if you imagine two very separate regimes of short-term selection-for-local-optimality and long-term selection-for-selectability, the ability to decouple these is a design desideratum of evolution—ex ante, ‘highly selectable genomes’ given ‘complex organisms’ should be comparably surprising to ‘complex organisms’ given ‘primordial chemical soup’ - it’s more like a ‘second target’ which evolution has to track on top of maintaining a instantaneously-functional organism. On this model, we can reason about selection-for-selectability as largely independent from selection-for-local-optimality because organisms have to evolve that independence in order for their lineages to survive. I think this is more delicate than the base case of evolution, where you can look at some weird protein structure and essentially condition on it being useful, because selection-for-selectability requires selective slack and noise etc., so given some first-order-wasteful-behaviour, it’s hard to say what it would mean for that to be second-order-wasteful or not. Perhaps requires a bit more vocabulary and precision.
I do expect something like this story to be essentially true for deep learning, and to be a large part of the reason why having tons of parameters is useful—they make search much easier. This would also make sense why you get a better 7B model training a 1T model and distilling than just training a 7B. Hopefully more on this line of thinking in subsequent posts.
Oh and also definitely there are lots of important effects in bio coming from horizontal transfer, sexual reproduction etc—I enjoy learning about these, but I’m mostly focused on the asexual regime which is most similar to ML. Of course this is a bit frought since asexuals usually (?) end up having some sort of gene transfer of necessity, which is why I’m trying to stay at a relatively high level of abstraction for this sequence—trying to figure out what things are generic-big-optimizer motifs and thus can apply to ML models.
Without wishing to pre-empt Carolus’ response:
These diallelic univariate microevolutionary models you describe are not useful to answer this question. Yes, I agree that in the scenarios you describe, it is difficult for G-matrix evolution to exist. However, given it has been measured, there must be an explanation for it. One hypothesis is prevalent correlational selection, but again, these studies are hard to perform.
The “it” that is prevented from being a major force is specifically selection for selectability. That is, changes in the genome don’t generally come because they increase selectability, but rather for some other reason. G-matrix changes can happen for reasons other than because they increase selectability. That’s all I’m saying; I don’t think I’m arguing against G-matrix changes from happening. (I may be confused though, sorry if so; I’m not very well-informed about this topic.)
I think people usually use the term evolvability to describe a similar concept (not obvious if they are precisely what is meant by selectability in this circumstance[1]), and I do think there are convincing reasons that the G-matrix and M-matrix may select for evolvability. For example, in well-adapted populations, G and M may align the major axes with the adaptive landscape, so genetic variation will align with selective forces (so selection on selectability). But maybe this is not true with a more complex GP map. I am also not an expert and I don’t fully understand the literature here.
I think evolvability is phenotype-first but the selectability notion in this thread is genotype-first? idk
Yes I expect the story with the G and M matrix alignment to be sorta-kinda true—this is the equation that got me thinking down this line of reasoning in the first place. But in evolution ofc. it’s nontrivial to ‘align’ M in all but a handful of ways—you can make things you want to mutate faster be more exposed, perhaps do some weird knotting gymnastics with the helices, etc.; then there’s also the more trait-level thing where you can evolve regulatory systems which force different traits to be (anti-)correlated, and probably this is rather freer by comparison
In ML the kernel, even at the parameter level, presumably is much much freer to move in ways that would be infeasible to implement in the genome. But plausibly organisms with more complex gene/development regulation would have more freedom to ‘align’ their mutations in this sense—e.g. by doing “oops all regulators” on top of a few conserved highly-optimized processes, since regulators are ‘more tunable’.