Views my own, not my employers.
cdt
This is not how a frontier lab should be interacting with the math community.
Why would we expect them to behave any differently? They want to be first, and they don’t want to share. In all things, not just math, the aim has been to reach the finish line, regardless of whether the outcome is good, useful, truthful, safe, honest, or thoughtful.
Even if the math community decided to punish them, what force could they muster? The power is in their hands now, and academia is increasingly weakened. When frontier labs use up the remainder of the math community and discard them like rubbish, I don’t think there will be much surprise.
I don’t really understand why, in all the graphs, the numbers are not the same each time. I mean, compare this graph to this graph. Why are the numbers not the same (even rounded)? Why are the time horizon bins different?
Strongly upvoted this because, while I don’t feel this argument is well-reasoned, the negative karma system has the unfortunate effect of hiding this reaction. It is important for LW to see responses to the changing culture.
(The post was largely meta, so this comment is meta-meta?)
Modelling variation in the METR Uplift Study
It strikes me as weird for employees to independently manage the public like that. Why stick your neck out and risk being off-message? Perhaps this is more indicative of their freethinking culture /OpenAI’s lack of discipline (which has been noted before).
another instance of … whatever
To be clear: do you think it is one of:
conscious coordination,
peer pressure,
having similar world views,
coincidence?
I get this. I find even for comments sometimes it’s just too short. What I really want to do is pass a post-it note saying “Read this! <link>” without any more fanfare.
Is there a popular resource which can give a layman a feel of this complexity?
Try KEGG, they have maps on the front page, just click the links.
This is a closely related question to this recent post, if you haven’t read it.
if pathogen-relevant capability comes from pathogen data, restricting access to some viral datasets could potentially reduce risk (while preserving most biological research).
This means some research may need restricted dissemination, and some questions may not be worth answering at all. We don’t think the default should be that everything produced by a BAIM-safety research program is published
There are twin problems with this:
A) It’s easier to generate data and train on prokaryotes, which are most likely to be pathogenic and thereby dangerous. Eukaryotes are safer but there is significantly less data of the kind that is useful for training.
B) I’ve seen a lot of project proposals that suggest safety classifiers that infer virulence properties directly from sequences. This functions the same way as a tool for finding novel weaponry. I think the dual use risks of these tools are extremely high. Working on them but holding them back from public sight is dangerous since you would essentially build a private repository of weaponry. The only way to get around this I can think of is to classify overbroadly, but that produces a lot of resentment and encourages end-users to break it.
I’m only speaking out about this here now because it has become popular in the public sphere, so infohazard << effectiveness of the warning IMO. I haven’t fully understood why this is different from classifiers for cyber-risk: maybe because we understand the types of dangers more completely.
I think people usually use the term evolvability to describe a similar concept (not obvious if they are precisely what is meant by selectability in this circumstance[1]), and I do think there are convincing reasons that the G-matrix and M-matrix may select for evolvability. For example, in well-adapted populations, G and M may align the major axes with the adaptive landscape, so genetic variation will align with selective forces (so selection on selectability). But maybe this is not true with a more complex GP map. I am also not an expert and I don’t fully understand the literature here.
- ^
I think evolvability is phenotype-first but the selectability notion in this thread is genotype-first? idk
- ^
> heal the souls of AI researchers and/or their surrounding culture, so that they can stop being so nihilistic or selfish or deluded or similar
I’ve often heard from the perspective of AI boosters that AGI research is to compensate for the lack of hope contemporary society has to solve the the modern-day crisis. AGI has to happen so we can solve… uh, “climate change (or lack thereof), death, wokeness” etc. I don’t want to argue whether their perspective is right or wrong, simply that this is a foundational belief.
But there is something I don’t understand: in 2015, when the first glimpses of this movement were happening, the dominant theme was not hopelessness, so how did we get here? How would we know that this is not a retroactive justification?
Can you describe what physically you want to happen? Is it wealth redistribution? Is it something else?
Without wishing to pre-empt Carolus’ response:
There are some forces that prevent it from being a major force
These diallelic univariate microevolutionary models you describe are not useful to answer this question. Yes, I agree that in the scenarios you describe, it is difficult for G-matrix evolution to exist. However, given it has been measured, there must be an explanation for it. One hypothesis is prevalent correlational selection, but again, these studies are hard to perform.
Hi, this was all very interesting to me, I have some thoughts for you:
This is what I mean by selection for selectability: the architecture of variation is shaped by which variation selection was able to see and use. Foreshadowing, we will call this genome–environment alignment—the genome structures itself so that common mutations align with common environmental variations.
I think it’s good and interesting to talk about the evolution of genetic architecture of phenotypic variation. I wonder how we can best split this apart from:
Changes in mutational spectra → changes in the distribution of phenotypic effects from denovo mutation
As an aside, changes in recombination can induce changes in mutational spectra because recombination hotspots can themselves increase mutation rates at a locus. I think this is one of the running theories for adaptive radiations.
Changes in the maintenance of standing variation → changes on phenotypes directly from standing variation
In other words, how can we split apart the effects of G-matrix (additive effect) evolution and M-matrix (mutation effect) evolution? One example would be different effects on phenotypic drift. But I don’t know—I haven’t read much prior work in the particular area you are discussing.
measured G-matrices are typically low-rank: most of the available variation lives in a small number of effective directions. Note the similarity to the low intrinsic dimensionality of fine-tuning, LoRA, and friends
Unfortunately we know very little about how the G-matrix changes across populations and in time. There’s a handful of works out there and they are quite old now—if you know more I’d love to hear about it! I think because this kind of work is incredibly manual and it doesn’t particularly jive with the modern sequencing-based paradigm.
I’ll save it for the followup speculation post, but I expect this sort of trait correlation to be half of an explanation of emergent misalignment.
Do you have any idea why these feature correlations / trait correlations might appear? There are many examples of biophysical trade-offs resulting in trait correlation in ecology and this is basically one of the modern frontiers of ecological science. I don’t know if any trade-offs exist for languages or information.
It could explain neuralese (i.e. too few words to produce effects, so words gain unusual meanings because they are correlated with multiple required outputs) but most people I have discussed this with doubt that LLMs are complex enough to produce this effect. What are your thoughts?
It’s not clear to me that public-facing posts about a two-person dispute generate anything but more heat, I’m afraid. You said yourself (emphasis mine):
(AFAICT this is the first time [Charge of the Hobby Horse] has been applied in real life, outside of your OP. Not a lot of examples to go on, but if we waited until it’s used more widely, it might be too late to stop at that point.)
The community clearly liked your comment because it was highly upvoted. But nobody can force any specific person to like your comment. Removing yourself from the community only punishes the people who upvoted your thoughts. As far as I can see, you are not banned by anyone currently, so why not just wait it out? If you need to critique something and you end up banned, why not turn it into a Quick Take?
Just anecdote, but I had difficulty with applications this year as a doctoral candidate with direct biosecurity experience. Some of this may be down to disagreement with house view, YMMV.
EDIT: I think it’s difficult if you’re someone with applicable skills, fellowships and formal roles can feel like “waiting to receive permission to do what you can do anyway”. Biosecurity is even harder to contribute to as an outsider because of the obvious moral and info risks. If you want to work on monitoring, that seems safe and plausible as an outsider.
In my experience the subjective feeling of being full or empty is not a good estimator of whether you are gaining or losing weight. Even switching choices of food without becoming vegan/vegetarian per se can prompt weight gain or loss (e.g. staying at a friend’s, and thereby eating like they do).
I suspect that internally AI companies have benchmarks similar to this to hill climb onto, but yeah, yeesh, this is the exact opposite of safe.
Looking at the karma bouncing up and down, it feels like this comment is being used as a proxy for some other argument: maybe to get out frustrations on OpenAI. I do feel strongly that the money and power involved will create bad outcomes, even if everyone has good intentions—it already happens in mundane academic scooping incidents. But I don’t like being in the middle of a fight. Perhaps I should have phrased the comment more neutrally, mea culpa.