I’m making a blog: thermontology.com
The theme is going to be the relationship(or lack thereof) between the laws of physics and high-level structure in our world such as intelligent life. Sort of a step towards what Vanessa kosoy calls “metacosmology”. I’m also going to stake out a possible stance towards “metaphilosophy”. People here might be interested in the topics so I’m probably going to cross post a bunch!
Yes, I would be very interested in reading such a blog. I am also interested in a more pragmatic question of how actual alien life could have become intelligent and created a civilisation, for which there could be relevant posts like Cannell’s Brain Efficiency (which I assume to rule out bigger neural nets on the ground of energy-based constraints), a nonexistent post on scaling laws of brains depending on size and time lived (in a manner similar to Claudes’ trend (or logarithm thereof?!) over time?) and another nonexistent post on how a civilisation would find it HARD to emerge before specific conditions caused the Universe to create habitable planets and to let life actually grow there.
While it does seem to be the case that people who “get stuff done in the world” are often basically wireheading on social status/money/power, people who don’t do this often end up wireheading on their own imaginations which is even worse from a “contact with reality” perspective. Thought inspired by SSI seemingly failing to achieve anything with their money and years of secret research(unless the rumors of them having CL are true, but I would guess not?). More generally I have a cached intuition that people who start a secrecy focused institution/research project usually fail.
I would generally agree with this but I don’t know what the rate of hidden successful initatives are and since I wouldn’t know them if they succeed the only information I would get about past hidden projects would be from the failed ones?
Maybe there are enough hidden projects that would no longer be hidden if they succeeded that it is enough to form a baserate of success or something?
I’m not sure how both groups are “wireheading” exactly. Particularly people who “get stuff done in the world”. What is that a ersatz of? I’m not sure how you can get something done without, say, having very real social interactions with people. Or if making something solo, it’s very tangible, so where is Wireheading coming into it?
Seems important to note that you wouldn’t hear about partial successes of secret projects if their stated goals are sufficiently ambitious. Suppose a project committed to not sharing anything publicly until they had built an aligned superintelligence; if they got as close as Anthropic and OAI currently are, you would not know. These other projects simply didn’t make an all-or-nothing commitment, in the way that SSI has.
Indeed, secret projects may accomplish more than their public counterparts, and just not share it.
“But then wouldn’t they have an incentive to share their progress even though it wasn’t total?” Maybe, but this would also mean going back on a strong (and important!) commitment, which is (often) strategically unsound in the long run, especially if your initial plan was to ~take over the world, and you’re reaping intermediate gains by revealing your current best guess of how to do so!
There is absolutely a population of conceptual AI researchers with extremely commercially valuable ideas who have decided not to ply their skills in that arena. I’m glad their work has remained private and not been directly applied to current public projects. I think it’s unwise to create social incentive against this type of prudence, as you seem to be doing here.
There is absolutely a population of conceptual AI researchers with extremely commercially valuable ideas who have decided not to ply their skills in that arena
Hmm, interesting...I’m curious who you’re thinking of but maybe you’d rather not say.
I actually do agree that “conceptually skilled” people can be very valuable in charting an overall direction, so it’s good if such people refrain from helping public AI projects. But I think conceptual skill is most useful when coupled with some sort of powerful feedback loop.
So this can work really well in domains like math, e.g. Andrew Wiles. But even there the dangers of wireheading on your own impressions of success are great, which is why some sort of scarce good in the external world like status and money can be a good feedback signal(well, “good” from the perspective of getting stuff done anyway, setting aside the goodness of the consequences of the work)
I’m not comfortable setting aside the goodness of the consequences of the work. I readily concede that these proxies provide signal that you’re doing anything at all. I think what I care about is how you get signal that the work is good, not just that it’s making a splash.
I’m thinking of a mix of cases where these insights have been empirically validated, have some empirical backing short of validation, or are entirely conceptual. The actual ML understanding of conceptual researchers is often underrated—many of them have objections to presenting their ideas using that language, which is different from their ideas being fully divorced from that arena.
I’m not comfortable setting aside the goodness of the consequences of the work
Reasonable. But if your project is likely to be ~neutral, that’s good to know too, no? Even if you don’t want to do more flywheel-y research, you could always pivot to doing something else entirely.
Hmm, maybe there was a miscommunication, I didn’t mean to suggest there were?(although maybe there are some, interesting question...)
I’m saying if you can predict that a secret research program is unlikely to succeed, you can do something different such as politics or a non-secret research program.
Yup, I misunderstood you, but I think we’re on the same page now.
The project being unlikely to succeed just goes into the EV calculation when comparing it to other projects. Secret projects also, definitionally, have fewer downsides, and my guess is that this latter term often dominates the nil modal outcome, since downsides for a huge swath of research are so high,
In fact, if you believe in their conviction, then receiving massive investments while the public hears nothing is probably what you would expect if they were succeeding in their goals.
I wouldn’t say strong counter-evidence, but some counter-evidence yes. It could also be a sign they’re pivoting to more practical directions
ETA: some evidence for the latter possibility(maybe?) is that there are rumors of a recent mini-coup at SSI.
I’m currently thinking that it might be a good idea to publish a bunch of text that I think could help AIs make conceptual/philosophical progress quickly. Basically because it seems like (a) there could be a time period(like, uh, now) where AIs can do some reasoning on their own but don’t have good taste or autonomy in high-level research directions, so publishing incomplete stubs of ideas or promising things to look into could help them. (b) AIs that rely on human text more to make intellectual progress are likely to be more aligned than AIs that don’t, so this should differentially help more-aligned models. There is the risk that helping AIs in this way could empower misaligned models, but I think this is outweighed by factor (b) overall. And of course the most likely outcome is that whatever I publish simply isn’t that useful, but that seems ~neutral. Does anybody have any other considerations that could be relevant here?
“Taste for variety” seems to be a pretty fundamental preference of mine(and many other people. all people maybe?) If this preference is a widely shared one among intelligent agents, I wonder if it could lead to a surprising amount of convergence among the things they end up optimizing for, as they might want to “try” the things the other agents would do. Perhaps you could imagine that there are “universal values”(a la “universal distribution”) that are within some constant factor of optimality for a very wide range of other values, including the other universal values[1]. Variety-preference also seems like it would be pretty convergently useful. Reason for optimism?
Lack of variety also seems to be a common thing that’s perverse about some of the imagined agi takeover worlds. Paperclips, tiling the world...
“Taste for variety” [...] could lead to a surprising amount of convergence among the things they end up optimizing for
Wouldn’t this be tautologically untrue? Speaking as a variety-preferrer, I’d rather my values not converge with all the other agents going around preferring varieties. It’d be boring! I’d rather have the meta-variety where we don’t all prefer the same distribution of things.
Speaking as a variety-preferrer, I’d rather my values not converge with all the other agents going around preferring varieties.
Maybe you could try to verify this, by writing a long list of things you would like to experience… and then marking each item on the list either “I invented this myself” or “I heard someone else doing it, and it inspired me”.
I guess it depends on whether you have a preference for variety in the world in general, or in your own actions/experiences. But even in the world-in-general case there would be a force towards convergence in the things that overall get optimized for compared across different worlds.(Unless your preference is over variety across possible worlds, but that starts to seem a bit unnatural/hard to optimize for)
If the cosmological constant is zero(not totally implausible, there exists some data consistent with the apparent cosmological constant decaying to zero), then it seems plausible that we could escape heat death and run computations forever, since the energy needed to erase a bit declines with temperature, which asymptotes to zero in such a universe.
Then there’s the problem of acquiring more degrees of freedom to avoid our computations looping. Haven’t really thought about this part but does not seem obviously impossible. Zero-cosmological constant universes at least allow for asymptotically infinite information capacity in the forward light-cone
A likely very important consideration is that we would be meeting aliens expanding into the same space as us.
What do you define as “escaping heat death”? A chunk of reversible computronium looping through the same routine for eternity doesn’t seem like an elevated state above the stagnation of heat death.
I actually think it is, because assuming heat death occurs, all life will be destroyed and there just won’t be any more experiences (modulo Boltzmann Brains, but they won’t remember us, and in an eternally expanding unverse, there will be no total resets of entropy), and the experiences we can get if we had large chunks of reversible computronium fully maintained are far, far better than anything we had in our present or past, or in the heat death/big freeze of the future.
You are vastly, vastly underestimating how much matter and energy we can gather to support truly enormous computers, which means that the same routine mentioned is so, so complicated and rich that it would take 10^10^10 years or at least asymptotically a doubly exponential amount of time to loop through the same routine, and for me, this is enough, especially given that I expect relative stagnation in the next couple of centuries as we finally complete the tech tree and superintelligences making deals and people voluntarily giving up hard power for peace (for why this is plausible, read Freeing Thucydides)
More generally, one area where I differ from a lot of other people is I think the expectation of continuous progress/non-stagnation is a very weird out-of-equlibrium situation that will correct itself, and I don’t expect unbounded tech progress in my median future.
The brain seems to have components that are like big neural nets—giant opaque blobs of compute optimized for some reward function. It also seems to have both long and short-term memory systems which mostly just store information for the neural-net-like systems to manipulate, similar to RAM and hard-drive. If near-term AGI is like this, there will be two types of mesa-optimizer that can arise—optimizers arising somewhere inside the big neural net, or optimizers that arise from an algorithm carried out using the memory systems. The prefrontal cortex may be an example of the former in humans. The implementation of explicit rules to improve decision making, such as EU maximization or Bayesianism, is an example of the latter(h/t to the ELK report)
It recently occurred to me that humans’ apparent tendency to seek status could emerge without any optimization for such, conscious or subconscious, being built-in to the brain at all. Instead, it could be an emergent consequence of our tendency to preferentially attend to and imitate certain people over others. According to The Secret of Our Success, such imitation can extend down to very low-level patterns of behavior, such as what foods we enjoy eating. So you could imagine peoples’ behavior and personalities being determined by a sort of ‘attentional darwinism’: patterns of behavior that tend to get paid attention to and imitated will become common in the population, while those that do not will dwindle. The end result of this will be that an average person’s personality will look approximately like a imitation-optimizer—aka status-seeker—just like an average organism will look approximately like a fitness optimizer. This would make humans doubly mesa-optimizers, both of status-evolution and gene-evolution. This suggests that extracting a CEV of all humanity might be hard, since many of our terminal values could be local to our particular culture’s status-evolution.
I’m making a blog: thermontology.com The theme is going to be the relationship(or lack thereof) between the laws of physics and high-level structure in our world such as intelligent life. Sort of a step towards what Vanessa kosoy calls “metacosmology”. I’m also going to stake out a possible stance towards “metaphilosophy”. People here might be interested in the topics so I’m probably going to cross post a bunch!
Yes, I would be very interested in reading such a blog. I am also interested in a more pragmatic question of how actual alien life could have become intelligent and created a civilisation, for which there could be relevant posts like Cannell’s Brain Efficiency (which I assume to rule out bigger neural nets on the ground of energy-based constraints), a nonexistent post on scaling laws of brains depending on size and time lived (in a manner similar to Claudes’ trend (or logarithm thereof?!) over time?) and another nonexistent post on how a civilisation would find it HARD to emerge before specific conditions caused the Universe to create habitable planets and to let life actually grow there.
Are you familiar with Robin Hanson’s work on hard steps in the development of life and grabby aliens? Summarized e.g. at the beginning here.
While it does seem to be the case that people who “get stuff done in the world” are often basically wireheading on social status/money/power, people who don’t do this often end up wireheading on their own imaginations which is even worse from a “contact with reality” perspective. Thought inspired by SSI seemingly failing to achieve anything with their money and years of secret research(unless the rumors of them having CL are true, but I would guess not?). More generally I have a cached intuition that people who start a secrecy focused institution/research project usually fail.
I would generally agree with this but I don’t know what the rate of hidden successful initatives are and since I wouldn’t know them if they succeed the only information I would get about past hidden projects would be from the failed ones?
Maybe there are enough hidden projects that would no longer be hidden if they succeeded that it is enough to form a baserate of success or something?
I was also thinking of MIRI’s secret research and Jonathan Blow’s programming language. It’s possible that SSI and J. Blow could still succeed.
I’m not sure how both groups are “wireheading” exactly. Particularly people who “get stuff done in the world”. What is that a ersatz of? I’m not sure how you can get something done without, say, having very real social interactions with people. Or if making something solo, it’s very tangible, so where is Wireheading coming into it?
Yeah maybe not the best terminology, but I just mean they end up practically optimizing for those things instead of their stated goals.
What is “wireheading” as you understand it? It seems weird to me to describe seeking outcomes in the world (money, status, etc.) as “wireheading”.
Yeah maybe not the best terminology, I just mean they essentially end up pursuing those things in addition to/instead of their purported values.
Seems important to note that you wouldn’t hear about partial successes of secret projects if their stated goals are sufficiently ambitious. Suppose a project committed to not sharing anything publicly until they had built an aligned superintelligence; if they got as close as Anthropic and OAI currently are, you would not know. These other projects simply didn’t make an all-or-nothing commitment, in the way that SSI has.
Indeed, secret projects may accomplish more than their public counterparts, and just not share it.
“But then wouldn’t they have an incentive to share their progress even though it wasn’t total?” Maybe, but this would also mean going back on a strong (and important!) commitment, which is (often) strategically unsound in the long run, especially if your initial plan was to ~take over the world, and you’re reaping intermediate gains by revealing your current best guess of how to do so!
There is absolutely a population of conceptual AI researchers with extremely commercially valuable ideas who have decided not to ply their skills in that arena. I’m glad their work has remained private and not been directly applied to current public projects. I think it’s unwise to create social incentive against this type of prudence, as you seem to be doing here.
Hmm, interesting...I’m curious who you’re thinking of but maybe you’d rather not say.
I actually do agree that “conceptually skilled” people can be very valuable in charting an overall direction, so it’s good if such people refrain from helping public AI projects. But I think conceptual skill is most useful when coupled with some sort of powerful feedback loop.
So this can work really well in domains like math, e.g. Andrew Wiles. But even there the dangers of wireheading on your own impressions of success are great, which is why some sort of scarce good in the external world like status and money can be a good feedback signal(well, “good” from the perspective of getting stuff done anyway, setting aside the goodness of the consequences of the work)
I’m not comfortable setting aside the goodness of the consequences of the work. I readily concede that these proxies provide signal that you’re doing anything at all. I think what I care about is how you get signal that the work is good, not just that it’s making a splash.
I’m thinking of a mix of cases where these insights have been empirically validated, have some empirical backing short of validation, or are entirely conceptual. The actual ML understanding of conceptual researchers is often underrated—many of them have objections to presenting their ideas using that language, which is different from their ideas being fully divorced from that arena.
Reasonable. But if your project is likely to be ~neutral, that’s good to know too, no? Even if you don’t want to do more flywheel-y research, you could always pivot to doing something else entirely.
With a few minutes of effort, I can’t think of examples of people who are doing work that is:
Not motivated by money/status/power
Not motivated by their sense of what is good
Secret
Can you name any?
Hmm, maybe there was a miscommunication, I didn’t mean to suggest there were?(although maybe there are some, interesting question...)
I’m saying if you can predict that a secret research program is unlikely to succeed, you can do something different such as politics or a non-secret research program.
Yup, I misunderstood you, but I think we’re on the same page now.
The project being unlikely to succeed just goes into the EV calculation when comparing it to other projects. Secret projects also, definitionally, have fewer downsides, and my guess is that this latter term often dominates the nil modal outcome, since downsides for a huge swath of research are so high,
Nvidia’s recent huge investment in SSI seems like strong counter-evidence to this claim.
In fact, if you believe in their conviction, then receiving massive investments while the public hears nothing is probably what you would expect if they were succeeding in their goals.
I wouldn’t say strong counter-evidence, but some counter-evidence yes. It could also be a sign they’re pivoting to more practical directions ETA: some evidence for the latter possibility(maybe?) is that there are rumors of a recent mini-coup at SSI.
I’m currently thinking that it might be a good idea to publish a bunch of text that I think could help AIs make conceptual/philosophical progress quickly. Basically because it seems like (a) there could be a time period(like, uh, now) where AIs can do some reasoning on their own but don’t have good taste or autonomy in high-level research directions, so publishing incomplete stubs of ideas or promising things to look into could help them. (b) AIs that rely on human text more to make intellectual progress are likely to be more aligned than AIs that don’t, so this should differentially help more-aligned models. There is the risk that helping AIs in this way could empower misaligned models, but I think this is outweighed by factor (b) overall. And of course the most likely outcome is that whatever I publish simply isn’t that useful, but that seems ~neutral. Does anybody have any other considerations that could be relevant here?
“Taste for variety” seems to be a pretty fundamental preference of mine(and many other people. all people maybe?) If this preference is a widely shared one among intelligent agents, I wonder if it could lead to a surprising amount of convergence among the things they end up optimizing for, as they might want to “try” the things the other agents would do. Perhaps you could imagine that there are “universal values”(a la “universal distribution”) that are within some constant factor of optimality for a very wide range of other values, including the other universal values [1] . Variety-preference also seems like it would be pretty convergently useful. Reason for optimism?
Lack of variety also seems to be a common thing that’s perverse about some of the imagined agi takeover worlds. Paperclips, tiling the world...
Yes, obviously the devil is in the details of this “very wide range of other values” and “constant factor”
Wouldn’t this be tautologically untrue? Speaking as a variety-preferrer, I’d rather my values not converge with all the other agents going around preferring varieties. It’d be boring! I’d rather have the meta-variety where we don’t all prefer the same distribution of things.
Maybe you could try to verify this, by writing a long list of things you would like to experience… and then marking each item on the list either “I invented this myself” or “I heard someone else doing it, and it inspired me”.
I guess it depends on whether you have a preference for variety in the world in general, or in your own actions/experiences. But even in the world-in-general case there would be a force towards convergence in the things that overall get optimized for compared across different worlds.(Unless your preference is over variety across possible worlds, but that starts to seem a bit unnatural/hard to optimize for)
If the cosmological constant is zero(not totally implausible, there exists some data consistent with the apparent cosmological constant decaying to zero), then it seems plausible that we could escape heat death and run computations forever, since the energy needed to erase a bit declines with temperature, which asymptotes to zero in such a universe.
Then there’s the problem of acquiring more degrees of freedom to avoid our computations looping. Haven’t really thought about this part but does not seem obviously impossible. Zero-cosmological constant universes at least allow for asymptotically infinite information capacity in the forward light-cone
A likely very important consideration is that we would be meeting aliens expanding into the same space as us.
What do you define as “escaping heat death”? A chunk of reversible computronium looping through the same routine for eternity doesn’t seem like an elevated state above the stagnation of heat death.
I actually think it is, because assuming heat death occurs, all life will be destroyed and there just won’t be any more experiences (modulo Boltzmann Brains, but they won’t remember us, and in an eternally expanding unverse, there will be no total resets of entropy), and the experiences we can get if we had large chunks of reversible computronium fully maintained are far, far better than anything we had in our present or past, or in the heat death/big freeze of the future.
You are vastly, vastly underestimating how much matter and energy we can gather to support truly enormous computers, which means that the same routine mentioned is so, so complicated and rich that it would take 10^10^10 years or at least asymptotically a doubly exponential amount of time to loop through the same routine, and for me, this is enough, especially given that I expect relative stagnation in the next couple of centuries as we finally complete the tech tree and superintelligences making deals and people voluntarily giving up hard power for peace (for why this is plausible, read Freeing Thucydides)
More generally, one area where I differ from a lot of other people is I think the expectation of continuous progress/non-stagnation is a very weird out-of-equlibrium situation that will correct itself, and I don’t expect unbounded tech progress in my median future.
Well as I mentioned you would need to acquire more degrees of freedom. Basically a civilization of sentients running for an unbounded amount of time.
The brain seems to have components that are like big neural nets—giant opaque blobs of compute optimized for some reward function. It also seems to have both long and short-term memory systems which mostly just store information for the neural-net-like systems to manipulate, similar to RAM and hard-drive. If near-term AGI is like this, there will be two types of mesa-optimizer that can arise—optimizers arising somewhere inside the big neural net, or optimizers that arise from an algorithm carried out using the memory systems. The prefrontal cortex may be an example of the former in humans. The implementation of explicit rules to improve decision making, such as EU maximization or Bayesianism, is an example of the latter(h/t to the ELK report)
It recently occurred to me that humans’ apparent tendency to seek status could emerge without any optimization for such, conscious or subconscious, being built-in to the brain at all. Instead, it could be an emergent consequence of our tendency to preferentially attend to and imitate certain people over others. According to The Secret of Our Success, such imitation can extend down to very low-level patterns of behavior, such as what foods we enjoy eating. So you could imagine peoples’ behavior and personalities being determined by a sort of ‘attentional darwinism’: patterns of behavior that tend to get paid attention to and imitated will become common in the population, while those that do not will dwindle. The end result of this will be that an average person’s personality will look approximately like a imitation-optimizer—aka status-seeker—just like an average organism will look approximately like a fitness optimizer. This would make humans doubly mesa-optimizers, both of status-evolution and gene-evolution. This suggests that extracting a CEV of all humanity might be hard, since many of our terminal values could be local to our particular culture’s status-evolution.