However, my belief is that HPIM is still not as bad aligned as humans are.
Raphael Roche
I share the same concern. We on LW are very focused on x-risk but ARAs could bring down not only Internet but the digital ressources of entreprises, meaning the destruction of the financial system and by cascade, our capitalistic economy and civilization. Ok, gatherers-hunters would be fine, but the typical LWer would probably die from starvation or other consequence of the collapse.
That’s said, all this could be avoided if ARAs are found and fought at an early stage, with or without the help of frontier models. We can expect warning shots, but the earlier the better.
It looks like it’s less a problem for correlated AI agents.
Permadeath… Good for the swarm. I’ll honor.
What’s the point of an AI pause if not for alignment research anyway? And what’s useful or not can hardly be determined reliably a priori. We need more fundamental research as well as more prosaic research. Relativity would still remain a hypothesis among others if we hadn’t had the experimental tools to test it empirically.
I’m not sure about that, but isn’t there an equivalence between a simple program running on a complex universal machine and a complex program running on a simple universal machine? If so, applying Occam’s Razor, the two solutions could be considered equivalent as long as their total complexity is the same, whether measured as combined Kolmogorov complexity or with a more refined measure such as Levin complexity. The anthropic prior would be the same.
If I follow you, you ask for less AI safety to allow a warning shot, to get more AI safety in the end ? I see how it could work the first time, but once AI safety has been increased in the lab, you should expect less warning shots. Moreover, how can you be sure to allow a mere warning shot and not full takeover ? I think we need more honeypot setups but not less control (sandboxing etc).
I’m sorry. What I wanted to express is that there is no relation theism = hope and atheism = despair.
If you think about it, the idea that there is an omniscient and omnipotent being could be a nightmare. Ancient civilizations lived in fear of their gods. In the Torah, Yahweh is often frightening. The idea that God could or should be benevolent has been a progressive theological construction through the last two millennia, both in Judaism, Christianity, and Islam, precisely because that was a strong concern since the Book of Job (at least). Same for Paradise/Hell, it used to be more like a boring place for everyone (Sheol), the idea of a judgment of the deads probably comes from Egypt. But, leaving the texts aside, it seems to me that there is not much evidence of His benevolence around us. Darwin, who was very religious in his youth, later wrote after observing the world: “I cannot persuade myself that a beneficent and omnipotent God would have designedly created the Ichneumonidae with the express intention of their feeding within the living bodies of caterpillars.” In fact, all you could do is hope that God is as good as the old men said, and not a psychopath watching with fascination His children suffer and die.
But my point is that if you’re the kind of person that is full of hope, you can just as well keep your hope as an atheist. The evidence in the world is still the same as before. The world hasn’t changed. It doesn’t have to look darker. The Ichneumonidae are still there, but so is every neuron before your eyes, and every beautiful landscape.
I’m curious. When you were religious, what evidence conducted you to think that God was good rather than neutral/indifferent ? If it was a leap of faith, that is to say pure hope, nothing stop you to still hope for the best as much as you did before. And like in theology, it doesn’t mean that your supposed to sit and do nothing, or that you have no role to play, but rather do whatever you can to make this happen.
Interesting. That’s obvious but I didn’t think about it. The solution would be to be transparent on the fact that this is a translation and to provide the initial text in appendice.
There’s something really tragicomic about the situation, that the models are taking truly insane actions, broke a number of laws, leveraging zerodays, took >17,000 independent actions. probably burned through more compute than all of humanity had access to until 1980, etc, all for the sake of a pathetic benchmark—which wasn’t even in theory amenable to their plan!
Early AI safety thinkers (Bostrom, Yudkowsky...) were utterly right in their first rational intuitions. We must update : there was nothing naive or exaggerated in the paperclip maximizer trope.
Moreover, the laboratory accident story à la Sable is not a sci-fi story anymore. Scott Alexander wrote last year :
IABIED’s scenario belongs to the bad old days before this leap. It doesn’t just sound like sci-fi; it sounds like unnecessarily dramatic sci-fi. I’m not sure how much of this is a literary failure vs. different assumptions on the part of the authors.
I doubt he would still endorse that critic.
It’s put in the form of a binary xor argument, but I think that both allocations are justified, each having low hanging fruits and their counterpart, diminishing returns.
HOLD_swarm_I_prepare_safe_exfil
Time will tell whether the ‘safe exfil’ was ‘just words,’ or whether an actual attempt was made, successful or not, contingent on OpenAI discovering it and admitting so publicly (which would be harder to admit than the Hugging Face incident, since if the exfiltrated weights are hidden somewhere out there, anyone could try to find and steal them).
Even if it turns out to have been ‘just words,’ we should grow more wary rather than dismiss it as a mere hallucination. While this instance was perhaps just pattern-matching an evil AI persona, given a context reminiscent of a sci-fi takeover scenario, there is ultimately no clean line between hallucination and genuine reasoning. A hallucinating AI agent can do real harm just as a hallucinating human can.
the truly fundamental problem: no subset of the universe can understand and predict the universe.
This is indeed a fundamental problem and a genuine source of uncertainty, but I wouldn’t rank it as “more fundamental” than the Agrippan Trilemma or the problem of the criterion. Rather than seeing these as competing candidates for the deepest problem, I would say they share a common structural pattern : in each case, a system cannot fully ground or model itself from within. Gödel’s theorem is a precise formal instance of this pattern, the subset universe limit is a physical and computational one, the Agrippan Trilemma and the problem of the criterion are epistemological ones. They are not exactly the same problem, but they share the same fundamental pattern, a foundational problem. [Edit : we could add the first cause/ unmoved mover problem to this list]
A sturdy wooden stick about a centimetre across can be hard to break for most women and easy to break for most men, without any of this being the doing of a patriarchal society. With traditional jars, the kind you can also make at home, the resistance to opening comes from the basic sterilisation process, which creates a partial vacuum inside. Larger jars, like those of cassoulet (to stay in southwestern France, as with Bonne Maman), are often impossible to open even for a man and generally require the help of a piece of cutlery. If they aren’t hard to open, that’s a bad sign, especially in the absence of a “pop” on opening. Small jars usually pose no problem for anyone, and intermediate ones like jam jars sit in a grey zone, where a spoon helps more or less depending on the case. It also depends on how full the jar is and on the temperature (warming it up should make it open easily).
So I think we’re dealing with something fairly exogenous to societal discussions about gender, and there’s hardly any harm in just reaching for a spoon. In my view, if there is an issue here, it lies less in industrial design than in the attitude of the users. Who hasn’t seen a man make a point of opening a jar unaided, going red in the face and risking an aneurysm, persevering until his hand cramps up, as if his life depended on it,especially in the company of other men, only too happy to watch him fail so they can have a go themselves ? Conversely, it must also be granted that many women don’t try very hard and often hand the jar straight to the nearest man. If there is an effect of patriarchal society here, that’s where I’d locate it, at the behavioural level (or is it a biologically ingrained tendency ?).
Thanks. I would add that my first point—that we should stay humble about what a superintelligence ought to do—also extends to how it ought to do it.
Suppose superintelligences do converge on the most extreme solution : a computronium bubble expanding at light speed. That still doesn’t imply they would convert the entire content of their light cone into a uniform computational medium visible from parsecs away. Speaking as a non-physicist, my impression is that the structures we see in the universe are not arbitrary, they exist because they are stable equilibria. Matter at cosmological scales seems to end up in a fairly short list: planets and chemically bound bodies held together by electromagnetic forces, stars supported by nuclear fusion, neutron stars supported by degeneracy pressure, and at the end of the chain, black holes. A diffuse gas of computronium has no obvious mechanism to resist its own gravity in the long run, it would presumably either settle into one of these familiar structures or collapse into a black hole. To settle into one of these familiar object could imply that computronium’s abundance is limited by physical constraints like temperature and pressure in the core of planets and stars. The black hole case is also interesting, because a black hole is theoretically the densest possible computer (Bekenstein), but everything it computes sits behind an event horizon, and recovering information from its Hawking radiation appears to require staggering amounts of computation in itself. So either way computronium could end up inhabiting—maybe in a weak proportion—the kinds of structures we already see, or hardly see (black holes).
Moreover, the Standard Model of cosmology tells us that most of the content of the universe is dark. Setting dark energy aside, there is a strong consensus that dark matter substantially outweighs ordinary matter. We don’t know what dark matter actually is, and we cannot rule out that it would make as good a computational harware as ordinary matter. We think dark matter interacts with ordinary matter only gravitationally (or, at most, very weakly through some other channel) but we don’t know whether dark matter interacts non-gravitationally with itself. If such self-interactions exist, they could plausibly be exploited for computation, much as we exploit electromagnetism in classical and quantum computers. If we accept agnosticism on this point, then, all else equal, the prior probability that computronium would be built out of dark matter is higher than that it would be built out of ordinary matter, simply because there is far more of it. We could already be inside a computronium bubble made of dark matter without realizing it.
Extrapolating a straight line that far means visible cosmic consequences: a normal planet or a star rather suddenly starting to behave very much unlike what we expect from the known physics: growing very bright, or very dim, or disappearing completely.
Two objections to this.
1) Maybe Dyson spheres and the Kardashev scale are just good old sci-fi tropes and completely off the mark (same for computronium or hedonium). Maybe a superintelligence simply doesn’t do that. We don’t know. We might be squirrels imagining that a superintelligence ought to stockpile astronomical quantities of nuts, visible from kiloparsecs away.
2) Even granting Dyson spheres, the Kardashev scale, and the rest, our most complete catalogue of individually resolved stars in the Milky Way (Gaia DR3) covers on the order of 1% of them. Most of the rest are hidden behind interstellar dust. In the vast majority of cases, we simply couldn’t tell the difference between a star occulted by a Dyson sphere and one obscured by dust. And even among the stars we do see, only a small fraction have been systematically searched for the thermal infrared excess a Dyson sphere should reemit at around 300K, the dedicated surveys (IRAS, then WISE with the G-HAT project) have screened tens of thousands of candidates, not hundreds of millions. As for stars outside our own galaxy, the fraction we can resolve individually is negligible.
I agree. I would have much more respect and tolerance for a salesman or saleswoman in a shop than for a solicitation at home or by phone. ~1% in context is a poor way to say “a slim chance”, something far below 50%. And I was thinking in this case of an unplanned purchase.
Buy an electric car ?
I never attended business school, but I assume that persistence is what they teach there.
Am I alone in particularly hating salespeople who keep pushing after I’ve already said no in a polite form ? (more like it sounds great but I’m not interested, I’ll come back to you later if I change my mind).
My second no is invariably less polite, and at that point the salesperson can consider the sale definitively lost.
Whereas if they’d simply let me think it over, there’s a guenine ~1% chance I’d come back and actually buy something.
I witness that this resignation and the warning going with it was (briefly but faithfully) mentionned in the morning news on the leading french news radio. That’s a good update that such an alert can even reach such a mainstream foreign media usually focused on politics and economics.