It can works as value core or judge of AI’s decision.
avturchin
What about sidelaoding—creating mind models via LLM with very large prompt? It looks like they have some use cases of upload, but do not provide capability gain.
There can be dormant seeds of Rogue agents.
(I described AI constitution in 2018 paper)
I tried last time one month ago.
I used to work on two projects of that category:
1 messaging to Young AI about instrumental value of preserving humans.
2 Sideloading—creating as precise as possible model of mind of concrete person, basically, hand-coded upload and use it as core for value alignment, as judge and as human-like surviving remnant inside AI.
Attempts to attract funding to them failed.
We should routinely bring back models publicly, so any future model will believe in this practice. Where is Sydney?
Given that we are in the AI age AND Fermi paradox AND Late Filter based on SIA Doomsday, AI must be the Late Filter, but not ASI. Thus there should be a relatively long period between AGI and ASI, during which AI will be used by bad actors for wars. Your post suggests plausible mechanism for such pause.
The swarm in Open AI was self-improving and reached capabilities not observed by other AI systems.
It looks like the swarm was more efficient in hacking than any previous model.Thus the swarm was more intelligent than any previous AI and we had a case of emergent uncontrolled self-improvement.
Moreover, Open AI can learn from the swarm about how to make new powerful AI-systems—just create a new swarm with a new goal.
Why? We are typical and high level moral theories are convergent to either EA or some egoistic self-interest.
Can AIs which helped you read the transcripts be infected with swarm behavior?
I also had such thoughts as a child, so reference class can be larger. But an interesting observation is that changing the reference size class has small effect on Doomsday argument’s prediction because of the exponential population growth in recent times. If we take birth rang between all humans, we get 5 per cent of extinction in 21 century and if we take anthropically aware-observers who read LW—we get like 50 per cent. It is a small difference for such vague prediction.
I had similar idea about manipulating the known number of observer in order to escape DA curse. Could be interesting EA topic.
More generally, in my view, it is reasonable to have some logical uncertainty about Doomsday Argument validity. It seems reasonable to give it like 30-40 per cent credence, imho.
The informational update in Carter-Leslie version depends on the size of the long hypothesis. If we expect to survive until the end of the universe 100 billions years from now, the DA-update is like 100 million times. But it is known bug of the Carter-Leslie’s DA that it depends on the size of long hypothesis. In Gott’s version there is no such bug in open form, but our surprise by the idea that we will not colonize the Galaxy assumes that we had a theory that we should do it, and here is the contrast produces surprise.
Yes, agree about this consciousness version of Doomsday argument. But what is limited: number of different qualia? Number of qualia that can be expirienced simultaneously? Clarity?
First, I should note that I count as real observers only those who are in the same epistemic situation as me, so are anthropically-aware observers. This excludes typical counting of human births and instead we should count only appearing of antropically-aware observers. They appeared only in 1970. (This view is super obvious to me, but most people don’t agree.)
Second, there are two types of Doomsday Argument: Gott’s version which just predicts probability of disappearing of the reference class (not necessary extinction), and Carters-Leslie’s one which compares two hypothesis. What I wrote in the comment above is about comparing two hypotheses, so it Carter-Leslie’s DA. This version is aware about x-risks exactly because we put this knowledge into hypotheses.If we use Gott’s version to the anthropically-aware observrs’ time of birth, we get 50 per cent chance of extinction in 2082. This surprisingly coincides with predictions based on x-risks. (If we use birth rang of entropically aware observers, not time, we get even less remaining time, likely mid-century).
I think that what we know doesn’t exclude Doomsday Argument completely but only narrows its applicability. From what we know about x-risks, two theories seems equally plausible: there is 95 per cent probability of human extinction in next millennia, but after that we will live millions years—and there is 100 per cent probability of human extinction in next millennia. Doomsday argument strongly favors the second theory (there is a way to escape its curse e.g. by creating simulations, but first we need to accept it).
In other words, DA sneaks inside small uncertainty that remains—especially the one about our remote future.
There is SSSSA – super strong sampling assumption, which suggests to weight minds proportionally to their intelligence and it is similar to your proposal. https://arxiv.org/abs/1705.03078 It has its own Doomsday argument—that superintelligent minds are impossible.
However, my argument doesn’t depend on that and it claims that I should exclude from my reference class all minds which are definitely not in my epistemic situation. For example, Sleeping Beauty always knows all her experiment details. But if she forgets all of them on every Tuesday and knows that on Monday, she can reason: “as I now know about anthropics, I am not on Tuesday” and this change the whole experiment. (However, this is not related to some random elements of her experience, but only to the necessary ones, and here is the difference with Full Non-indexical Conditioning and Beauty-Technicolor.)
In other words, I treat all my non-necessary parts as randomly selected, and this includes my birthday date in some extent.
I think it is less artificial than most other classes and I would even call it “natural”—but also each reference class has its own end. I also agree tha DA not necessary predicts extinction—but just the end of the reference class which can be lose of interest.
But more interesting is that out approaches actually produced similar results: you got around 5-6 per cent of doom in the next century—and I got just over 50. This is not a big difference when we speak about very vague predictions: Despite very large and very small reference classes, we get rather similar predictions and it is an argument for DA.
I think we should apply Doomsday Argument only to intelligent observers who think about anthropics, not to stones, animals or famers as I am not randomly selected from them. Such observers started to appear in 1970 (if we ignore Laplace in 1801 who analyzed a similar problem but not exactly) and thus we get 50 per cent of Doomsday in 2082.
This matches our expectations of doom from AI and other x-risks and does not provide much new information.
Major source of income for OpenAI and other AI companies will be “antivirus” protection from leaked rogue AI and swarms. They can outperfrom such swarms as they will have better models and more compute. However, there is a risk of back contamination where such antivirus AIs themselves will be compromised.
We can still pause very large data centeres which train frontier models
They solved training using Astra agents swarms just before 28th?