I don’t trust myself as sole dictator (my ethics might hold, but my judgement might be impaired), so the structures will have to be able to override me if I go unbenevolent.
Stuart_Armstrong
I think LLMs would be great if the description it was given included all the information that was relevant, and disastrous if the description didn’t include that (and that the LLM would not be able to tell the difference).
See e.g. the plucked chicken example here (this a relatively weak example, but it gives the flavour): https://www.lesswrong.com/posts/TZgezuYjkfMQxyqJC/value-generalisation-2-the-missing-hole-in-ais-abilities#fn-LgYydes692nevpruq-3
Thus their values generalize to roughly as wide a range of situations as humans’ own values would: they get uncertain when faced with a situation no human society has yet faced, where humans would also have to think about it for a while.
Or, put another way, they fail out of distribution from the values inherent in their textual training data, while humans (while having to think for a while) extend beyond.
But it’s wrong to think of value generalisation as just generalising values; it also needs to generalise the concepts underlying them (what is human? what entities can be conscious? what entities can feel pain? what does respecting the law or respecting society’s values mean when you can change these?), an out-of-distribution situation that occurs far more often.
Thanks!
Anchoring is about the moment you apply the “anthropic correction” (moving from standard probabilities to anthropic probabilities) - this moment might be far in the past or the future, it doesn’t have to be a current moment. But it seems that for D-SIA, such a moment has to be chosen. C-SIA is a bit more robust, but not fully so: choosing the moment before or after a duplication event changes things; C-SIA is only independent of anchoring between anchoring events.
Thus if you consider D-SIA as a) definition of the anthropic correction plus b) chosen anchor, you have a fully defined probability system, which updates in the usual way, no Dutch books.
And the choice of anchor point for D-SIA only changes the weight of a possible universe in an expectation-zero way—except across duplication events, where D-SIA suffers like C-SIA.
You can consider conditional probabilities like “given that the universe is roughly how it looks, what is the probability of X” which are coherent under D-SIA in infinite universes, but not C-SIA. Or you can choose your definition of agent type to include the whole agent history; in that case, the brief-and-almost-historyless Boltzmann brains can be a very small subset of the number of possible agents; essentially, duplicate Boltzmann brains don’t count much. Or you can choose agent type to be every agent, but use a density measure over time, which will crush the the number of future Boltzmann brains down.
Basically, D-SIA doesn’t tell you how to choose your agent prior; but, when the agent prior has been chosen, it remains coherent, unlike C-SIA which breaks on infinite numbers of agents or (or finite number of agents but infinite expected numbers of agents).
D-SIA never has a problem in a single universe, even if it contains an infinite number of agents. But it can have a problem with infinite expected number of agents across different universes (e.g. you can always pick a uniform prior to make D-SIA=C-SIA within any universe with finite number of agents, making D-SIA break just as C-SIA does). Given that D-SIA breaks in that situation, can we still say sensible things? Well, the “modal trick” allows us to approximate the infinite expected number of agents across different universes “from below” with a family of probability distributions where D-SIA is well-defined. So what we would want to say (have not fully checked the maths and convergence results on this, sorry) is that any probability number that converges to a single value however \lambda goes to zero, should be taken as the value of D-SIA in the collection of infinite universes where it isn’t defined.
So, careful legal analysis before doing any of this, got it :-)
Thanks for the feedback. Feasibility conversations at the beginning are always useful, even if legal experts would catch the issues later on.
Nothing is ever guaranteed, but we would recruit moral people. One possible approach would be to have contracts that are much more restrictive (in terms of IP and working on the ideas at a rival) in the case of ethics board blocking, rather than otherwise. That’s a good idea; thanks for prompting the thought.
Recall that this would be under the assumption that: “it is deemed safe (by us and the AI ethics board) to forge ahead commercially (or at least safer than not doing so)”. If that assumption fails, then we’d have shut down or pivoted to immediate profit instead of doing this.
So if we go ahead commercially, we don’t have to resist excessive pressure—we just need to be commercially successful within the standard parameters. We also don’t need to be strictly better than OpenAI or Anthropic—just successful in specific sub-markets or sub-uses. One option is to build on that success with a possible licensing system—“you can use our technology, for a reasonable fee, as long as it’s also using value-aligned models”. I’m hoping that pre-aligned models will be so successful in their own niche that they set the design for that these models should be.
That’s why part of the plan is to sell the models (or use of the models), not to sell the research https://www.lesswrong.com/s/EYgCdcxsn73fKWWra/p/uMKGaEKRDpoqnZyBh
And that’s the main reason that I’m considering the commercial path in the first place, to get some control over the use of the technology.
Dividing the possible world into “value generalisation has strong capability increases”: then I can get powerful models with much less investment (today’s models are ridiculously overtrained to compensate for their poor generalisation). And “value generalisation doesn’t have strong capability increases”: then the risk is lower.
The bad spot would be “value generalisation has strong capability increases that only large established companies can take advantage of”.
Simple Bayes means that, if there is no longer any anthropic angle to your problem, you should be able to apply non-anthropic updating.
Equivalently-ish [1] it implies that it doesn’t matter when you started thinking anthropically; if you started doing it yesterday and then updated on one day’s observation, it would be the same as if you started doing it today, or twenty years ago and updated on twenty years of observations. And it also doesn’t matter that you could have been X, if you know that you’re not X.
Formally, this is anchor-free: there isn’t an anchor moment that defines the beginning of when you’re thinking anthropically [2] .
- ↩︎
Anchor-free is stronger than simple Bayes; simple Bayes is just the equivalent of anchor-free in situations where there is no anthropic uncertainty. Anchor-free is the real condition I want, but simple Bayes is enough for the impossibility result.
- ↩︎
One needs to be careful here; technically, any method with a specified anchor is “anchor-free” in a certain sense if it says “whenever you actually started thinking anthropically, then go back to this moment in the past (or future) and anchor then”. If the specified anchor is part of the definition of the method, then it’s “anchor-free” in that it formally doesn’t matter when you started thinking anthropically, because in all cases you just go to a specified anchor [3] and role out everything backwards/forwards from that I don’t consider “anchor specified” to be anchor-free.
- ↩︎
“Go to a specified anchor” is not actually that trivial to define. If the “anchor specified” truly behaves like “anchor-free” then it has to work from the very moment of first consciousness (as that’s certainly a place you could start anchoring from), when you could have been, potentially, any being in the universe. So there’s a universal anthropic probability that depends potentially on the details of all agents who ever were or ever could be.
- ↩︎
Society has a terrible track record with this kind of technology (see what liberal eugenics ended up looking like in practise—it wasn’t nazism, but it wasn’t anything positive, either).
Human intelligence is a very complicated system, playing out over the course of an individual’s life. All we have for the moment is evidence of small, local genetic effects. It would be a grave error to assume that these effects remain additive, without side-effects, when you combine them all. And measuring these side effects takes decades. Genetics isn’t like physics: you can’t go off-distribution (like sending a rocket to the moon) and assume the results hold up.
I think it wasn’t a coincidence that liberal eugenics went so wrong. The theory overpromised and was unsound, especially as you got into more details. It also played to people’s prejudices. This seems to be repeating the exact pattern—overpromising, unsound on the details, playing into people’s prejudices. Of course you don’t intend for things to go sour, and are advocating this for the best of motives—but the liberal eugenicists didn’t intend that, either, and good motives don’t guarantee good results.
At the very least, I’d want to see a detailed argument for why this won’t go wrong the way it’s gone wrong in the past. “This technology will only be used sensibly by sensible people under sensible governments” is not an assumption you get to make for free.
So on Sunday her epistemic state contains the statement “I am the original”, on Monday it does not.
Epistemic states can change freely without breaking martingales; on Sunday, she’s in room 1; put her to sleep and wake her up in room 1 or 2, randomly chosen; no duplications. Her epistemic state has moved from “I’m in room 1” to “I don’t know what room I’m in”, and yet the martingale works perfectly across this.
In any case, the statement is not “I am the original” but “on Monday, I was the original”, which is preserved.
I’ve also weakened the requirements; now it’s the law of total probability that is violated, not the martingale.
If you get
, you’re violating simple Bayes on the initial Sunday, which forces Q to equal the prior on worlds.
Yep, seems right.
Personally, I consider false memory to be a form of anthropic scenarios—you don’t know exactly who you are.
Yep!
I suppose it’s back to my old conclusion—decision theory is more fundamental that probability theory. https://arxiv.org/abs/1110.6437
That’s standard Sleeping Beauty; I crafted duplicate Sleeping Beauty so that there was no memory-loss issue. Sleeping Beauty always knows what day it is.
Are you saying that SIA+SSSA suffers from these problems?
SIA is well defined, so you can calculate with it, and notice the probability of Heads go from 1⁄2 on Sunday to 1⁄3 on Monday to 0 or 1⁄2 on Tuesday. The Monday → Tuesday transition is consistent with probability rules, the Sunday → Monday is not. Inclusion-exclusion (now replaced with the law of total probability) is one of the easiest ways to pinpoint the inconsistency.
SSA has multiple definitions, but in it’s “halfer” format, with reference class being Sleeping Beauty copies currently in that universe, it assigns 1⁄2 to Heads on Sunday, 1⁄2 on Monday, and 2⁄3 or 0 on Tuesday. It is Bayes consistent, but violates simple Bayes on Tuesday. Other SSAs may behave differently.
It assign probabilities for the agents inside to know their centered world (that is, all physical facts plus “what is here + now”), that’s enough for things like “estimating size of universe” or “figuring out if today is Tuesday” or “figuring out if coin is heads”, which people disagree on.
I feel that it probably has timeline inconsistencies, such as assigning a different size to the universe at different moments, without any observations to update on. The source of this feeling is that if it were consistent on timelines, then it would assign a reasonable internal agent distribution to future observations.
After all, if you’ve assigned probabilities to the day, the coin, how many agents there are, and the room, this should be now constraining your expected observations.
That design is SIA over full histories (see the end of the post), which breaks simple Bayes on Sunday.
The simplest version of SIA (C-SIA, counting SIA) upweighs worlds by the number of copies of you there is in them at any one time and is otherwise fully Bayesian. This is consistent with deaths, divergence of copies, and the initial creation of you(s). The impossibility result shows that it isn’t consistent with duplication, though.
A potential criticism of SSSA+SIA: if it doesn’t intrinsically assign subjective probabilities to agents inside the problem, what is it for? These problems aren’t difficult to analyse probabilistically from the outside; it’s from the inside that we need an effective tool.
What does this give as P(Monday|Sunday)? Seems it has to be zero.
Now obviously P(Monday|Sunday)=0 in the sense that “If it’s Sunday, it’s not Monday”, but how does SSSA+SIA encode “If it’s Sunday, what’s the probability that it will be Monday?” or “If I’ve observed Sunday, what’s the probability I will observe Monday?” It feels that it’s given up causality.
Unless I’ve made a mistake (very possible), my result shows that if SSSA+SIA can be turned into a personal probability distribution for Sleeping Beauty to use for predicting her experiences as well as world-facts, it will then break [1] in that form.
EDIT: I’ve replaced inclusion-exclusion with the law of total probability and simplified the proof further.
- ↩︎
The fundamental reason it will break is that h_1 and h_2 are mutually exclusive and exhaustive full histories for Sleeping Beauty, but any sensible probability system wants to put P(h_1)=1 and P(h_2)=1/2 (which are the correct probabilities for the existence of these histories, since they are not mutually exclusive from the outside).
The system has too few degrees of freedom; you can’t take the outside-observer probability distribution (which is what simple Bayes forces on Sunday and Tuesday), substitute in mutual exclusivity, and not have something break. Inclusion-exclusion is just the simplest thing that shows the break, currently.
- ↩︎
Because it’s not a case of virtue with a clear line (“never accept a bribe”) or virtue with an unclear line (“at what point does networking turn into nepotism”), but a case of careful judgement in the presence of strong confounding incentives. In that case, outside opinion is invaluable, but needs teeth.
Though thank you for the implied trust :-)