Respect
Ramana Kumar
Let me know when you can receive donations via a UK charity.
Vaguely related perhaps is the work on Decoupled Approval: https://arxiv.org/abs/2011.08827
Dialogue on What It Means For Something to Have A Function/Purpose
Thanks for this! I think the categories of morality is a useful framework. I am very wary of the judgement that care-morality is appropriate for less capable subjects—basically because of paternalism.
Just to confirm that this is a great example and wasn’t deliberately left out.
Consent across power differentials
I found this post to be a clear and reasonable-sounding articulation of one of the main arguments for there being catastrophic risk from AI development. It helped me with my own thinking to an extent. I think it has a lot of shareability value.
I think this is basically correct and I’m glad to see someone saying it clearly.
I agree with this post. However, I think it’s common amongst ML enthusiasts to eschew specification and defer to statistics on everything. (Or datapoints trying to capture an “I know it when I see it” “specification”.)
The trick is that for some of the optimisations, a mind is not necessary. There is a sense perhaps in which the whole history of the universe (or life on earth, or evolution, or whatever is appropriate) will become implicated for some questions, though.
I think https://www.alignmentforum.org/posts/TATWqHvxKEpL34yKz/intelligence-or-evolution is somewhat related in case you haven’t seen it.
I’ll add $500 to the pot.
Interesting—it’s not so obvious to me that it’s safe. Maybe it is because avoiding POUDA is such a low bar. But the sped up human can do the reflection thing, and plausibly with enough speed up can be superintelligent wrt everyone else.
A possibly helpful—because starker—hypothetical training approach you could try for thinking about these arguments is make an instance of the imitatee that has all their (at least cognitive) actions sped up by some large factor (e.g. 100x), e.g., via brain emulation (or just “by magic” for the purpose of the hypothetical).
It means f(x) = 1 is true for some particular x’s, e.g., f(x_1) = 1 and f(x_2) = 1, there are distinct mechanisms for why f(x_1) = 1 compared to why f(x_2) = 1, and there’s no efficient discriminator that can take two instances f(x_1) = 1 and f(x_2) = 1 and tell you whether they are due to the same mechanism or not.
Will the discussion be recorded?
(Bold direct claims, not super confident—criticism welcome.)
The approach to ELK in this post is unfalsifiable.
A counterexample to the approach would need to be a test-time situation in which:
The predictor correctly predicts a safe-looking diamond.
The predictor “knows” that the diamond is unsafe.
The usual “explanation” (e.g., heuristic argument) for safe-looking-diamond predictions on the training data applies.
Points 2 and 3 are in direct conflict: the predictor knowing that the diamond is unsafe rules out the usual explanation for the safe-looking predictions.
So now I’m unclear what progress has been made. This looks like simply defining “the predictor knows P” as “there is a mechanistic explanation of the outputs starting from an assumption of P in the predictor’s world model”, then declaring ELK solved by noting we can search over and compare mechanistic explanations.
I haven’t been very much engaged with this community in a while. I appreciate what Richard is doing here. My 2c for now is that what the EA/rationalist/alignment folks need to, and struggle to, fully digest and abandon is eugenics (and its milder cousin meritocracy). Because this “community” (used loosely) typically comes from a starting point that holds entrenched strongly favourable views and beliefs about eugenics etc. (even though I wouldn’t think they endorse them explicitly in those terms for the most part) they are taking a very long time to take in and/or rediscover the criticism and ultimate demolition of that whole belief and value system. I think it’s highly related to the more obvious (and described in this sequence) propensity to align with power, fail to critically evaluate power structures and systems, or recognise concentration of power itself as the main cause of harm.