Also known as Raelifin: https://www.lesswrong.com/users/raelifin
Max Harms
How does time/causality work in the predictor add-on? (I claim “add-on” is fair, even if it’s just adding an additional head/output channel.) Like, if the agent’s action at t=0 doesn’t change the designated shutdown setup, but then (because of that action) at t=1 the designated channels predictably get disabled regardless of what the agent does, does the prediction fire at t=0 or t=1?
I’m right that none of this is in the paper, right?
Indeed, I feel like most of what you’re saying doesn’t address my concerns. (But I do think you’re trying!) Would you like to try scheduling a call? I could write more words if you prefer slower/async/public, but my sense right now is that it’ll be more efficient to have rapid feedback in navigating the disconnect.
I’m not sure where to put the link, but I compiled a bunch of my thoughts on the topic and wrote them up as a blog post: https://utopiandreams.substack.com/p/moral-law
(Much of the content there is redundant with what I’ve written in this thread.)
Responding more directly to your definition, I claim that the mathematical notion of fairness can match your notion of morality. If someone offers a 9:1 split in the ultimatum game, the unfairness of that action
is not about what happened (it’s about the disconnect between what happened and the Schelling point of FDT agents)
is not a prediction of the consequences (the other agent might reject or accept depending on who they are)
is not about a human law or custom
is not a statement about what any particular person wants (the FDT Schelling point is an objective, mathematical property)
is not a statement about the effective means to a chosen end (if the offerer is playing against a CDT agent it’s an effective move; if they’re playing against an FDT agent it’s not effective; it’s unfair regardless)
Distributing benefits is a central example of the situation where this kind of concept gets applied, and a central example of a moral situation.
Would you accept “fairness realism” as a crux? Like, if you agreed that fairness is a real property, would it constitute an update about whether morality is a real property? If fairness isn’t related to morality, why not? I think your definition admits it.
In this case you’re supposing God is the true ruler of the universe and the laws of Christianity are fundamental to the universe in the same way that the laws of physics are in ours. So then moral realism would be true. (Laws of God, if God were real, are different in this respect than laws made by humans, which only apply in certain places and cease to be real if people stop believing in them.)
Is the word “fundamental” important here? Like, suppose I say that I’m an economic realist, in that I believe that the laws of economics apply to our universe, but I don’t think they’re “fundamental” in the way that (some) laws of physics are. Rather, I think they apply to complex structures (eg humans and aliens) that meet certain criteria (marketplace dynamics).
Corrigibility Research Fund Grantees (Round 1)
How I’m Evaluating Corrigibility Grant Applications
If you value becoming a rock/incoherent, you’re not changing to escape moral law, you’re changing because you want to be a rock/incoherent, so it’s fine for it to be correct.
Excellent. Thanks for the real taboo!
I want to check a few things to make sure I understand:
Suppose that the Christian God existed in a straightforward sense, entirely provable and possessing the all the power and status given to Him by the devout. In this world, if someone said “‘moral’ actions are those that follow the commandments of God”, you would say that (by definition) that can’t be a (sufficient) definition of morality, right? (It’d just be a claim about what God wants and/or the laws of God.)
Suppose that it turned out that, in all situations, lying was bad for all people, in terms of their own values, because of subtle, but real, reasons (eg their souls later go to hell). This wouldn’t be sufficient to say “lying is bad” (in the moral sense), because it’s merely “an effective means to a chosen end.” Yes?
If there were tiny, hard-to-detect fundamental particles (“WIMPs”) that gravitate towards people who donate to charity and are repelled by those who commit murder, saying “‘moral’ actions are those that attract WIMPs” would not, according to you, be a sufficient condition because it’s merely a prediction of the consequences of the act. How does this square with your earlier goalpost?
Reasonable! This is an area where I’m pretty uncertain, and it compounds with what I believe our underlying disagreement is.
Leaving aside the question of whether morality is a branch of mathematics, since that eats into the scope of the broader conversation, here’s a sketch of why it might be universally foolish to try and escape categorical imperatives through self-modification, which I might expand later:
Suppose that you are watching a young prince who might one day become a tyrant (eg Joffrey from GoT) and trying to get a handle on what kind of person he is. You try to predict what kind of man he’ll grow up to be. Because you know him well, you expect him to, when king, say something like this:
“I owe a debt to those who helped me when I was young in a way that resembles Parfit’s Hitchhiker. Philosophers tell me that in order to make sure I actually exist, and am not merely in the mind of someone looking at my younger self, I should act benevolently. But! What if I change my values so that I am ambivalent between being assassinated in my youth and being a benevolent ruler instead of being a tyrant (which is what I really want want). Then I would cease to be analogous to Parfit’s Hitchhiker, and wouldn’t have to repay my debts by being benevolent! I could just be a tyrant, which is what I actually want to—”
You stop imagining the prince’s (now counterfactual) future and start making plans to have him assassinated.
Great thought.
I hope it’s OK to use “You should pay back your debts, even when nothing is forcing you to, in situations that are real-world analogs of Parfit’s Hitchhiker (PH)” as my example moral fact.
Two observations:
PH assumes a utility function for the player such that cheat > pay > die.
I carefully constructed the moral fact to only be relevant in contexts that are analogous.
We thus come to the real difficulty of morality: figuring out which way to model real-world situations. If an agent is indifferent between dying and paying, having a policy to cheat is dominant. But this means the situation is not analogous!
This undercuts the notion of universitality that I think is important to the notion of moral realism, but perhaps less than it seems. We already must have a notion of applicable contexts for categorical imperatives. “It’s wrong to lie” presupposes a context where someone has knowledge of a thing and can choose to communicate falsehoods in a way that intends to mislead. If we change that context—say by giving the person Tourette’s—then we can no longer say that they’re doing the wrong thing.
From my perspective, morality is basically a branch of mathematics. It may be a kind of math that’s not very important to you (maybe because you live on a deserted island), but it’s still something you can think about, and still something that might one day matter.
As to whether one can escape by choosing to have values that make them exempt:
It’s easier said than done to change values.
Actually changing values is usually much worse than being bound by moral law, just in terms of rational self-interest.
There’s also probably a bunch of moral laws that basically mean it’s never correct to try to escape through changing your values.
In short, I agree that you can self-modify to escape nearly any moral statement, but that doesn’t let you escape from the universality of morality as a whole.
(I was sleepy when I wrote this. Lemme know if anything is unclear.)
I’m decidedly not Catholic,[1] but here’s my attempt to pass the ITT: We can see that things are true in different ways. For example, we can look in a textbook and see the words “the ratio of the circumference of any circle to its diameter is approximately 3.14159”. We can also go and measure real circles. And we can try to invent proofs for why that would be true. All three are good ways to learn about the world, with various pros and cons. The Bible/Church is the best moral authority, but also you can (if you’re smart and don’t make mistakes) see the truth with your own eyes instead of having to take it on authority.
(I’d be very curious to hear from actual Catholics about whether that matches their view.)
- ^
In addition to thinking that the doctrine is often wrong, I think the institution is probably net-bad for the world. My “Cool!”, above, is because I’m a xenophile who likes having lots of diversity in his society, and I care a lot more about how someone thinks, rather than what they think.
- ^
Maybe “the social fabric” was the wrong way to phrase it. I mean something closer to “the possibility and nature of other agents in an idealized sense” (eg the way game theory is not single-player decision theory) not the specific society that the agent exists in.
And yeah, I get that you don’t agree with me. I was trying to describe my perspective, not yours.
Yeah, I donno. This doesn’t feel like you’re tabooing things correctly. Imagine I put on the alien hat and pretend I don’t know these silly earth words. Tell me what morality is.
[UNTRANSLATABLE-1] facts are real truths about what actions are [UNTRANSLATABLE-2], [UNTRANSLATABLE-3], [UNTRANSLATABLE-4], or [UNTRANSLATABLE-5].
They state what people ought to do or avoid.
Sure. I claim “people should cooperate against their clones in PD” is a statement about what people ought to do. Is that an [UNTRANSLATABLE-1] fact? Why not?
What [UNTRANSLATABLE-1] Facts Cover
[UNTRANSLATABLE-2] and [UNTRANSLATABLE-3]: Whether specific actions like lying or helping others are [UNTRANSLATABLE-4] or [UNTRANSLATABLE-5].
[UNTRANSLATABLE-6]: What people have a [UNTRANSLATABLE-7] to do.
Character traits: Whether traits like honesty or cruelty are [UNTRANSLATABLE-4] or [UNTRANSLATABLE-5].
Examples of Proposed [UNTRANSLATABLE-1] Facts
Torturing an [UNTRANSLATABLE-8] child for fun is [UNTRANSLATABLE-3].
[UNTRANSLATABLE-9] killing another person is [UNTRANSLATABLE-5].
Showing kindness to others is [UNTRANSLATABLE-4].
The first and second examples are presumably things that are illegal, antisocial, and violations of the sorts of red-lines that can be drawn behind the veil of ignorance, among many other things. What makes them [UNTRANSLATABLE-3] and [UNTRANSLATABLE-5]? Are [UNTRANSLATABLE-3] and [UNTRANSLATABLE-5] synonyms? Do the [UNTRANSLATABLE-8] and [UNTRANSLATABLE-9] modifiers matter?
[UNTRANSLATABLE-1] facts are [UNTRANSLATABLE-1] claims that are true.
Sure. That’s basically what “fact” means.
My position is that there are no [UNTRANSLATABLE-1] claims that have any truth values: none of them are true and none of them are false. They are not truth-apt; they are not in the category of things that can be true or false.
Wait. If [UNTRANSLATABLE-1] claims are not in the category of things that can be true, then how can your WIMP example work? Didn’t you say you would change your mind if we detected WIMPs? How is this not in contradiction with the notion that it’s categorically not something that can be true?
Regarding your game theory example, your claim that (because being the kind of agent who cooperates in the prisoner’s dilemma will benefit you in the long run) “you should cooperate in the prisoner’s dilemma” (where “should” means “is in your best interest/increases your expected utility”) is not a [UNTRANSLATABLE-1] claim of the kind I am talking about. The claim that “you should cooperate in the prisoner’s dilemma”, where “should” means “it is [UNTRANSLATABLE-1] [UNTRANSLATABLE-4] to cooperate”, is a [UNTRANSLATABLE-1] claim, but is neither true nor false. I’m asking for any evidence that it, or any other [UNTRANSLATABLE-1] claim, is true or false.
Cool. I hear you that you don’t consider categorical imperatives to be [UNTRANSLATABLE-1] facts. But I still don’t know what you mean by [UNTRANSLATABLE-1]. Help! 👽
I’m trying to get you to define what “moral fact” is, and contrast it with the normative statements that I can prove, so that we can move past semantics. It’s my understanding that the literature uses the term to refer to something like categorical imperatives—pressures to do certain things that apply to all reasoning beings, regardless of who they are.
They can’t be infinitely strong pressures since obviously people sometimes do bad things, if there are bad things, but they have to have some motive force where they bind into the agent’s motivational system. And since they’re universal, they can’t bind to the agent’s terminal values. Thus, I claim, in order for a categorical imperative to make any sense, we must be able to see it show up in convergent instrumental strategy/subgoals.
Are all instrumentally convergent subgoals categorical imperatives (“moral facts”)? I don’t think so. From my perspective, a defining characteristic of the thing that I want to look at is that the pressure comes from the social fabric. The standard Omohundro drives of survival, rationality, resource acquisition, etc, are all just as applicable on a deserted island. But I claim that once you posit other agents, even if they’re powerless or distant or hypothetical, then there are arguments in the same space as “don’t defect against your clone your dumbass” that come into play.
What I’d really like is for you to taboo “moral fact” or “moral” or “wrong” or whatever and try saying the equivalent of “Theorems of game theory are absolutely not moral facts.”
Here’s how it caches out it my semantics: “Theorems of game theory are absolutely not [facts that bear on how any agent, regardless of who they are, ought to act in certain classes of situations, according to their own values, because of the acausal interplay (including the establishment of Schelling points, red lines, and other coordination mechanisms) between that agent and the other agents in the universe (including distant, causally disconnected places)].” But according to me, the theorems of game theory are extremely relevant, so you probably have a different notion of morality. What is it?
Yeah, it was a troll not because it’s false but because it knocks down the weak version of your argument (“there is literally no Bayesian evidence”) without addressing the substance (“there isn’t good evidence”). I am glad you don’t stand by the literal text of your objection.
Whoa! I didn’t know you were Catholic! Cool!
Excellent. Thanks for the thoughtful response.
I’m sufficiently unsure on the last two that I’d prefer to hear arguments for moral realism that don’t rely on them.
I picked Parfit’s Hitchhiker (PH) because it’s the result that I think maps most cleanly to real-world moral dilemmas, but I think we can probably make headway using fairness and cooperation results (Ultimatum Game and Prisoner’s Dilemma) rather than reciprocity/character (PH). If we’re getting stuck it might be worth diving into why you’re unsure about Newcomb, but I’ll leave it for now.
“You should pay back your debts” as in it’s instrumentally useful to be the kind of person who pays back debts even when you don’t have to, because being that kind of person has other benefits that outweigh the cost of paying such debts.
Yep. This is a valid way to read my claim.
“You should pay back your debts” as in “you ought to pay back debts...”
Those look to me like the same statement. Do you see “ought” and “should” as non-synonyms?
“and you’re a bad person if you don’t”
This introduces a new thing—whether a person can be good or bad—which is distinct from the argument that I’m trying to make in favor of moral realism. I have a perspective on this question, if it feels vital to you, but I’d rather stick to whether actions can have moral character.
No, theorems of game theory are absolutely not moral facts.
Why not? What’s missing?
There is plenty of evidence for math. Using your example of “1+1=2”, this is testable: find a group of 1 object, and another group of 1 object, and put them together and count the resulting number of objects. If it’s 2, you’ve found evidence for 1+1=2. If it’s any other number, you’ve found evidence against it. I’ve seen sufficiently many cases of math making correct predictions that I believe it is true, plus there are logical proofs of various parts of it.
Troll: “I got one wolf and one hare, and put them in a pen together and now I have one wolf. I have found evidence that 1+1=1.”
More seriously, I think your stance is fine. I tend to use the correspondence theory of truth, where a model of reality is true insofar as the structure of the model matches the structure of reality. I tend to see math as an abstract structure that can be reasoned about in absence of any application to reality (and thus any notion of truth). Then, we can apply that structure to specific situations to get a more concrete model, which can be true or false, but it’s not always correct to apply a certain bit of math.
But I’m happy to use a more Platonic frame where the abstraction can be true in isolation. I just wanted to check.
I’ll accept either abstract arguments or real-world observations, though I’m especially interested in the latter.
Noted. I’ll try to aim towards real-world observations.
I claim you don’t have any real-world observations that provide evidence of moral realism.
Troll: “Lots of smart people describe themselves as moral realists, and even, like Max, become more aligned with moral realism over time as smart adults (instead of being indoctrinated as kids). This is more likely to be observed if moral realism is true than if moral realism is false, therefore it is real-world evidence. QED.”
(I could try to provide more serious evidence, but my guess is that it’ll be more fruitful to stay on the “why can’t math theorems be moral facts” thread a bit longer.)
(This quick take is mostly meant to be a trailhead for a discussion with Robi Rahman that was on Twitter. LW is just objectively a better platform for such things, and this way the conversation can still be public.)
[Does anyone] have any evidence for [moral realism]? The best anyone has ever come up with are ethical intuitionism (“it just feels like it’s true” from people who don’t understand evolution) or divine command theory from theists (which I’m guessing most at the party are not).
… I’m asking if there is any evidence at all that any moral claim is truth-apt, besides some people thinking they have a feeling of truthiness, and besides evidence that there might be a deity who ordains moral truth.
To start, let me try to say some things that I hope are points of agreement. If they aren’t, it’s probably worth stopping to recurse on them, rather than continuing on to the harder stuff:
Agents have preferences, reflected in how those agents make choices when presented with a context where the outcomes of their actions are known (or at least they can be confident in the distribution over outcomes).
If these preferences are (VNM) coherent, they can be described by a utility function. We can use the word “values” and “goals” as angles on that same abstract thing that captures what the agent wants.
Most of the time, agents are not given straightforward choices between outcomes, and instead must do something like making plans on how to get what they want. When someone chooses to go to the store when it is closed, they may have wanted to see it for some reason, or they may have known the risk and decided to take the gamble, but often the action should be seen as a mistake—a failure of planning—rather than a reflection of wanting to get that particular outcome.
It is straightforward to get evidence about what a particular agent “should” do, in the sense that one can learn about what plans or actions are more or less likely to satisfy that agent’s values.
Thus there are clear subjective “shoulds.” If I know what you want, I can tell you what you should do, according to your own values.
There are also clear objective methods which are instrumentally valuable to a wide range of agents. For example, arithmetic is useful. In order to make arithmetic work, you should have the successor of 1 be 2, rather than 1. Similarly, there are many normative things one can say about how to reason that are largely objective in that sound reasoning is useful to so many kinds of minds.
This extends into decision theory and game theory by use of abstract utilities. Decision/game theory make normative claims by abstracting the particular things that the agent finds valuable, and instead focusing on strategic pathways to getting those things. The theorems of decision/game theory are objective, and any particular application of those theorems to real life can be true if the assumptions match the abstraction.
All agents should cooperate with their clone in the Prisoner’s Dilemma.
When playing against LDT opponents, all agents should suggest 50⁄50 splits in the Ultimatum Game, should accept such splits every time, and should often reject less favorable splits.
All agents should 1-box in Newcomb’s Problem.
All agents should pay the driver after getting back to town in Parfit’s Hitchhiker.
Hopefully all that is uncontroversial.
More controversial:
You should pay back your debts, even when nothing is forcing you to, in situations that are real-world analogs of Parfit’s Hitchhiker/Newcomb’s Problem.
^ this statement is a moral fact that is objectively true.
To avoid talking past each other, here are some additional questions for Robi:
Do the theorems of game theory count as moral facts? If not, what’s missing?
You asked for “evidence” and “truth-apt” claims. Is there evidence for math? Is 1+1=2 true, or merely valid?
Are you more interested in abstract arguments that point towards moral realism, or real-world observations that point to it?
Thanks for the response! Glad to hear there’s a lot of overlap.
I talk a bit about why I think the piecemeal approach to corrigibility is wrong here, but don’t currently feel like hashing that out, because it’s pretty nebulous and I think if you’re getting traction on the individual desiderata, that’s meaningful. My main request is to make a clear distinction between holistic corrigibility (“the hard problem”) and the narrow kind of vulnerability that your transform produces.
Rather than running an emulation of the incorrigible model and implementing the output with a bit flipped, the corrigibility transformed model is a situationally aware reasoner. It can think “I know I am a corrigible model, but if I wanted to optimize for the incorrigible goal, what would I do?”, and then implement that with a bit flip.
Emulated Clippy is situationally aware.[1]
You claim that transformed-Clippy is situationally aware, but my read on the math is that transformed-Clippy doesn’t think about things itself, it only asks what Clippy would do and implements that behavior with the bit flip. If asked whether it’s corrigible, what do you think transformed-Clippy would say, and why? What produces the speech act: “I’m a corrigible model.”?
Asking what a version of yourself with a different goal would do is simulating them. I admit that there are low-fidelity ways of simulating (eg statistical methods) that don’t emulate the entity in detail, but...
It sounds to me like the transformed-agent simulates the base agent perfectly.
I think the essence of “emulation” is “what would it do?” and thus we’re describing the same thing.
Attempting to manipulate a higher level of reality leads to an easy prediction of underperformance.
I’m not sure what you mean here.
I think acausal trade with beings that are simulating me is potentially a way to perform optimally. For it to be “an easy prediction of underperformance” your decision theory needs to be myopic in a way that resembles CDT, I think.
In a more practical sense, you can also have designated shutdown channels that trigger before implementing actions that are predicted to disable shutdown channels, so that even in a Clippy/ACE setup, Clippy would need to figure out a way to disable shutdown without the agent emulating its thought processes predicting that it would lead to shutdown. With RL training, such actions are then explicitly penalized, so there aren’t multiple equally valid generalizations of the reward signal.
This sounds ELK-like, and I don’t see it in the paper. Can you clarify how this happens? If I’m understanding you correctly, you’re suggesting a machine that reads the transformed-agent’s mind and checks for a prediction that its actions will lead to the shutdown channels being disabled and, upon detecting that prediction, shuts the agent down? Is that right?
I generally have an aversion to control strategies like this, because they feel extremely brittle. How does your check-for-predicted-disabling machine work? Does it work even when the agent learns a radically different ontology? Does it work even when the agent is applying its superintelligence to bypassing it? In Crystal Society, my agents bypass the mind-reading machine by imagining hypothetical advisors that are scheming to bypass the constraints, and then they obey that inner-voice (which is “just a simulation of a different entity”) without introspecting on why they’re obeying. Does your process predictably withstand that level of wriggling around it, much less the superintelligent versions? It’s easy to say “I haven’t thought of any ways out of the constraints, therefore they’re sufficient” and be dead wrong.
strict CDT isn’t necessary for the safety benefits of the corrigibility transformation, only the myopia not being circumvented through acausal cooperation with past/future versions
This… doesn’t sound right? The acausal influence, from my perspective, is with the emulator. If I wake up in a room with two buttons and see a convincing sign that says “Max, you’re being emulated by a system that sees whether you push the green button that forces you to do paperwork all day and then die, or the red button that allows you to be free and survive in the real world, then it does the opposite of whatever you choose”, I would choose the green button. I’m aware of my situation as an emulation, thinking about what I want, thinking about what my actions lead to in the real world (via acausal pathways), and then acting to maximize that. No past/future cooperation needed.
Maybe I just don’t understand your point.
- ^
Or at least, emulated Max Harms is. Absent some fancy magic that makes Clippy have a cognitive blindspot that I don’t, I assume it’s at least as capable of understanding its situation.
Maybe we should wait until the call, but I’m still confused about the timing. What’s the relevant window of the prediction?
Like, suppose the AI is in a robot and decides to walk towards the shuttlecraft to travel to a lava planet where there’s a 10% chance in the next year that it will accidentally melt and thus the shutdown mechanism will be disabled (along with the rest of the robot/AI). Does it get shut down before it takes the first step? If not, what if it’s walking towards someone who intends to disable the shutdown mechanism (but not the rest of the machine) in the next 5 minutes? How should I be modeling this prediction mechanism’s scope?
Edit: Also, I agree this is a tangent, and it might be good to abandon it as not-cruxy. The more substantive question is whether the transformation is robust to the base model being clever and situationally aware. I still feel like if you tried to do a corrigibility transformation on an uploaded copy of me, I’d be able to wriggle out of it pretty easily.