Chris van Merwijk(Chris van Merwijk)

Karma: 692

Chris van Merwijk 11 Jul 2018 0:22 UTC
2 points
in reply to: Vaniver’s comment on: Two agents can have the same source code and optimise different utility functions
You don’t necessarily need “explicit self-reference”. The difference in utility functions can also be obtained due to a difference in the location of the agent in the universe. Two identical worms placed in different locations will have different utility functions due to their atoms being not exactly in the same location, despite not having explicit self-reference. Similarly, in a computer simulation, the agents with the same source code will be called by the universe-program in different contexts (if they weren’t, I don’t see how it makes sense to even speak of them as being “different instances of the same source code”. There would just be one instance of the source code.).
So in fact, I think that this is probably a property of almost all possible agents. It seems to me that you need a very complex and specific ontological model in the agent to prevent these effects and have the two agents have the same utility function.

Chris van Merwijk 11 Jul 2018 0:44 UTC
1 point
in reply to: cousin_it’s comment on: Two agents can have the same source code and optimise different utility functions
Here are some counterarguments:
There can be scenario’s where the agent cannot change his source code without processing observations. e.g. the agent may need to reprogram himself via some external device.
The agent may not be aware that there are multiple copies of him.
It seems that for many plausible agent designs, it would require a significant change in the architecture to change his utility function. E.g. if two human sociopaths would want to change their utility function into a weighted average of the two, they couldn’t do so without significantly changing their brain architecture. A TDT agent could do this, but I think it is not prudent to assume that all actually future existing AGI’s we will deal with will be TDT’s (in fact, most likely most of them won’t be it seems to me).
So I don’t think your comment invalidates the relevance of the point made by the poster.

Chris van Merwijk 23 Aug 2019 6:21 UTC
4 points
on: Tabooing ‘Agent’ for Prosaic Alignment
I endorse this. I like the framing, and it’s very much in line with how I think about the problem. One point I’d make is: I’d replace the word “model” with “algorithm”, to be even more agnostic. “Model” seems for many people already to carry an implicit intuitive interpretation of what the learned algorithm is doing, namely “trying to faithfully represent the problem”, or something similar.

Chris van Merwijk 16 May 2020 11:06 UTC
1 point
in reply to: Pongo’s comment on: Subspace optima
I made up the term on the spot, so I don’t think so.

Chris van Merwijk 18 Oct 2020 9:02 UTC
3 points
in reply to: Vanessa Kosoy’s comment on: Moloch games
Yes that’s what I meant, thanks.

Chris van Merwijk 5 Apr 2021 14:54 UTC
9 points
on: Predictive Coding has been Unified with Backpropagation
I suspect a better title would be “Here is a proposed unification of a particular formalization of predictive coding, with backprop”

Chris van Merwijk 8 Apr 2021 7:35 UTC
8 points
on: Violating the EMH—Prediction Markets
Is there currently a way to pool money on the trades you’re suggesting? In general it seems like there is some economies of scale to be gained by creating some kind of rationalist fund

Chris van Merwijk 18 Jun 2021 14:27 UTC
3 points
on: How to Throw Away Information
I might misunderstand something or made a mistake and I’m not gonna try to figure it out since the post is old and maybe not alive anymore, but isn’t the following a counter-example to the claim that the method of constructing S described above does what it’s supposed to do?

Let X and Y be independent coin flips. Then S will be computed as follows:
X=0, Y=0 maps to uniform distribution on {{0:0, 1:0}, {0:0, 1:1}}
X=0, Y=1 maps to uniform distribution on {{0:0, 1:0}, {0:1, 1:0}}
X=1, Y=0 maps to uniform distribution on {{0:1, 1:0}, {0:1, 1:1}}
X=1, Y=1 maps to uniform distribution on {{0:0, 1:1}, {0:1, 1:1}}
But since X is independent from Y, we want S to contain full information about X (plus some noise perhaps), so the support of S given X=1 must not overlap with the support of S given X=0. But it does. For example, {0:0, 1:1} has positive probability of occurring both for X=1, Y=1 and for X=0, Y=0. i.e. conditional on S={0:0, 1:1}, X is still bernoulli.

Chris van Merwijk 23 Aug 2021 10:28 UTC
1 point
AF
on: Finite Factored Sets: Orthogonality and Time
Can’t you define $C ⊢^{S} X$ for any set $C$ of partitions of $X$ , rather than $C ⊢^{F} X$ w.r.t. a specific factorization $F$ , simply as $C ⊢^{S} X$ iff $⋁_{S} (C) \geq_{S} X$ ? If so, it would seem to me to be clearer to define $⊢$ that way (i.e. make 7 rather than 2 from proposition 10 the definition), and then basically proposition 10 says “if $C$ is a subset of factors of a partition then here are a set of equivalent definitions in terms of chimera”. Also I would guess that proposition 11 is still true for $⊢^{S}$ rather than just for $⊢^{F}$ , though I haven’t checked that 11.6 would still work, but it seems like it should.

Chris van Merwijk 23 Aug 2021 13:19 UTC
1 point
on: Finite Factored Sets: Orthogonality and Time
In the proof of proposition 18, “part 3” should be “part 4″.

Chris van Merwijk 23 Aug 2021 13:42 UTC
4 points
on: Finite Factored Sets: Orthogonality and Time
I just want to point out some interesting properties of this definition of time: Let time_C refer to the classical notion of time in a dynamical system, and time_FFS the notion defined in this article.

1. Suppose we have a field on space-time generated by a typical differential dynamical law that satisfies time_C-reversal symmetry, and suppose we factorize its histories according to the states of the system at time_C t=0. Then time_FFS doesn’t make a distinction between the “positive” and “negative” part of the time_C. That is, if x is some position (choose a reference frame), then the position (x,2) in space-time (i.e. the value of the field at position x at time_C 2) is later in time_FFS than (x,1), but (x,-2) is also later in time than (x,-1). In this sense, the time_FFS notion of time seems to naturally capture the time-reversal symmetry in the laws of physics: Intuitively, if we start at the big bang, and go “backward in time” we are just as much going into the future as we are if we would go “forward in time”. Both directions are the future.

2. However, more weirdly, time_FFS also allows a comparison between the negative-time_C and positive-time_C events. Namely, (x,1) happens before_FFS (x,-2) while (x, −1) happens before_FFS (x,2). I am not sure what to make of this, or whether we should make anything of it.

3. Suppose a computer is implemented in the physical world and implements a deterministic function $f : X \to Y$ , AND we restrict to the set of histories in which this computer actually does this computation. Now let x denote the variable that captures what input is given to this computer (meaning, the data stored in the input register at one particular instance of running this algorithm), and y similarly denote the variable that captures what the output is, then y occurs (weakly) earlier_FFS than x, even though the variable x is defined to be earlier than y (more precisely, to directly apply the definitions of x and y to check their value in a particular history h would involve doing a check for x at a time_C that is earlier_C than the check for y). I’m not sure what to make of this, though it kind of seems like a feature not a bug. If we don’t restrict to the set of histories in which the computer does the computation, I’m pretty sure this result disappears, which makes me think this is actually a desirable property of the theory.

Chris van Merwijk 24 Aug 2021 8:08 UTC
1 point
AF
on: Finite Factored Sets: Conditional Orthogonality
I think a subpartition of S can be thought of as a partial function on S, or equivalently, a variable on S that has the possible value “Null”/”undefined”.

Chris van Merwijk 28 Sep 2021 9:23 UTC
2 points
in reply to: VojtaKovarik’s comment on: Moloch games
Yeah I suppose that you’re taking an essential property of a Moloch to be that it wants something other than the sum of utilities. That’s a reasonable terminological condition I suppose, but I’m addressing the question of “what does it even mean for ‘society’ to want anything at all?” Then whatever that is, it might be that (e.g. by some coincidence, or by good coordination mechanisms, or because everyone wants the same thing) what society wants is the same as what would be good for the sum of individual utilities. It seems to me that the question of “what does society want?” is more fundamental than “how does that which society want deviate from what would be good for its individuals?”

Chris van Merwijk 28 Sep 2021 10:04 UTC
2 points
in reply to: VojtaKovarik’s comment on: Moloch games
Maybe. I actually don’t think the term “Moloch” is very important. What I think is important is getting a good conceptual understanding of the behavioural notion of “what society wants”, behavioural in the sense that it is independent of idealized notions of what would be good or what individuals imagine society wants but depends on how the collection of agents behaves/is incentivized to behave. I view the fact that this ends up deviating from what would be good for the sum of utilities, as essentially the motivation for this topic, but not the core conceptual problem. So I’d want to nudge people who want to clarify “Molochs” to focus mostly on conceptually clarifying (1) and only secondarily on clarifying (2).

Secondarily, just to push back against your point that “Moloch” is historically more connotated with (2). This is sort of true, but on the other hand, what does the concept of “Moloch” add to our conceptual toolbox, above and beyond the bag of more standard concepts like “collective action problem” and “externalities” and so forth? I’d say that it is already well-understood that collections of individuals can end up interacting in ways that is globally pareto-suboptimal. I think the additions to this analysis made in SSC are something like: conceptualizing various processes as optimizing in a certain direction/looking at the system-level for optimization processes. The core point to get clarity on here is I think (1), and then (2) should fall out of that.

Chris van Merwijk 24 Oct 2021 10:15 UTC
18 points
on: Countable Factored Spaces
I suggest renaming this to “countably factored spaces”. Countably being a property of the factorization rather than the space.

Also I suggest adding an actual self-contained definition of countable factored space to make it more readable.

Chris van Merwijk 21 Mar 2022 7:28 UTC
3 points
in reply to: Dagon’s comment on: Natural Value Learning
The thing underlying the intuition is more something like: We have a method of feedback that humans understand and that works fairly well, and is adapted to the way values are stored in human brains. If we try to have humans give feedback in ways that are not adapted to that, I expect information to be lost. The fact that it “feels natural” is a proxy for “the method of feedback to machines is adapted to the way humans normally give feedback to other humans” without which I am at least concerned about information loss (not claiming it’s inevitable). I don’t inherently care about the “feeling” of naturalness.
Regarding no Safe Natural Intelligence: I agree that there is no such thing, but this is not really a strong argument against? This doesn’t make me somehow suddenly feel comfortable about “unnatural” (I need a better term) methods for humans to provide feedback to AI agents. The fact that there are bad people doesn’t negate the idea that the only source of information about what is good seems to be stored in brains and that we need to extract this information in a way that is adapted to how those brains normally express that information.
Maybe I should have called it “human-adapted methods of human feedback” or something.

Chris van Merwijk 21 Mar 2022 7:28 UTC
3 points
in reply to: Charlie Steiner’s comment on: Natural Value Learning
I haven’t specified anything about the algorithms, but they will maybe somehow have to be different. The point is that the format of the human feedback is different. Really this post is about the format in which humans provide feedback rather than about the structure of the AI systems (i.e. a difference in method of generating the training signal rather than a difference in learning algorithm).

Chris van Merwijk 21 Mar 2022 7:41 UTC
3 points
in reply to: Steven Byrnes’s comment on: Natural Value Learning
Not really a fair characterization I think: 2 mostly seems orthogonal to me (though I probably disagree with your claim. i.e. most important things are passed from previous generations. e.g. children learn that theft is bad, racism is bad etc, all of those things are passed from either parents or other adults. I don’t care a lot about the distinction parents vs other adults/society in this case. I know about the research that parenting has little influence, I don’t want to go into it preferably). 1 seems more relevant. In fact maybe the main reason for me to think this post is irrelevant is that the inductive biases in AI systems will be too different from that of humans (although note, genes still alow for a lot of variability in ethics and so on). But I still think it might be a good idea to keep in mind that “information in the brain about values has a higher risk to not get communicated into the training signal if the method of elliciting that information is not adapted to the way humans normally express the information”, if indeed it is true.

Chris van Merwijk 21 Mar 2022 16:28 UTC
3 points
in reply to: Steven Byrnes’s comment on: Natural Value Learning
“I tend to think that learning and following the norms of a particular culture (further discussion) isn’t too hard a problem for an AGI which is motivated to do so”. If the AGI is motivated to do so then the value learning problem is already solved and nothing else matters (in particular my post becomes irrelevant), because indeed it can learn the further details in whichever way it wants. We somehow already managed to create an agent with an internal objective that points to Bedouin culture (human values), which is the whole/complete problem.

I could say more about the rest of your comment but just checking if the above changes your model of my model significantly?

Also regarding “I think I’m much more open-minded than you to …”: to be clear, I’m not at all convinced about this I’m open to this distinction not mattering at all. I hope I didn’t come accross as not open minded about this.

Chris van Merwijk 28 Mar 2022 5:17 UTC
1 point
in reply to: Yitz’s comment on: Manhattan project for aligned AI
I’m not sure I understand the motivation behind question. How much of my modern knowledge am I supposed to throw away? Note I am not in fact an atomic theorist who has the state of knowledge of atomic theory in 1942 so it’s hard to know what I’d think, but I can imagine assigning somewhere between 5% and 95% depending on how informed of an atomic theorist I actually was and what it was actually like in 1942. Maybe I could give a better answer if you clarify the motivation behind the question?