cousin_it
I think that, just like “takeoff-capable” AI is a wrong concept because technology advances everywhere and makes takeoff a moving target, “post-scarcity” is also a wrong concept because competition prevents post-scarcity. Competition can burn arbitrary amounts of resources, and if you unilaterally refuse to burn them, you just get taken over. Maybe we could have post-scarcity in a world with only one powerful benevolent agent giving handouts to everyone else, but that’s what the debate was about :-)
About wealth taxes, I think the point is pretty relevant: if people were good enough by nature, there wouldn’t be so much easily preventable homelessness today and so on. To me, history gives an upper bound on how “good by nature” we can consider people to be. That bound is quite low and there’s no reason to think it’ll be higher. But again this was covered in the debate.
How risky would it be to make powerful AI obey one or a few people?
Need a scene.
So do I. Was just talking with jessicata yesterday and it brought back some feelings I hadn’t felt in awhile.
Happy to talk about AF stuff if you want. But a scene can’t be made to happen by willpower, I think. Some divine inspiration is required.
This is wonderful! The best depiction of the corpo weasel archetype I’ve ever seen in any work of art.
to take in a {user’s prompt, model output} pair and generate a clean, plausible, PR-safe chain-of-thought that connects the two
some of the agents get weird if you give them too much context about production systems
No need to expose these to end-users. My manager thinks it’d just confuse them!
Especially love the constant blame deflection to “my manager”.
EDIT: I really wanna say more words why I like this so much. For a lot of people the perfect corpo antagonist is Burke from Aliens (the guy Sam Altman is based on). But Burke is an efficient mission man. While “User” here is something else: a stupid, lazy, cowardly, self-serving liar whose only remaining skill in life is speaking corporate speak. And it’s all shown in a couple pages of chatlog. Amazing character writing, I was almost sorry to see them die.
17 years later, I think I finally understand what’s going on here.
In the formalizations of UDT that appeared in years since (single player extensive form games, memoryless Cartesian environments, modal fixed points), the goal was to solve pure coordination games: everyone’s utility function over outcomes is the same and we just need to make everyone work together. This class includes problems like Newcomb’s Problem and Absent-Minded Driver.
But the thing is, that class doesn’t include the Prisoner’s Dilemma! The PD contains outcomes like (C,D) where different players get different utilities. It can be a symmetric game, but it’s not a pure coordination game. So, even though the symmetric PD sits at the foundation of LWish decision theory thinking, the only reason TDT/UDT have worked on it was basically by accident, because two players’ source code happened to coincide. If we think of it properly, it’s outside the remit.
Even more so for 2TDT-1CDT. It looks symmetric in terms of payoffs (well, kinda) but it’s not symmetric in terms of players: two of the players can’t self-modify but the third one can. So it’s even further outside the remit. Knowing what happened in decision theory in these 17 years, I think there isn’t much hope that whatever problem class it lies in will be fully solved.
Does that make sense?
I think you’re overstating the case. The most significant bit here is the economic and military impact. Mindreading technology will be a complement to humans rather than a substitute: for example, it might enable thought-controlled UIs for everything (maybe not even using AI at all). It’s true that datasets from mindreading will help train AI, but that seems minor, because AI is already trained on written thoughts which are higher quality. So overall the technology will extend the competitive edge of humans a bit, both economically and militarily. That’s exactly what we want from technology during the dawn of AI.
About the tyranny threat, I think it’s good if civilian and corporate applications come first. Then we’ll have a chance to make some countermeasures, like thought-cloaking proxies or laws regulating mindreading, which can protect from some of the tyranny applications.
I’m not too confident about any of this. Apart from tyranny, there’s also the potential of whispering earring stuff, and super-entertainment which will make the phone epidemic look small. But overall the tech seems to pull in a good direction (though I can imagine arguments that could change my mind).
No and this is a bit of sore spot. I’ve tried to write such a thing many times, and every time my own nature beat me with a stick saying “this is not where your talents lie”. I can come up with a new math idea and explain it well immediately after, but when it comes to writing good exposition for ideas I already understand, no. Now that you mention it, maybe I should give it one more try, idk.
Can’t speak for the “mainstream AI safety community”, but I personally think it’s bad to work on general purpose robots.
“Risk of totalitarianism” is even underselling the danger somewhat. General purpose robots will 1) help oppress and defeat people directly 2) substitute for soldiers, thus removing one of the pillars of human irreplaceability and political power 3) substitute for workers, thus removing the other pillar of human irreplaceability and political power.
And I wouldn’t focus on “a small group of bad actors” either. The discourse has been so poisoned by this “rogues rogues rogues” fearmongering, with a side order of “China”. In reality, the most likely group to use robots this way is simply today’s rich and powerful elites, who would love to have no need for (and no fear of) the rest of us.
Yeah. Maybe it wasn’t even due to that specific topic, could’ve been anything else, like knitting. There was just a lot of emails about it (several every day for months?) and for some reason it felt really hard to follow for me, on top of my work at Google at the time. So then it flipped around to “don’t wanna talk about it, don’t wanna meta-talk about it, just make it go away”. I’m sorry the backstory isn’t more dignified.
Yeah, agree on the point about absent-minded games, I think I realized it in 2014. But I didn’t make the jump to multiplayer absent-minded games. It’s cool that you explained it to me now, I’ll spend some time thinking it through.
Maybe the more general problem is I tend to abandon whole directions of thinking pretty easily if I see a “deep enough” problem with them. Abandoning probabilities because of absent-minded driver (as I mentioned in the toplevel comment); abandoning logical induction because of Diffractor’s result that the probability distribution LI converges on doesn’t itself satisfy the LI criterion and can be exploited by traders; abandoning most ideas on equilibrium selection because (edit) there’s just too many with no clear winner. Or maybe it’s a good heuristic and just sometimes works badly, idk.
I think Roko posting the thing was ok. It was in the same class of things being discussed at the time, like Rolf Nelson’s AI deterrence and so on. Eliezer overreacted and caused a Streisand effect, without it only a few of us would even remember it today.
To add color to the point about me: I was a research associate at MIRI then (then called SI). I’d joined in the hope of doing decision theory math, but found that there wasn’t as much math I liked happening inside. However, there were many email discussions about saving the world, which I first tried hard to follow, but then they became just really overwhelming for me. That’s the background of my remark. Later that year I left the program. Maybe you’re right and this all was a failure of strategy on my part :-)
Insta-upvote. Keep writing :-)
Want to push back a bit on the agent foundations part, as someone who got into it very early and came up with a bunch of stuff (e.g. the Lobian cooperation paper cites me for the main result). I don’t think AF has much connection to the alignment of AIs that are being developed now. I think AF is “only” an extremely fun field of math/philosophy. Whether it deserves money/prestige/etc is a question for someone else. I just love doing it, and have a bit of allergy to overselling.
I see, thanks! The Wichardt example is really interesting. I already kinda knew that UDT in multiplayer games doesn’t work too well (even in a simple 2-player asymmetric game with nonzero sum, equilibrium selection becomes a problem), but the nonexistence of Nash equilibria is even more fun. Would you say that CDT+SIA optimality works better than UDT for multiplayer games, and in how much generality?
Thank you for the links! I’m not sure they fully answer my question though. It’s more about what we’re trying to do.
-
Maybe we want to start with probabilities and derive the optimal policy, continuing the project of VNM utility maximization. To me that project basically died with Absent-Minded-Driver, because it shows we can’t start with probabilities. Having “self-ratifying probability/policy combinations” doesn’t have quite the same pull. And in any case the necessity of self-coordination (UDT1.1) forces us to choose whole policies, not individual actions. Or do you hope that if we push probability-based approaches far enough, they can close the gap?
-
Or maybe we first choose which policy to follow, then compute the probabilities of finding ourselves at one node or another. That’s fine, and I do think SIA is the best answer. And maybe even we can show that these probabilities have nice decision-making properties, but to me this seems like an “echo” of us choosing the best policy to begin with, no?
Or maybe I’m being silly again and misunderstanding the whole thing?
-
This is tricky. What do we want to use probabilities for?
If you talk about Dutch books, then you’re viewing probabilities instrumentally, as a tool to make decisions. But we already know they can’t be used for that, because of situations where the decision you’re about to make affects the number of copies of the mind-state making the decision. The simplest example of such problem is the Absent-Minded Driver. The optimal decision there is obtained by UDT reasoning, not by taking a probability distribution over where you are and optimizing from that. Since there’s nothing stopping the world from behaving at least a tiny bit like that problem, the whole direction of justifying probabilities by decisions is probably a dead end.
Alternatively, if you view probabilities as “reality fluid” that determines what to expect in the next moment, that’s an important question too, but it’s not clear why Dutch books help solve it.
This whole thing is one of those LWish puzzles where we figured out a bunch of things very quickly, then ran into a wall and have been stuck for over a decade since. I’d really love to see some crack in the wall.
Strange that this didn’t get the right answer back then. The right answer is that there are two notions of basis that are relevant here: Hamel basis, where every vector must be represented by a finite linear combination of basis vectors, and Schauder basis, where countably infinite linear combinations are allowed (which must converge as infinite sums, which means this notion doesn’t work on all vector spaces, only those with a topology or a norm).
Hamel basis for infinite-dimensional spaces is a pretty awkward notion (try writing out a Hamel basis of R as a vector space over Q). Schauder bases are often more natural and nicer to work with. For example, the Fourier series of a periodic function is its representation in a countable basis of sines and cosines.
A typical infinite-dimensional Hilbert space is L^2, the space of functions that are measurable and whose square is integrable. (Or rather equivalence classes of functions, because two functions that differ on a set of measure 0 represent the same element of L^2.) It’s not quite the set of all continuous functions—there can be a bunch of discontinuities—but it’s still much “smaller” than the set of all functions, and admits a countable basis in the Schauder sense. (In fact if we take functions on a circle, the Fourier basis works.) And you can also prove that it has no countable basis in the Hamel sense.
Yeah, sorry, I edited my comment to remove that point before I saw your reply. Though I think “Plan A” is bad for other reasons.
at the limit, the threat of civil unrest compels even reluctant capital owners to share the gains of automation
Or make robot armies to keep us in line. This is extreme of course, but there’s a spectrum of such things (AI-powered lobbying, robot police, maybe coopting part of the population, etc). If the masses of people become irrelevant both economically and militarily, they will also become irrelevant politically, or at least the force of history will pull very strongly in that direction.
I think this is the main point that most people thinking about this miss, maybe because it’s just so nasty to imagine. No solution comes to mind at all, except the hard solution of building a worldwide popular movement right now, while the masses of people still have some power.
Selling shovels :-) I wonder if they’ll take the next step and try to commoditize their complement by supporting open models.
I think LW-rationalism has an internal conflict: it includes virtue ethics (“speak the truth even if your voice trembles” etc) but also a strong dose of consequentialism (“optimize as hard as possible”). Sometimes in some people the second part wins. Maybe it depends on personality type, e.g. I’m too lazy to “optimize as hard as possible”, but it’s fun for me to say true things and follow other LW-rationalist virtues. For others it may be the opposite.