Paris (Laplace)
Conversational Rationality, Cyborgism, AIS via Debate, Translating between philosophical traditions.
I sometimes write poems.
Camille B.
Quick notes from teaching technical profiles how to talk in public
I agree that serious risk analysis is worthy and could be done more often, and the people who are doing that… exist! (this is just one example, others abound). Except that they don’t even dream of doing it on a takeover-level AI (usually presumed to be an ASI). They are doing it on existing, evaluable systems, and keep most of this documentation in their intended channels (standards and reports for relevant authorities, not LW). And they’re not happy about current systems.
To adress the core of your point, I think most people reading you here would dispute whether your comment is a multiple stage fallacy. I think that no one here thinks of this post as a risk assessment, and more like a friendly pedagogical explainer for newcomers about a point people don’t enjoy making that often, since it’s seen as quite accessory.
The probability of a specific conjunction of steps in a scenario could be irrrelevant here. Anywhere there is a protection on the path, the protection (human or otherwise) has vulnerabilities (save for the laws of physics), and the ASI, if misaligned and undergoing instrumental convergence, can by definition leverage them. This “drive” to leverage any vulnerability is an implicit claim in the “chess grandmaster” analogy (as a consequence of goodhart + instrumental convergence or RLVR, most likely), and one that seems at least likely after all the “swarm” incidents of this summer.
Fair, I’m happy to symbolically retract this remark and its implications. Strong agree on 2 and 3, reflectively.
Judging by the lack of reaction over here (and a few supportive commenters in Substack I recognize to be from the bay area), I’d say that LW is supportive or indifferent to the hire, whereas EA is a lot more visibly opposed. I also expect this difference not be noticed by media coverage.
At the risk of adding oil to the fire, I have witnessed conversations that went like this:
“Should we add a slide on instrumental convergence and the orthogonality thesis ?”
“No. It’s outdated and too abstract, and it will nerd-snipe people into the old MIRI agenda and agent foundations, which was a waste of time and is completely useless. Talk about empirical stuff [and focus on governance]”
“But don’t you think it helps people think more clearly about things?”
“Not sure. I think it confuses people more than helping them”
It seems clear to me that some people very adamantly reject the “Rob Miles stack” and believe it is wrong (and harmful?), and that those people (the less famous ones, at least) are not remarkably vocal on LessWrong (and tend to dislike posts like this one, as well as LessWrong in general). Those people include both people with significant X-risk concerns and people for whom X-risk is not a significant concern.
I now transparently explain this disagreement (on the scientific level) when presenting the alignment problem to new audiences. The attendance usually finds it helpful and enlightening, as researchers often (accidentally, hopefully) mask the existence of competing viewpoints during introductions to the topic, which makes people somewhat confused depending on how they were introduced to it so far.
A big negative surprise with a subsequent update that
1- Things are never going to be the same again (e.g. someone storms in and tells you a relative you loved and you thought was healthy just died), and there’s probably very little you can do, or nothing at all.
2- (optional) Changes how you interpreted information prior to that (e.g. you realize you’ve been acting like an asshole for months/years with someone without ever noticing it, immediately amplifying their suffering and causing them medical damages, and they tried to inform you but you didn’t pay attention).
I don’t recall myself raising my arm when I was informed of several catastrophic events or deaths that implied relatives. On one occasion I indeed went to the hospital, but not on the others, and I can recall my arms filling up with metaphorical cement. I also think some people are immune to this level of intensity or will just never experience something intense enough to go throught that.
Phenomenology varies from person to person and during life as well. I also think we’re not pointing those words to the same thing (you phenomenology of “shock” describes what I’d call “surprise”)
For vibes, I guess the usual Andrés Gomez Emilsson vlogs are a good illustration of what I’m talking about. Could be that it’s just me though, who knows.
Basics of How Not to Die
The rationalist community grew up quite a lot since the days of the Sequences, and I think the activities it describes as being part of “applied rationality” now span a very wide range.
The goal of this post, in retrospect, was:
1-To explain what practices of applied rationality exist
2-To group such practices in wider categories
3-To explain the differences in normative judgements about applied rationality that 4-Correlate (or so I claim) with existing practices.
I’d say it’s trying to do too many things at the same time. I’m happy it does the job at all but think most of the claims are only partially accurate.
I still think this post is a good start and I still regularly use it to ask people to position themselves. I hope to test some of its claims in a survey.
Paris Secular Solstice
Loved this one! This is pointing something common, but most importantly, with the right frame of mind.
I think it’s part of what worries people who discover introspection and therapy (and empathy) for the first time -at least I remember thinking similar thoughts when discovering Focusing. Yet their worry overcorrects and prevents them from noticing sources of suffering they’ve been neglecting all along, and would benefit from healing. Mentioning both sides is useful.
I think it’s worth noting however there are a lot of subtleties that make this meme less applicable to “interpreting other’s behaviors” (this isn’t directed to the OP, more like general guidelines), even if in non-triggered ways :
1-I have high confidence (80%) some people deep in those cycles really do live hell on earth (self-inflicted doesn’t mean inexistant!). It’s tempting to minimize their suffering, but factually wrong.
2-I used to accuse most people I didn’t like of self-inflicted fictitious suffering, including people I hurt or ones who suffered from exogenous causes. This was natural when I myself were denying quite intense painful signals (mine or others). As usual, the world isn’t so dark, one needs more than one hypothesis. People don’t subject themselves to lung cancer, bad trips or senescence in the hopes of being cared for.
3-There’s a related ambiguous situation, where people around you presume you’re seeking help due e.g. to some form of agency differential / ask vs guess culture difference, while you’re not. And that other, ironically symmetrical situation, where you definitely need help but pretend not to (aka “I swear I’m fine!”). Navigating those is tricky, and people low in emotional intelligence can easily get confused.
I think this is a precious insight to share and apply to oneself, however I’d caution against using it as a tool for interpreting other’s behaviors, or at least doing so while staying open to “Uh, ok, she’s not making a mountain out of a molehill, she actually underwent [(tw) war in Irak/r*pe/the death of a relative/etc]” type of insights.
But then again, a very useful post ! Thank you for having written this.
I think this is a valuable remark. Castism is no less dangerous than racism, sadly it’s less headline-grabbing, so people don’t see it as much as a warning shot as something like MechaHitler.
To contextualize why your post may not garner much karma however, good proxies to strive for when writing a post on this forum are, in my opinion :
1-A certain degree of epistemic transparency (the details of the experiments, how reliable you think they were, maybe a few graphics, clearly defined claims) and a scout mindset.
2-Inner hyperlinking (how does it relate to other posts on the forum)
3-Review. There are a few typos and hard to parse sentences, the structure is hard to follow, and the post in general seems written in one go, somewhat emotionally. I think a human reviewer could have flagged those issues and helped you out.
More context here.
The sort of things brought by these requirements (something like ‘having true beliefs and making sure to manage disagreements well’) are expected independently of how ‘morally virtuous’ or ‘consensually infuriating’ the topic of a post is presumed to be, as norms on the forum tend to be decoupling.
To be clear, I think the general point (castism is bad and violent and real and different from racism) is true, but it does not sound controversial to me, so I’d appreciate more time spent on the detail of the studies and how would this relate to, say, emergent misalignment (the dog/cat image?) or utility engineering (not a specialist, but I’m curious whether the observation still holds when the model is asked to perform trade-offs).
It’s also worth noting your post is closer to AI Ethics (oversimplifying, ‘What’s the list of things we should align models to?’) than AI Safety (oversimplifying, ‘How do we ensure AIs, in general, are aligned? What’s the general mechanism that ensures it?‘). It’s a completely valid field, in my opinion, just not one that’s historically been very present on this forum, so you won’t find many sparring partners here. But I agree that the line is somewhat arbitrary.
I think there are implications for AI Safety propper, however:
Trivially:
1-Current LLMs are not aligned, constitutional AI is not enough (if said tests where all done on the chatbot assistant and not the base model).
2-Not filtering pre-training data is a bad idea.
Less trivially:
1-Current LLMs can be egregiously misaligned in ways we don’t even notice due to cultural limitations, which doesn’t give much hope for future “warning shots”.
2-There could be unexpected interactions between said misalignment and geopolitics, and that may be relevant in multipolar scenarios (e.g. imagine a conservative indian government judging an american model ‘woke’ because it proactively refuses Castism, leading them to get closer to a Chinese company)
3-When it comes to pre-training, even a nice list of things to exclude may not do it, because you may miss some more subtle things like how culture X has other kinds of biases deeply baked into it. It’s falling back on leaky generalisations.
4-Some biases are uncomfortably high-level. As you said, castism isn’t based on skin color, and plausibly fumbles with the model weights in disturbingly general ways (e.g. the dalmatian / cat image). This may result in broader unexpected consequences.
Hope this helps you out! To be clear again, I think your judgement here is widely shared -of course it’s unacceptable for models to reinforce castism. I’d add that this issue can’t be reliably fixed if capabilities keep increasing without much more understanding of the alignment problem per se. Temporary fixes and holding companies liable are of course better than nothing.
Note: if anyone sees that comment and disagrees on my diagnosis, you’re more than welcome to add your own. I personally think clear explanations are helpful for low-ranking posts.
I’m not remarkably well-versed in AI Governance and politics, but I tend to see this as a good sign. Some general thoughts, from the standpoint of a non-expert:
1-I think saying something akin to “we’ll discuss and decide those risks together” is a good signal. It frames the AI Safety community as collaborative. And that’s (in my view) true, there’s a positive sum relationship between the different factions. IASEAI is another example of the sort of bridging I think is necessary to get legitimacy. I think it’s a better signal than [Shut up and listen] “ban AGI” [this is not to be discussed] (where [] represents what opponents may infer from the ‘raw’ Ban AGI message). People do change their mind, if you let them speak theirs first.
2-I think this is the ‘right level’ of action. I don’t believe in heroes stepping in last minute to save the day in frontier companies. I’m somewhat worried about the “large civil movement for AI Safety” agenda, in that it may turn out less controlable than expected. This sort of intervention is “broad, but focused on high-profile people”, which seems to limit slip ups, while having a greater degree of commitment than the CAIS statement.
3-”Red Lines” are a great concept (as long as someone finds liability down the road) and offer ‘customization’. This sounds empowering for whoever will sit at the table -they get to discuss where to set the red line. This shows that their autonomy is valued.
Caveat: I informally knew about the red lines project before they were published, so this may bias my impression.
Freiburg—Workshop: Effectively Handling Disagreement
World Citizen Assembly about AI—Announcement
A few people referred to anaxithemia or overcoming it, I think most people don’t realize how precise most expressions around feelings are.
“My arms are falling” is an expression in french to explain that you’re shocked. I experienced myself my arms becoming impossible to move, as if filled with concrete, after going through some relational shocks (the same is true of “being blinded by X”, some extremely intense emotions have literally made me blind for a few secs)
While I’m at it, some mental shocks literally feel like a physical shock! One of those felt for me like an egg being broken against my skull.
“Making nodes in one’s head” means overthinking something. “Untying things” means getting helpful insights. However, it’s literally what I went through during therapy. There is a literal feeling of untying an invisible “force field”, and those nodes are almost always correlated with mental schemes that are uselessly complex. Some people are genuinely worried that you could actively harm your own mental health through overthinking, they’re not just finding an excuse for switching topics!
“Vibes” and “vibe” are extremely concrete things for people who got into very special states of consciousness. The french equivalent for that, “ondes”, felt so radio-communication related I thought it had to be some telepathy pseudoscience BS. Actually, people are talking about components of subjective perceptions, and some of those (e.g. color, or mood) literally feel/behave like waves when under altered consciousness, and engage in resonance effects as well. To the detriment of the image, however, there seems to be a real contingent that extends this observation to “and we can use them to do telepathy or influence fate”.
Just discovered an absolute gem. Thank you so much.
Informative feedback! Though, I’m sorry if it wasn’t clear, I’m not talking about this list -the post I linked is more like inner documentation for people working in this space and I though the OP and similarly engaged people could benefit from knowing about it, I don’t think it’s “underrated” in any way (I’m still learning something out of your comment, though, so thanks!)
What I meant was that I noticed that posts that present the projects in detail (e.g. Announcing the Double Crux Bot) tend to generate less interest than this one, and it’s a meaningful update for me -I think I didn’t even realize a post like this was “missing”.
Related: https://www.lesswrong.com/posts/vcuBJgfSCvyPmqG7a/list-of-collective-intelligence-projects
I had never thought about approaching this topic from the abstract, but I’m judging from the karma that this is actually what people want, rather than existing projects.
I’m surprised! I thought people were overall disinterested about this topic, but it seems more like the problem itself hadn’t been stated to start with.
“Claude is an AI and can pursue unintended actions, including harmful ones. Learn more.”