an lwer who is making an extraordinary effort to get her own things out there. DMs open. Linktree
jelly
I recently came up with a particular way of thinking about my mind: there are many relatively independent subsystems in my mind that
are particular quirks in human cognition that I can take advantage of, and
always watch and react to observations and sometimes my own thoughts; I usually can’t hide from them
For example, here are some of mine, with made-up names:
the “Magic Brain Juice” one, which automates any repetitive enough trigger-action pairs
the “naive RL learner” one, which reinforces rewarding experiences and stuff
the “associative thinking” one, which makes random associations of various concepts in my mind whenever I’m not focusedly thinking about something, and notifies my consciousness when it finds something novel
the “competence-identity” one, which tries to maintain (the appearance of) my competence, and emotionally fuels me whenever it doesn’t match that
the “emotional validation” one, which craves social attention
… and so on and so forth
I expect this framing to generalize to some extent to other people, and I also expect it to matter even more to rationalists, because some rats do a kind of renovation in their mind, and in the process ignore these “subsystems”
Inspirations:
Kaj Sotala’s IFS (which helps me think about “subsystems” generally)
“We are agents who cannot simply act because every action is accompanied by self-modification.”—Magic Brain Juice
… which motivates “some subsystems are always in my mind observing me doing stuff”
I made a Manifold market about this:
Some generalizable piece of advice from HOW TO SUCCEED IN MRBEAST PRODUCTION:
USE CONSULTANTS
Consultants are literally cheat codes. Need to make the world’s largest slice of cake? Start off by calling the person who made the previous world’s largest slice of cake lol. He’s already done countless tests and can save you weeks worth of work. I really want to drill this point home because I’m a massive believer in consultants. Because I’ve spent almost a decade of my life hyper obsessing over youtube, I can show a brand new creator how to go from 100 subscribers to 10,000 in a month. On their own it would take them years to do it. Consults are a gift from god, please take advantage of them. In every single freakin task assigned to you, always always always ask yourself first if you can find a consultant to help you. This is so important that I am demanding you repeat this three times in your head “I will always check for consultants when i’m assigned a task”
Counterexample, if Q = “true”, then it becomes “it’s true to believe that if it’s true to believe P, then P”, which means Prov(Prov(P) → P), which implies Prov(P) because of Lob’s theorem
I found that listing things as a hierarchy, like a taxonomy, immensely dissolved my confusion about what to do in both AI safety and my personal life. I had been picturing the AI safety field as something vague and huge and confusing, but breaking it down into categories and subcategories made it much easier to think about. Not only that made it clear simply because writing things down is helpful, but the hierarchy feature itself chunks/abstracts things to make it easier on my working memory. In hindsight it feels really obvious, but the effect was much larger than I expected.
Also, here’s my AI safety taxonomy. I wrote it only for myself so there may be some confusing parts and some missing here and there
taxonomy
ai safety
technical ai safety
prosaic ai safety
persona selection model
interp
mech interp
activation oracles
j-space
model diffing
scalable oversight
automated ai alignment
agent foundations
learning-theoretic agenda
natural abstractions
synthesizing standalone world-models
embedded agency
decision theory
ai control
evals
meta-philosophy
brain-like agi
corrigibility
non-technical ai safety work
forecasting
buy time
politics …
impact on outside society
recruit more people
persuade more people, i.e. advocacy
politics
international coordination
funding
awareness of cybersecurity
awareness of what other “ai safety” people are doing?
awareness of which ai lab is winning the race?
taking care of the rationality community
useful questions to ask are:
which of these is [some content about ai safety] related to?
which is missing from this list because of [some content about ai safety]?
when i contribute to ai safety, which of these do i target?
Definition 4: A COD Q weakly leads to a COD Q′ iff, for all m1,n1,m2,n2∈M,p∈P:
PQ(mO1()=1∧nO1()=1)PQ(nO2()=1)>PQ(mO2()=1∧nO2()=1)PQ(nO()=1)→PQ′(O(m1,n1,m2,n2)=1)=1
PQ(mO1()=1∧nO1()=1)PQ(nO2()=1)<PQ(mO2()=1∧nO2()=1)PQ(nO()=1)→P
I think there’s a typo here,
appears in and is undefined
I mean reaching out to strangers in general. And also, I expect (75%) the problem to be solved by the time I reach 10 cycles in the “actively reaching out to a stranger” <-> “hearing from them” feedback loop.
I’m very bad at writing and reaching out to people, and this is the single reason stopping me from making an impact in this space. I have been trying hard to figure out why I’m bad at writing and talking and reaching out to people for a very long time (even writing this shortform was difficult for me) and here are some hypothesized reasons why:
I’m not a native English user
I’m perfectionist about writing style, trying to make it so that the writing style is “native” enough according to my intuition
I haven’t talked to English people that much so I don’t really know the conversational style
I haven’t talked much to people in general so I don’t even know which information in particular is important to deliver
I’m generally very anxious about actively talking to people I don’t know
I procrastinate a lot recently
Having written the reasons out I expect them to be a one-time issue, and that someone helping me out to resolve these issues has a chance of being counterfactually impactful without spending much time
Unless I’m reading this wrong somehow, I think you’re excluding people who think something along the lines of “current alignment techniques work great in the current regime but won’t generalize to superintelligence, and the hope instead is to use the best AI that can still be aligned to automate AI alignment”.
I have an intuition that the mud-rock spectrum is a very important concept to pay attention to, and the rationality community leaned too hard on muddiness and muddy rationalist techniques, and that this is underemphasized among rationalists. (for example, metacognitive strategies, a CFAR situation, the fact that you have to quickly replace some of your assumptions as you become a rationalist, the quick development of rationality and foundational beliefs on LessWrong in general, …) I personally feel like some of the framing of the concept in the mud-rock post isn’t quite right (muddiness/rockiness probably isn’t best described as a state of mind), but the conceptual understanding behind it basically is. In a very rough summary, things are “muddier” when they change deeper/more foundational beliefs/assumptions, and things are “rockier” when they harden them instead. I think that paranoia is a good step in the right direction here, and that people should develop more rationality techniques in the general rocky direction. (something something Chesterton’s fence?)
You can take a look at the agent foundations wikitag
it just isn’t clear to me why that text should have any meaning to humans reading it that necessarily relates to what the activation means
To my knowledge, the hope is that the model being trained will improve its own explanations without destroying the association between the explanations and reality, or making its explanations illegible, or using steganography, … because it’s the “simplest” way for the model to improve. It’s the same rationale behind using chain-of-thought to monitor LLM behavior; iirc research does show LLMs keep chain-of-thought legibility under various circumstances, though there are edge cases.
I think that the amount of contributions a person can contribute to discussions like on LessWrong, and cognitive interpersonal activities in general, is not only determined by intelligence, but also how unique their perspective of the world is, or how much thought-patterns they have that others don’t, or how different they think from the others, etc. Audrey Tang joining the AI safety field is an example (it feels like to me she does have some wacky intuitions that could help the field see things in more different ways, aside from being very smart).
Related: The bar is lower than you think
I don’t necessarily agree or disagree with you, but you might be interested in reading Formal Methods are not Slopless
A Manifold market suggests an 8% chance of Hantavirus causing a pandemic in 2026
A better way to frame it is that the example treated the two hydrogen atoms in H-O-H as the same thing, when in fact they are not, in the same way that there are three fruits in a collection with 2 apples and 1 orange, not two, because the two apples aren’t the same thing. You can say that the set of atoms in H-O-H is {the first H, the second H, the O}
I once said that I’m bad at writing, and recently I realized it’s probably because my natural writing style is a strange kind of rambling, and it takes a significant amount of effort for me to polish it to the writing-style-I-want-to-have, and a part of me doesn’t want to publish the rambling writing style I have, because it stands out too much, at least in a LessWrong shortform feed.[1]
Within the past ~3-6 months, I did a bunch of stream-of-consciousness daily, as a huge part of how I think, and without external feedback on my writing style (like a blogger would get), I developed a bunch of internal jargon, and also possibly drifted away from what I mean by things in some respects, etc. Which means I don’t get to practice normal prose much, even if I still technically write a lot.
no, I’m not using the rambling style here, if you’re wondering