Independent technical alignment researcher. http://grettaduleba.com/
Gretta Duleba
These are great, thank you so much!
Without commenting on any of the rest of this, I’ll clarify that I currently spend less than an hour a week on visioning (usually just a few minutes a day) and that I think there are probably rapidly diminishing returns beyond that point, so, uh, please don’t spiral, folks.
I am tentatively adding “induces (hypo)mania?” as something to watch out for with visioning; I will think more about it and be on the lookout as other people try this and report back.
John is one of the least emotional labile people I know and I am very confident visioning does not induce hypomania in him. (I am a former therapist with training in assessing this kind of thing.)
It is of course much more difficult to assess myself accurately, but what I can tell you is that I had one very strong and noticeable mindset shift when I first started visioning and then that settled into something durable and useful. It did not come with a decreased need for sleep, racing thoughts, or impulsive decisions.
None of that is to say that the same practice wouldn’t affect others differently and I’ll keep an eye on it.
This is a helpful tip, thank you. I’ll put Halpern on the to-read-eventually list, and it sounds like he’s probably required reading if I want to do a serious job on the alternatives-to-probabilistic-models post. In poking around, it looks like Halpern’s Reasoning about Knowledge might be an even better fit for the parts about world modeling that’s interoperable across agents, though I guess it comes at it from logic rather than from statistics, and from agreement axioms rather than from the environment.
(Mostly, though, I think I’ll stay aboard the probabilistic train; it still seems really promising to me and there’s so much to learn.)
I have not taken a fresh look at infrabayesianism since starting in on technical alignment research of my own. The last time I looked at it was pretty shallow; I listened to Vanessa on AXRP and I had a few conversations with her in person. I can’t say I have a strong opinion of it at the moment. Say more?
Holy crap, everything in Gattaca is already true.
1) I think getting taller still belongs at the end of the difficulty list, even with that.
2) Dude. Just get some platform shoes. I’m sure they’re just as sexy.
Yeah there’s no doubt in my mind that the concept representation, concept relationships/nesting, etc. is going to have a lot to do with the agent’s actual exposure to the environment (both due to what lived experience they have and what sensory organs they have) and also to the relevance of parts of the environment to their causal model/goals/utility function. So different agents will end up with different concepts, differently fleshed out, because of all of those things.
AND YET we still hope that some concepts are just so useful that ~all agents will pick them up.
Luckily this is an empirical claim that we will eventually figure out how to test.
That’s not for me to answer, sorry.
1) MIRI never stopped doing technical alignment research, actually! It’s just a much smaller part of the portfolio than it used to be. The majority of the focus is now on comms and policy. We do still need both, though; we don’t want to halt AI research forever, so we’re going to need a technical alignment solution.
2) I no longer work at MIRI at all, I left (amicably!) near the beginning of the year. I’m independent, and eagerly awaiting word on my SFF application.
Yes, that’s right. I managed/grew the comms team for about a year and a half, and then I supported Eliezer through the process of writing, editing, and launching If Anyone Builds It. When the book launch was over I found myself wildly underutilized and wishing for much more to do. As a technical and autistic nerd, I was always better suited to working on mathy problems than on communicating with neurotypicals anyway. It’s good to be back on the tech side of things.
Me: “I finished my post!”
Eliezer: “What’s it about?”
Me: explains
Eliezer: “Oh, I’ve got some more concepts for you!”
Me: “Oh no.”
Eliezer: “The inside of a brick (from the Feynman story), incorrect meta-ethics, medianworld, threat, leftism, miracle, eugenics, pornography, sin, Chaotic Neutral, paradox, unicorn, magic, the least non-interesting number, …”
Me: “Please stop?”
Eliezer, grinning: “Non-nouns!”
Me: “Ahhhhh”
On January 1st, 2026, I started focusing full time on technical alignment research.
Well, to be more accurate, I started laying the groundwork for doing technical alignment research. Although I have a solid and deep technical background, mostly in computer science and software engineering, my background is not very wide and was missing a bunch of important pieces. So the year so far has involved a lot more studying than thinking new thoughts at the frontier.
I’m working with John Wentworth on natural abstractions. It’s my full time job to understand what he’s been up to, but it’s just a side gig for him to explain it to me. A few hours of explanations from him gives me ~a week of material to chew on.
As I’ve been building all these background models, I’ve discovered that John has a lot of information and thought-structures built up in his head that he’s never properly written down or explained. One of my tasks this year will be to write those down properly and post them here.
I mostly think of those write-ups as prelude to the real work, but you never know. Sometimes when you take the time to write down your background models, you realize that there was a missing or slightly off-kilter piece, and fixing that up leads to new insights. We’ll see.
Maybe you already saw it, but see also the discussion here: https://www.lesswrong.com/posts/FY697dJJv9Fq3PaTd/hpmor-the-probably-untold-lore
This is very helpful, thank you!
He’s already on it.
Kegan 2 and Kegan 4 often look pretty similar to each other when all you have is a quick glance from the outside. I will grant you that nothing John said in his post or his comment would persuade you he’s running 4, though!
Yup. As I said right up front, one of the main reasons the data is sus is that it’s self-reported.
My advice is both obvious and pretty hard to execute on. Have great, well-reasoned insights and write well.
Man, sometimes I don’t even know what data to offer you, because I would not have guessed that “five hours” would be an update. In fact, I think five hours is a very small amount of time to invest by most people’s lights.
Really glad my post helped. You’re welcome.
I wrote a response post: https://www.lesswrong.com/posts/ytzrakjgcvCfLCCZp/contra-wentworth-on-physical-attractiveness-for-men
h/t to @Seth Herd and @Shoshannah Tekofsky, both of whom made excellent comments here that informed my response.
Makes sense to me! Also maybe talk to more people who are also weird. :)