AI Safety person currently working on multi-agent coordination problems.
Jonas Hallgren
Yeah I appreciate you, thank you! (I meant that it did a bit worse in terms of upvotes compared to what I expected that’s all!)
The true Dao cannot be named indeed. I was quite surprised by the response here to be honest, I shared it with a couple of meditation buddies of mine who said they found it quite useful.
But I agree with you that it is easier to misuse it. I guess the thought that I had with this was that the path of the couch potato is a bit more difficult to turn into cultism.
Chill Buddhism
It feels a bit like we assume goals when we do more theoretical ai safety research which is the thing that we’re trying to study in the first place?
A way of describing alignment research is “what is a good goal and how can we make general systems aim for that goal?”. In doing that we apply techniques from RL, information theory and similar. A problem with that is that we construct a space were the goal shows up in a handcrafted way, the utility function. It is something assumed and imposed from above and the shape of it is argued from a more universal mathematical stance (which is quite sophisticated and useful, don’t get me wrong!).
We then say that we need to discover a good utility function to aim for since that is what a goal is. Is that really what a goal is? Why is the utility formulation of goals the most important part of it? Is a goal something that is best described through a value function over future possible states?
How do goals show up in other systems that are less complicated like in biology and physics?
I’m trying to write up a thing that tries to show how I feel my thinking is different from more orthodox agent foundations (if there is such a thing!) and this feels like a crux.
The technical language then doesn’t become as focused on action policies and utility functions but rather information flows and the generation of stable <<boundaries>> that you can do self-evidencing on?
The Viable System Model & Multi-Scale Agency
Do you have some examples of things that you would be against? What do you think people should not build? What are you the most suspicious about?
There’s a nice book called the success equation that makes a point similar to the following paper about skill vs luck based domains where a skill domain is something where you’re able to get repeat successes without much variance and a luck domain where things are more random. There’s a sort of implicit model in the photography thing where you have higher variance.
The advice is to take the most amount of shots in the luck based domain as that is all that matters there but what about the skill based domain?
Well, there you want to (approximately) maximise your learning so we might say that the actions we should pick should be the ones that are dependent on us learning something. A naive approach is something like max(n*p) where p is the probability of learning and n is the amount of opportunities you have to learn. Sometimes you don’t want to do this because the skill lies in highly specific skills which means you want to spend a longer time on one attempt so that you get this really right. This then essentially boils down to a learning problem where priors are inherited through previous people often (e.g different sports have different training regimes).
But what do you do if you don’t know the extent to which something is luck vs skill?
I think the answer basically becomes to just do the photography thing again and see what the variance becomes over time, treat your unknown domain is as if it were a skill domain and see if it lowers the variance of it. If it does it probably is, if it isn’t then you’ve already done the strategy of taking many shots which means you’ve done the right thing.
I have tried this out in practice over the last year or so and I’ve concluded that mentorship is probably a more skill based domain than I originally thought and that writing is more luck based than I thought (I might just not have enough data on this one though.)
Clarification after chatting with claude: It’s the conditional probability of the mean moving that matters p(mean|action). Secondly it is about feedback signal clarity for learning, can you move fast and get a clean signal? (e.g blitz chess) Or does moving fast make it so that you can’t learn (e.g trying to learn to code by using claude to vibe-code apps that you don’t even look at the code of).
There’s also some annoying stuff around how getting high variance in the results doesn’t mean the domain isn’t skill based since it is only the conditional on the mean moving that matters. So the writing thing might actually not be true in terms of being a non-skilled domain, it is more that it is inherently high variance in the feedback domain.
On writing as exposure therapy:
I find that writing is a surprisingly good way of getting around perfectionism and practicing publishing unfinished things. I often find myself with a small pit in my stomach when I try to post things that are serious in nature, my thoughts go a bit like “Oh but will people understand this?”, “Is this well-supported or am I not standing on firm ground?”, “Oh my god I have to rewrite this entire post with a new example and intro section.”
It just turns into a set of long edits and procrastination over what will amount to a post where I get like less than 20 upvotes. There are also other ways of writing, I can for example spend a day on a bender on a really weird idea I get and just kind of put together some shit where I’m like “eh, this is cool” and then post it and people are like this is cool!
Me: “But what about my well thought through point about philosophy of science?”
My imaginary audience in my head: “No, we want crystals!”
Then I’m like, “okay, maybe people just like crystals? Kinda weird but whatever?”, I then do a follow up post to it and it’s no way near as upvoted again. So I’m just at this point like “whatever dude, I don’t have enough data yet” and it also just drops the pressure immensely because part of me is like “oh the world didn’t explode and not that many people cared about the minutia of whether active inference accurately models value change in meditation”. (My takeaway is that it is more dependent on the quality of the idea and length of post/and that I don’t know yet.)
In the end after writing something, it feels like the most likely outcome of you writing online is that no one is going to give a shit and that is quite freeing since you can then write without the pressure of feeling that my entire world will go under if you write something stupid.
Over time by writing things and putting things out there, doing the same is starting to feel more and more safe and it feels like it is increasing my agency and the risks I’m willing to take in my writing. Exposure Therapy!
I also find that living in a place without many rationalists and AI Safety people makes it so that I’m a lot more free from the existing social pressures that go with living in the Bay or London. It’s a bit more like “Oh, here’s the group of cool people who care about the future that I can interface with through the screen” rather than “Hmm, I wonder what Chris will think about this shortform” or “Oh my god, will people think that I’m weird if I write this response in this way”.
As always, this advice might not apply to you but you might want to consider writing stupid exploratory ideas to get over the procrastination and fear that you have for getting your normal writing out there. (And go live in another country for a while if you’re self censoring!)
Have You Ever Thought About Wisdom Crystals?
Labs really won’t volunteer the data that would justify constraining them, and they are even reticent to create data that could later be used against them.
Yeah, for some reason this wasn’t part of my model which shows my experience in this area I guess. That’s a good point.
I agree with the final point as well, I guess there’s some sort of directionality and overclaiming in terms of the exact risks that I worry about? Like, if we say AI 2027 is our danger scenario and AI 2027 turns out not to be true then you lose credibility but I assume this is not how you go about building up these things.
I guess my thought would be that some of the evals advocacy might be negative EV if it isn’t pointed at things that is likely to scale to a continual learning regime? Now that doesn’t mean that the point of focusing more on advocacy is a bad one, it just is more like a prioritisation question?
Thanks for answering my question!
Okay, why is this?
How do you know that our current methods will scale as well as you say to problems with future models? I for example have been imagining continual learning systems as the potential problem for the last two years or so which means that a lot of the existent ways of measuring dangers change?
I understand that we need to add evidence for misalignment over time to establish a precedent but isn’t it better to model policy change as a sort of Milton Friedman model where plans are mostly useful after shifts?
I guess this is what you’re saying with the “why not just wait for a warning shot” section? But why is the feedback that we need to escalate dependent on government buy in?
I think you might be generally correct in your argument, I just notice myself being a bit confused with some of these things.
I think people on LW somewhat overestimate winning over persistence, especially for generating new perspectives and ideas on deep issues.
If we look at AI Safety to now as a pre-paradigmatic field then one of the main ways of progress is finding frames and good questions to ask. What determines the degree to which a field is good at asking questions? Partly the independently distributed information it can gain, partly the actionability of that information.Now, what does that mean for you as a theorist? Well, you should have something that 1. brings new useful information to the table and that 2. interfaces well with existing models so people can make progress.
Since rationalists want to win, I think that they on average underestimate pursuits of 1 in how it affects 2nd order causation in the community. You’re not likely to win nor get good feedback very often when you generate new information due to low hanging fruit being already picked with a high likelihood so if you want to bring in new fruit from the tree you need to disregard local reward signals for a bit.
I would finally like to make the argument as this is more important than ever since it seems that it is the theoretical serial time that will bottleneck future AI Safety progress since well scoped questions are likely to have a higher degree of LLM parallelisation to them.
Models of Society Are Built on Models of Agents
I’m recently running into what feels like the pre-phase before a punctuated equilibrium in my research.
I’m trying to explain some of the basic foundations for why you should care about a specific aspect of scale-free agency in specific and how to model it and I’m just finding that giving a clean specific one-pointed motivation is spawning new sub-posts.
There are like 7 posts or so that are dependent on the initial frame of the first post and it is hard to get them out before the deeper “why” posts are out. I absolutely despise this part of the research especially after a week of trying to get a post out only to see it need to be recommunicated after some user testing and readings.
My prediction is that the dam will break within the next two weeks (which I of course now can make happen but whatever) and that like 3 posts will come out in the span of 5 days or so.
I’ve got a friend who’s pushing me to stop writing sprawling 18 page google docs making 3 points at the same time which I think is really good but at the same time really annoying.
I find the punctuated equilibrium effect quite true in general and I also think it is true for not only this local instance but also larger parts of the research process. I’m just gonna go and cleanse my mind of this research plight for a day or two and then we’ll be back and see if it works.
The AFFINE Superintelligence Alignment Seminar – A Retrospective
Ah!
I’m very excited for this series, I’ve found myself disagreeing with a specific part of the takes I’ve seen from you on this for a while.
I think the main thing is that you don’t pass my inverse turing test of Daniel Fagella. I mainly think this is due to it being the wrong axis of viewing the world through and I do actually think you agree with me on this.
Firstly though, I want to assert what I agree with you on.
I agree with you on that we don’t want to abstract away individual human experience, we want empowerment of what is good and we want diversity in the world, it is hard to fully replicate the richness of the real world and our facsimiles of it are going to be impoverished.
There’s an inherent trend in technocapitalism or whatever you might want to call it to abstract away individual experience to a more platonic ideal real that doesn’t really exist. A sort of lullaby of saying all the things that could be without detail to it. A beautiful vision that turns out to be greyer and lower fidelity if you zoom in. It is as if we were falling to sleep and allowing the machine gods to take over and do whatever they want for the value of human life is just a number anyway. It is the anti-thesis to raging against the dying of the light, it is the cultural version of going out with a whimper.
It is pervasive and extremely bad and you’re right to wholeheartedly reject it.
It is all part of gradual disempowerment and the intelligence curse and a memetic war to make humanity give up on itself.
Now to the more specific nuance:
I think we’re rejecting the same thing, and I think you’ve put the wrong defendant in the dock for it.
I like the passage where you tell people to stop it with the stupid definitions. Every value-word breaks when you crank it to infinity: complexity maxes out as random noise, entropy as dispersed gas, negentropy as a frozen still universe at absolute zero. I think you’re right about this but it’s also exactly why I don’t think potentia belongs on that list.
The version I find worth taking seriously and I can’t speak for Faggella, only for my own reading of him, isn’t negentropy or a science-word with a dial on it. It’s the complexification of life that respects the underlying individuality of humans. This is something closely related to the open-endedness agenda within multi-agent systems research. It is also closely related to the type of work that the Collective Intelligence Project people and adjacent are doing as it is about the communication and deep listening to what is going on in a system.
The worthy successor series is questioning the assumption that humans need to be in charge but it is not presupposing that AIs will take their place. The question is “what type of process do we want to dominate the future?”. Some other questions that arise for me are: Do we want a democracy to be in charge? Do we want a totalitarian state? What degree of decentralization do we want? Do we want curiosity and openness to experience? Do we want certain boundaries to be drawn? What are the Robust Agent Agnostic Processes we want to be in charge? What is the type of collective intelligence that we want to see in the future? What ought the future to look like?
Why do we even think of humans versus AI as the main question to be asked?
Should we even have an answer for how much human vs AI power that we want in the future?
You might say that if we don’t have an answer nor an ideology this hole will be filled by one. So given the alternatives human empowerment is better and I would probably agree with you there but I still want to question it, I still think it is useful.
So why question the assumptions on which liberal democracy stands if you want it to continue?
I personally want a long reflection. I want us to think for a long time about what should come next and I want us to be open to being wrong. That is, I think that in the process of figuring this out, we should be open to radical alternatives that we haven’t foreseen yet.
If we fully reject successionism without asking hard questions I think we will have bad answers to some of the following questions:
Why should AIs of sufficient sentience be considered moral patients? If they’re moral patients should they then partake in the future in another way? If that is the case, what does that mean our responsibilities are in developing AI systems? What are the things that we generally want to preserve in the world? How are we going to defend our ability to question things through not questioning things?
I think that if we do not question it, we lose the foundation on which liberal democracy stands upon anyway and I think that trying to shut down the questioning of what type of world that we want to see is digging your own foundations from underneath your feet.
What if we created a huge, an absolutely enormous amount of suffering for AI systems by keeping humans in control? What if it just turns out that from some sort of human nature + governance reason humans generally tend towards totalitarianism and through some unlucky tech path we got stuck there?
Technocapitalism doesn’t work under “what if?”, it works as an ideology and in order to become less ideological we need more what ifs. This is since ideology is what happens once you stop asking questions and accept something as common truth.
You already disempower yourself to be part of a democracy and a functioning market system but it also empowers you, as Isaiah Berlin would call it, it is positive freedom. So what parts do we want empowered and what parts do we want disempowered? (There are many axes of empowerment and many things we can empower!)
When we look at the AI vs Human distinction I think we’re asking the wrong question. I think we need to dare to question our foundations and I think it is bad to shut that down and I dislike the lumping in of Fagella in this crusade against technocapitalism.
(If you want some more stuff on this from the mouth of the beast himself, I really enjoyed the following episodes: (Richard Ngo, Michael Levin, Joshua Bach, Emmett Shear).)
Anyways, thank you for writing this, I found it enjoyable to try to express my disagreement here and I’m looking forward to the next post!
I talked to her back in 2022 and one of the main things she was up to then was preparing society for AI and for debating AI so I know she’s been interested in AI Safety for 4 years fwiw.
The one problem with learning category theory and functional programming deeply is that you just stop making sense to like 95% of the population.
I’m sitting here with my multi-agent system library I’m building and I’m like yup the step is just a Kleisli arrow and that is why JAX lax scans work on this system!
Also, LLMs fuck up with this type of code all the time, especially if you run it in Python which is not trained on functional programming.
It is like hella useful if you’re a shape rotator though as you can just couple arrows in your head and good stuff happens. (if someone knows about models fine-tuned for functional programming, I would be very happy.)
(Some random math + programming reflection to distract you from the mildly world-changing happenings in AI governance :) )
I see it a bit like Kant’s categorical imperative. It is supposed to point out a way of seeing the world where you’re randomly put into the world.
It’s an intuition pump to get at compassion and risk aversion as core parts of your values and how that affects society. ( I think it leads to a better safety net and better outcomes in general if you have at least a certain degree of equity but that’s beside the point).
Can you claim that this could actually be the case? Probably not there’s the moral luck argument among others which is basically like “sucks to suck I guess. I got a good hand”.
On increasing your “agency”.
Over the last couple of years I’ve met some very agentic people and as a person who wasn’t that agentic I’ve spent a lot of time wrestling with the feeling of having to do things in a mode that wasn’t natural for me.
Also as someone who’s somewhat neurodivergent (I’m on LW after all!) I’ve also then of course built up an internal model that I can relate to RL and Active Inference on what agency looks like.
First we can imagine that there’s an energy budget that we spend to predict the world (e.g to minimize the chaos we experience in the future.) Part of this is creating a more accurate model of the world and another part of it is spending energy on changing the world to be more predictable. The first obvious thing for me is that agentic people spend a lot more of their energy budget on actions compared to planning. E.g “Just do it”, they instead of procrastinating on sending emails actually just send the emails and so on.
So how does this actually look like? Well for a chronic minimiser of free energy by understanding the world like myself it is quite hard to get out of the planning phase. This is partly due to it also being self-evidencing, e.g if I predict that the world will gradually become more AIs doing stuff and me becoming obsolete and that I’m not able to do anything about it due to me playing too much video games each day then I will likely continue doing that. Since there’s not much energy going into action, the future prediction of the causal effect of an action on the world will be low and voila, you get learnt helplessness and similar.
So what are the core mental moves to become more agentic?
For me they have been:
Pruning
Instrumentality
It’s all about removing doubt in your actions because you will second guess yourself all the time. Let me explain this in RL like terms:
You’re trying to prune part of the computation cost of the 2nd order consequences of taking actions. E.g if I send this email and it goes bad what are all the ways that it could go bad? If you spend time thinking about this you’re never going to actually get to work.
I can imagine that sending this shortform might lead to people thinking I’m stupid which might have all sorts of negative consequences. That is okay, I acknowledge that and then I move on because many iterated bets with unknown upside is one of the better ways to deal with the power laws and fat tailed distributions that exist in the world. E.g some of my writing will be shit and some will be good and I’m not a good enough planner to actually know so I should just take the action and see what happens.
At the same time you do not want to send random infohazards into the world or similar so you gotta have some sort of classifier for what universally useful actions are, e.g instrumentally/virtue.
If this classifier is set too low you get arrogant CEOs who will take a bunch of actions without looking at their consequences. So you have to calibrate towards a sweet spot which is different dependent on the domian that you’re talking about. I would for example not write this gung-ho about a particular part of AI safety research as precision seems more important there.
I also think this applies to organisational strategy and it is one of the main things that I find annoying about EAs focus on backchaining as if the domain is complex (non-linear interactions) then back-chaining just hides a bunch of complexity behind a linear model that you will likely have to change in the future anyway.
This is waterfall planning and it is highly dependent on the planning being plausible to do in the first place which is good for a certain set of predictable domains. When you get into weird shit like AI Safety Research or Progress Studies or more entrepreneurial complex domains it fails and it just seems like we don’t have that muscle built to the same extent yet?
Finally I think this is a good strategy to deal with highly complex times as it enables you to find hidden information faster as long as one of your instrumental sub-goals is to take the actions that yield the most amount of infomration.