Hi, I am a Physicist, an Effective Altruist and AI Safety researcher.
Linda Linsefors
MSE loss does not generate superposition
Challenge: Hand coding weights for efficient sequence memorisation
Do you worry about Torchbearers being a high demand group? ControlAI (the parent organization) recommended it to me after I did an interview with them. I looked at the signup funnel, and it was “read these two documents about our ideals and how we want to run this community, then commit to producing legible AI safety output in a way you will report on regularly. Also, attend an all-hands video call meeting every two weeks.” It kind of gave me a funny feeling, and I was reminded of the history of some of the more notorious high demand EA adjacent groups, so I bounced off it.
I think one reason Torchbearers can be supportive and invest in their members, is because they ask for this type of engagement up front.
I’m not very good at volunteer management, but one thing I learned is that if you make it too easy to join, you get a bunch of (probably well meaning) people who don’t do anything, which is worse than no help at all. These people will take up your time, and derail plans by promise to do things and then don’t do them.
Lot’s of people will just sign up for stuff, without thinking it through.But this also goes both ways. As someone looking to volunteer or join some group, you want the person or org to prove that they respect your time, that you will actually get to do meaningful stuff, and not have your time wasted. I think the best way to do this is to ask potential new volunteers/member to do some task that is intrinsically valuable to them, given that they are the sort of person who would be a good fit.
Probably the sort of person who are a good fit for Torchbearer would like to know about the organisation their joining (this reading will probably not be useful for you if you decide not to join, but it’s useful for your decision of weather you want to join), and also wants some accountability to produce AI safety outputs. If these sounds like burdensome demands to you, then you’re probably not a good fit. Either that or their pipeline is bad. But just from what you wrote, these seems like reasonable requests, if what they want to build is a highly engaged community.I don’t know what “more notorious high demand EA adjacent” you’re talking about? The only thing fitting that description I can think of is Leverage, and that was a full time commitment while living together, so a completely different level. The level of commitment Torchbearers asks for seams similar to joining a non professional but still somewhat ambitious sports team.
4. The EA movement was pretty useless to me, the other community (Torchbearer community) I was in was significantly more supportive, helpful, etc. despite having been in it for a few months and having been in the EA movement for ~8 years. This has basically solidified that I won’t be broadly participating in EA anymore at least relating to AI safety stuff.
This is not surprising to me.
EA as a whole is not a supportive community, partly because of structural reasons and partly because of the priorities of the people who controls the funding.
Structural reasons: EA is too big, to diverse, and has no clear boundary. I think that to have strong ingroup support, it both has to be clear that who is ingroup, and the group has to be homogeneous enough that everyone is pulling in the same direction, so you know that if you support another member, this will result in the type of thing you care about. EA as a whole does not have this.
Some EA subgroups have this. E.g. a local EA meetup can have a clear ingroup (the people who show up to the meetups regularly) and and a shared goal (helping each other be effective). Or some specific EA org, or group that has settled on some specific goal and method, arguably Torchbearer is in this category?
My working definition for “EA” is that a person is an EA if they think they are (I respect peoples right to self identify), and an org is an EA org, if they would fit in at an EAG.
Priority reasons: Even though the all over structure of EA makes it hard to make it a strongly supportive community, it still falling way short of what it could be. I’ve seen people trying to make it better, and making a difference for people around them, but these initiatives don’t get funded.
I’m happy you found a group that supports you.
Thank you for doing AI safety outreach.
How much is this our brain doing lazy reasoning, and how much is this strategically correct reasoning under cultural constraints.
E.g. if I admit that I believe allegations against A, then I the norms of our culture demand that I must stop associating with A. But if A has a lots of good qualities, then the cost of stopping associating with them is high, so I might want to take that into account.
I.e, the debate that is superficially about [are the allegations about A true] is actually about [should we kick out A], and most people know this on some level and act accordingly.
If this is what is going on, then the only way to stopp this “fallacy” is to change the incentive some how. This would include making it common knowledge that after we find out the truth of the allegations, there is a second step of waying the pros and cons of having this person around. But that can get into very taboo territory.
I started lookin at what it doesn’t react to, too.
It reacted to “evolution”, “mating”, “evolved”, “genetics”, “condoms”, “erect penis” but not to “sex” or “sexuality”.
It reacts to “intermittent fasting”, “body fat”, “eating” but not consistently or strongly to “food”.
What would you say is the pattern?
To me it looks like a mix of cities, sex, food, politics and measurements. It doesn’t look random but does look polysemantic.
I looked at “Full”, which showed a lot more activations and variety of activations than “Snippet”, including strong green ones. I don’t know why this is, since I don’t know how neuronpedia works.
I head that mono semanticity is having a comeback, so I recently had a look at a newer (non public) neuron database. It was not mono semantic.
But you are making some good points. Thanks.
Thoughts on: When is it even possible (or likely) for single neurons to encode single concepts, based on model architecture
As part of my mech-interp research, I’m thinking a lot about the question “how would I do this if I was a neural network?”. Specifically, given some toy model architecture, how can I set the weights so that it actives some task.
Todays conclusion is that single ReLU neurons are almost useless. To do identify any non-linear pattern (e.g. XOR), you need at least two ReLUs. We should not expect any feature to be represented by a single ReLU neuron, since anything that can be picked out by a single ReLU, was already linearly accessible in the first place, so why use the ReLU at all. Therefore, anything meaningful in a ReLU-MLP will use multiple neurons.
GELUs are a bit more expressive, i.e. not completely monotonic, but I don’t think they are differens enough from ReLUs is enough to matter.
I have not though about SwiGLUs, and all the other ones MLP variants. I might get back to that, or share your thoughts in the comments.
Residual stream neurons will not be interpretable, for the simple reason that the residual stream has no privileged basis.
Convolutional neurons seems unusually well suited for single concepts, I think, which is why we see interpretable neuons in AlexNet. Although caveat that this is a post-diction, that I haven’t though super hard about.
No.
If you read it that way, then that is a communication failure, but I don’t know what part you are missing.
I think the easiest to point to difference (though not the full difference) is that my way of doing it gives the recipient the information they need to evaluate my (sometimes implicit) suggestion.
Maybe related: I’ve met people who will, what ever I say, try to interpret my words though the lens of [what action is Linda trying to get me to do? / what is the hidden request here?]. I can see that though that lens there is less of a difference. However, that’s a very limiting lens, that will miss most of what I’m attempting to say.
Often when I give advice I do it as an anecdote. This especially applies to:
Giving advice to anyone who isn’t a close friend (i.e. someone I know well).
Giving unsolicited advice.
Instead of suggesting what someone else should do, I tell a story of a similar problem I had and solved, and then let the other person pull out what ever part of the lesson applies to them. I find this works better for a number of reasons.
I don’t know their context enough to know if what worked for me will work for them. E.g. I don’t know if the direction I needed adjustment is the same as the direction they need adjustment. (https://slatestarcodex.com/2014/03/24/should-you-reverse-any-advice-you-hear/)
It avoids the sazen problem. (https://www.lesswrong.com/s/CkphjEuLfGnsYEzan)
This style works great in text forums (anything from FB to LW) where there is low feedback and more space. I.e. you’re more likely to want to give unsolicited advice, since asking first is a long delay. And you’re less likely to bother anyone by going on a long monolog, because they can read it in their own time, or not.
But it also works well in 1-on-1 or very small group conversations. Although if it’s a long story, you should ask first if it’s welcome.
I think I developed this style of advice because I hate when people who don’t know me super well try to tell me what to do, but I also have the normal human urge to offer advice. So I tried to find a way to share my wisdom that does not rout though the thing that I don’t like people doing to me. And then I just got positive feedback from doing it this way, so now I do it more.
Looking back at this short form, I notice that I never wrote “you should”, or even “you” (in previous paragraphs). I’m not telling you what to do, because I don’t know what constraints you have. But I would like if more advice directed at me, were in this form.
Also, there are many spaces in math. The J-space is a topological subspace for example.
You are right.
(But I’m still going with my terminology in my head, since that works better with my other verbal concepts.)
Why would it being a proper (vector) subspace be more interesting?
1) What they’ve found is a new method for finding meaningful linear directions in the residual stream. This is actually great! Specifically they are able to find linear directions that the model are able to verbalize in a single token, and these also seem to double as internal representation for a lot of thinking.
2) What they haven’t found is some separation between the conscious-like and the unconscious-like information channels.
Point 1 is actually really great, and I’m appreciating it more as I get further into the paper.
Bur for some reason they are trying to make 2 the title claim, probably because 2 would tell us something fundamentally new about these models, not “only” give us a better tool. But it dosen’t work. Taken as a linear space, the J-space is just the residual stream, which is not a new discovery.
They find that they can disrupt/redirect some operation by messing with the J-directions, but not others. Ok, that’s interesting, good to know what the tool works on. But that could be the difference between using internal states that are aligned with single tokens or not. Or having other redundancy or not.
If they had found that the J-vectors spanned less than the entire residual stream, that would be very surprising, and surprising=interesting.
(I’m still working though the paper, so I’ll update this comment if I change my mind.)
It’s not a space though. Unless you count the whole residual stream. We already knew the residual stream exists, and that it holds causally meaningful linear representations.
We already knew that there where concepts with the same direction(ish) across context. That’s why steering vectors work.
It’s a new method for identifying meaningful directions in the residual stream, which is great. But I don’t understand in what way you think they found some new space?
Sorry for putting you on the spot for this. A collogue told me in private conversation that everyone is making claims like this about this paper. But I’m not talking to a lot of people, and you happened to have said it officially in your commentary.
I’m noticing that the term “J-space” is actively bad for me trying to interpret these results. I’m currently going with the term “Sparse J-vector encoding” in my head.
I’ve discovered one bed and one pillow hack!
Problem 1: I’ve been having trouble finding a mattress I like. Most of them are too hard. I like soft. We currently have a memory foam mattress topper, that is very soft and nice, but it has too good memory, so there is a significant dipp in the middle of where I sleep, making me sleep all bent. We got a new one, and it only took months for the problem to reappear. I could try more combinations of mattresses and mattress toppers, but that would cost too much time and money, since just seeing what I like in the shop is not reliable.
Solution 1: I put a thin pillow between the mattress and the mattress topper, in the position between my hipp and my shoulder. It works great!
I’m a side sleeper, and my sides are not flat. My hipp and shoulder sticks out. Therefore a flat mattress is not the ideal shape for me. But most other humans are not very flat too, so why are non-flat mattresses not a common thing?
Problem 2: I don’t like the pillow I’ve been sleeping on. It’s too hard, and sometimes my ear hurts a bit when I wake up. But it’s the only one I have that is just the right height, and finding any pillow that is the right height is really hard. If it’s to high or to low, my neck will start hurting after a few days.
Solution: I filled a pillow protector[1] with shredded foam. I don’t know yet if it’s the perfect thickness, but if it isn’t I’ll put some more in or take some out tomorrow, until it’s exactly right. (I also considered cutting open another pillow and take some stuffing out of it, but since I had the ingredients to make a new one, I went with this.)
I have had these problems for a while, and apparently just failed to apply thoughts to try to solve them. To my (partial) defense there where reasons it wasn’t too bad until recently, and I had other things taking up my problem solving capacity.
I’m writing this down in case someone else could benefit from these solutions and hasn’t thought about it them selves yet.
And also, does anyone know why there is not a market for non-flat mattresses? Or is there one, and I just missed it? @Lucius Bushnaq told me about a mattress store where they made individual adjustments to the wooden bords under the mattress, which is in this direction, so that’s something.- ^
Like a pillow case, but thicker, and with a zipper.
- ^
I’m reading your commentary.
What claims is this paper making? In my opinion this paper makes 4 significant claims:
● Scientific claim: There exists a “cognitive space” inside the model, where (some) intermediate variables are stored during a forward pass
● Methodological claim: Logit and J-Lens both work for finding this cognitive space, and J-Lens is better
● Pragmatic claim: J-Lens is a practically useful interpretability technique, e.g. for alignment audits
● Philosophical claim: This cognitive space is analogous to a global workspace
● I think the scientific claim is by far the most interesting, and I am persuaded by it. The paper provides an overwhelming amount of evidence for the existence of this cognitive space—even if I quibbled over many details, there’s enough hard-to-fake evidence that clearly something important is going on.
Why do you think this paper is relevant for the Scientific claim? What did the paper show that we didn’t already know? You argue yourself, in the same document that something like this must exist, and that it has already been observed in many cases. Why give credit to this paper for something you already knew?
(Having a better alternative to Logit lens is useful though, no argument there.)
I agree that calling it “J-space” is bad. From the name I initially assumed it was a subspace, which by the way would have been a much more interesting result than what it seems to be.
(I’m still working though the full paper, and will update this comment if I change my mind.)
Why are the tokens so cheap in China? Who is subsidizing this and why?
What does “gentleman’s handshake data collection method between relay operators” mean?
Yes, those are two examples of situations with less race dynamics than just no regulated open market.
These are also examples of situation which are worse starting points for getting good regulations in the first place. But if we’re not going to try to regulate, then this does not matter?
I know that the text you originally quoted only argues against having a social movement, and not any regulation attempt. But the arguments do apply to any regulation attempt.
Any regulation attempt requires building blob of political will, somewhere.
The one example they give, the Interstate Commerce Commission, is an example of hostile take over of a regulatory institution. This mean that this is a cautionary tail of having a regulatory institution, regales of if it’s pushed for by popular demands or anything else.
If someone want’s to argue that we should not try to regulate AI by any means, then they would have to argue that that would lead to outcomes that are worse than free open competition.
If the bad outcome is just that this would make it harder for future regulatory attempts, then the conclusion is we should try but also be aware of the pitfalls and try to avoid them. The quoted critique has some value in pointing out dangers, but I don’t think it holds up as an argument not to try at all.
I think you read something into my comment that I did not mean. I did not mean to say that less people should be kicked out. I’m making no comment on that either way. I’m just saying that if someone disagrees with the criteria for what is a unforgivable act, and it’s taboo to argue over what should be forgivable, then they may (on the surface) argue over the facts instead, which may look like halo defense.
Specifically, there is a discussion about if person A has done [unforgivable thing]. Person B don’t think [unforgivable thing] should be unforgivable, but just a normal bad thing, that can be forgiven if A has enough other good qualities. Person B thinks that probably a lot of people agree with them, but no-one can admit that they think [unforgivable thing] is not infinitely bad, without large social risk. So instead person B gestures at all the reason we would all like to keep person A around, and suggest we pretend that person A did not do [unforgivable thing].
I’m not saying halo-defense is healthy. It’s not. I’m saying it might be a symptom of a different problem, which means you’d have to solve that problem to get rid of halo-defense.