Open to talk
rkt
Interpretability idea (Has it been done?)
Train a self-attention layer to predict one section (split for QKV) of an activation based on all other remaining sections.
Then, train the initial activation to resist this prediction.
Result (?): seperable features
I’m likely not going to test it, because I don’t have the interpretability knowledge to evaluate the result.
My world model here is admittedly weak; I feel that accurate forecasting on this wouldn’t change my own decision process.
As an oversimplified model, you will get a economic/societal dystopia when:
AI becomes transformative
There isn’t enough good policy
A potential global crisis doesn’t stop the most powerful player/s
Likewise, my take on environmental dystopia doesn’t go beyond what I read on AI 2040.
Despite not being my priority, this is an important topic for sure.
Hello there
I’m 17.5 y/o. Have been thinking about AGI research for 2.5 years. Of course at first my ideas were bad, but my strength is noticing contradictions in my world model, so over time they improved. I have converged on [removed to my own accord—this is not info to be shared publicly] being the most worthwhile research directions when it comes to reaching AGI. So I’ve been thinking impulsively on and off, not starting with implementation until recently, when I gained a high enough confidence in my ideas. edit: while my ideas changed, the vision behind them remains the same
However, I found out about LW 1.25 years ago (via AI 2027), and learned about the unfortunate reality of AGI. I expect there to be a connection between my ideas
, jailbreaking, data poisoning, and alignment, which is why I will research them (hopefully from the alignment angle) despite their AGI-oriented origin. It’s a common perspective in the alignment space that you should not, under any circumstances, increase capabilities. I disagree: Regulatory work will most likely not happen in time. As I see it, we’re left with 2 futures: global disaster and/or a dystopia. Alignment and AI Safety help either of those timelines, while capabilities are growing at a pace that leaves us with too little time to intervene, whether alignment folks accelerate it or not. Though I don’t know much about policy, so this belief might shift.I think the alignment space has to invest more into training methods than interpretability. Also from intuition, generalization and data poisoning seem to be very closely connected to train-time alignment. I think it’s a double-edged sword at its core.edit: While I believe training alignment methods are neccessary, they shouldn’t be shared publically so as not to accelerate ASI even faster
Ultimately, humanity is failing collectively. I don’t feel like talking to everyday people, who don’t subscribe to the ideas of (human or AI) instrumental convergence, rationality, goal-seeking (my strongest trait). So I’m potentially open to talk, mainly about world model stuff, not much else seems worth talking about to me.
Another thing worth mentioning, there is some evidence for UFOs (3-part Colares documentary(1,2,3), unrelated civilian videos(1,2), Trans-en-Provence, foo fighters, interesting cases(1,2,3,4,5)). I find it strange that this is still a fringe topic on LW. Even with no single definitive proof, we should model extraterrestrial intentions. I might make a short post about it.
Unluckily, there aren’t any good surveys on this. But I think that the distribution of rationality on LW isn’t that skewed. People with the will and capacity will figure out these concepts with time. But there is a variety of people here. The thing that varies is what they personally do or consider important. Overall, a quite good level of rationality is mentained. Some controversial topics being avoided helps.