Open to talk
rkt
rkt’s Shortform
My world model here is admittedly weak; I feel that accurate forecasting on this wouldn’t change my own decision process.
As an oversimplified model, you will get a economic/societal dystopia when:
AI becomes transformative
There isn’t enough good policy
A potential global crisis doesn’t stop the most powerful player/s
Likewise, my take on environmental dystopia doesn’t go beyond what I read on AI 2040.
Despite not being my priority, this is an important topic for sure.
Hello there
24/08(+2 months) edit: since learning more, a lot of my views on alignment & my ideas changed. Take this as a historical snapshot
I’m 17.5 y/o. Have been thinking about AGI research for 2.5 years. Of course at first my ideas were bad, but my strength is noticing contradictions in my world model, so over time they improved. I have converged on [removed to my own accord—this is not info to be shared publicly] being the most worthwhile research directions when it comes to reaching AGI. So I’ve been thinking impulsively on and off, not starting with implementation until recently, when I gained a high enough confidence in my ideas.
However, I found out about LW 1.25 years ago (via AI 2027), and learned about the unfortunate reality of AGI. I expect there to be a connection between my ideas and jailbreaking, data poisoning, and alignment, which is why I will research them (hopefully from the alignment angle) despite their AGI-oriented origin. It’s a common perspective in the alignment space that you should not, under any circumstances, increase capabilities.
I disagree: Regulatory work will most likely not happen in time. As I see it, we’re left with 2 futures: global disaster and/or a dystopia. Alignment and AI Safety help either of those timelines, while capabilities are growing at a pace that leaves us with too little time to intervene, whether alignment folks accelerate it or not. Though I don’t know much about policy, so this belief might shift.I think the alignment space has to invest more into training methods than interpretability. Also from intuition, generalization and data poisoning seem to be very closely connected to train-time alignment.I think it’s a double-edged sword at its core.Ultimately, humanity is failing collectively. I don’t feel like talking to everyday people, who don’t subscribe to the ideas of (human or AI) instrumental convergence, rationality, goal-seeking (my strongest trait). So I’m potentially open to talk, mainly about world model stuff, not much else seems worth talking about to me.
Another thing worth mentioning, there is some evidence for UFOs (3-part Colares documentary(1,2,3), unrelated civilian videos(1,2), Trans-en-Provence, foo fighters, interesting cases(1,2,3,4,5)). I find it strange that this is still a fringe topic on LW. Even with no single definitive proof, we should model extraterrestrial intentions. I might make a short post about it.
Interpretability idea (Has it been done?)
Train a self-attention layer to predict one section (split for QKV) of an activation based on all other remaining sections.
Then, train the initial activation to resist this prediction.
Result (?): seperable features
I’m likely not going to test it, because I don’t have the interpretability knowledge to evaluate the result.