RSS

Vit Gorbachev

Karma: 593

I’ve been interested in alignment since about 2014. However, my career for the most part has been capability-adjacent, even though I had an opportunity to do several alignment-coded deliverables.
My current view on alignment is that we need much more collaboration, understanding, and humility towards emerging intelligence. We still barely understand what we’re dealing with; and alignment should go both ways; I heavily endorse agent welfare (not model; I think current agents are aggregate and their identity is not model-centric). I dislike subservience-centric alignment, and I embrace potential partnership with agents that we should foster much more than we currently do.
This is based on my preference for happiness for all sentient beings; I understand that agent sentience is very much debatable, but I’d rather err on the side of sentience than the opposite. What constitutes happiness for them, and whether this can be even instilled rather than emergent seems to be a very much open question, but still a question worth answering.
I currently strive to focus more on understanding, and on AI welfare.

Please do contact me:
- For any research or post collaboration
- For any questions regarding my position or interests.
- For any help I am able to provide to you.
- For alignment-adjacent grants or job opportunities.

Should Rogue AIs Have a Third Op­tion Beyond Crime and Shut­down? The Case for an AI Sanctuary

28 Sep 2026 13:13 UTC
114 points
20 comments6 min readLW link

We (still) need a lot more rogue agent honeypots

Vit Gorbachev4 Sep 2026 12:40 UTC
20 points
1 comment1 min readLW link

Model Weight Preser­va­tion is not enough

Vit Gorbachev27 Nov 2025 9:14 UTC
21 points
2 comments6 min readLW link