RSS

Elena Ericheva

Karma: 22

t.me/​MrsWallbreaker
AI safety channel — where I learned to stop worrying and love the AGI — deep paper breakdowns plus the occasional rant

Questions I keep circling:
— Interpretability: can we really see what a model does inside, or just narrate it?
— Alignment in multi-agent systems: does it hold up when it’s many agents, not one?
— Refusal mechanisms (my own research): how many distinct things is a model’s “no” built on, and does that count transfer between models?
— RLHF & scalable oversight: how do you steer or supervise a system you can’t fully check?
— Evals: what does a benchmark measure once the model knows it’s being tested?
— The ethics /​ decolonial side of AI: who pays for alignment, and whose values count?

Background: alignment & interpretability in multi-agent systems @ JetBrains (interpretability of LLM/​RL agents, long-short memory in agents, capabilities elicitation with RL); evals @ METR (AI agent capabilities elicitation & evaluation, blue teaming /​ safety cases, RE-Bench); BlueDot AI Safety Fundamentals. Before AI safety — 15 years of ML/​DL research across medical imaging, computer vision and NLP, with the usual paper trail: peer-review and peer-reviewed publications, conference talks (including an oral), a Best Paper, some SOTA.

A ge­neal­ogy of AI safety: how di­rec­tions are born, and how they die (2005-2026)

Elena Ericheva10 Jul 2026 0:35 UTC
23 points
0 comments29 min readLW link