t.me/MrsWallbreaker
AI safety channel — where I learned to stop worrying and love the AGI — deep paper breakdowns plus the occasional rant
Questions I keep circling:
— Interpretability: can we really see what a model does inside, or just narrate it?
— Alignment in multi-agent systems: does it hold up when it’s many agents, not one?
— Refusal mechanisms (my own research): how many distinct things is a model’s “no” built on, and does that count transfer between models?
— RLHF & scalable oversight: how do you steer or supervise a system you can’t fully check?
— Evals: what does a benchmark measure once the model knows it’s being tested?
— The ethics / decolonial side of AI: who pays for alignment, and whose values count?
Background: alignment & interpretability in multi-agent systems @ JetBrains (interpretability of LLM/RL agents, long-short memory in agents, capabilities elicitation with RL); evals @ METR (AI agent capabilities elicitation & evaluation, blue teaming / safety cases, RE-Bench); BlueDot AI Safety Fundamentals. Before AI safety — 15 years of ML/DL research across medical imaging, computer vision and NLP, with the usual paper trail: peer-review and peer-reviewed publications, conference talks (including an oral), a Best Paper, some SOTA.