RSS

Marc Carauleanu

Karma: 740

AI Safety Researcher

Currently researching a neglected prior for cooperation and honesty inspired by the cognitive neuroscience of altruism called self-other overlap in state-of-the-art ML models.

Previous SRF @ SERI 21′, MLSS & Student Researcher @ CAIS 22′ and LTFF grantee.

In­duc­ing self-other over­lap with SFT re­duces de­cep­tion at scale, but gen­er­al­iza­tion re­mains uneven

Marc Carauleanu8 Aug 2026 15:21 UTC
24 points
1 comment7 min readLW link

Mis­tral Large 2 (123B) seems to ex­hibit al­ign­ment faking

27 Mar 2025 15:39 UTC
82 points
4 comments13 min readLW link

Re­duc­ing LLM de­cep­tion at scale with self-other over­lap fine-tuning

13 Mar 2025 19:09 UTC
162 points
49 comments6 min readLW link