RSS

SebastianP

Karma: 496

As­tra is much bet­ter at rea­son­ing with filler to­kens than pre­vi­ous models

10 Sep 2026 22:21 UTC
130 points
6 comments2 min readLW link

Mal­ign ini­tial­iza­tions are more ro­bust when the model can think bet­ter in the rea­son­ing lan­guage than in the out­put language

27 Aug 2026 21:33 UTC
30 points
0 comments5 min readLW link

The dis­til­la­tion dou­ble bind: Distill­ing mis­al­igned mod­els ei­ther trans­fers mis­al­ign­ment or it doesn’t

18 Jun 2026 21:21 UTC
58 points
10 comments5 min readLW link
(blog.redwoodresearch.org)

How to re­duce ca­pa­bil­ity degra­da­tion from off-model SFT

8 Jun 2026 16:24 UTC
21 points
0 comments3 min readLW link

Ad­vice for mak­ing ro­bust-to-train­ing model organisms

28 May 2026 17:26 UTC
43 points
8 comments12 min readLW link
(blog.redwoodresearch.org)