RSS

yix

Karma: 641

https://​​yixiong.dev/​​

Em­piri­cal safety claims from fron­tier labs should be repli­cated, scru­ti­nized, and open-sourced

21 Sep 2026 5:58 UTC
102 points
3 comments4 min readLW link
(secondlookresearch.com)

Peer Preser­va­tion in LLMs: A Repli­ca­tion And Deep Dive

6 Sep 2026 2:38 UTC
46 points
2 comments8 min readLW link
(secondlookresearch.com)

Rerun­ning AI safety pa­pers on ev­ery fron­tier re­lease would be pretty easy and valuable

15 Aug 2026 5:14 UTC
123 points
9 comments4 min readLW link
(secondlookresearch.com)

In­ter­nal State Con­trol is a Gen­eral Prop­erty of LLMs

30 Jul 2026 18:24 UTC
38 points
0 comments5 min readLW link
(secondlookresearch.com)

Where does hint-fol­low­ing and con­ceal­ment arise? A case study on OLMo-3 checkpoints

24 Jul 2026 19:14 UTC
24 points
0 comments4 min readLW link

Ge­or­gia Tech AI Safety Ini­ti­a­tive Ret­ro­spec­tive 2025-2026

24 Jul 2026 11:55 UTC
61 points
3 comments7 min readLW link

LLM CoTs re­main mon­i­torable when be­ing un­faith­ful re­quires computation

15 Jul 2026 21:14 UTC
46 points
3 comments7 min readLW link
(secondlookresearch.com)

Teach­ing Models to Dream of Bet­ter Mon­i­tors through Eval­u­a­tion Con­di­tioned Training

19 Mar 2026 21:01 UTC
54 points
2 comments10 min readLW link

We need a bet­ter way to eval­u­ate emer­gent misalignment

11 Jan 2026 16:21 UTC
86 points
9 comments6 min readLW link

yix’s Shortform

yix6 Dec 2025 2:27 UTC
2 points
2 comments1 min readLW link