RSS

yix

Karma: 389

https://​​yixiong.dev/​​

In­ter­nal State Con­trol is a Gen­eral Prop­erty of LLMs

30 Jul 2026 18:24 UTC
36 points
0 comments5 min readLW link
(secondlookresearch.com)

Where does hint-fol­low­ing and con­ceal­ment arise? A case study on OLMo-3 checkpoints

24 Jul 2026 19:14 UTC
24 points
0 comments4 min readLW link

Ge­or­gia Tech AI Safety Ini­ti­a­tive Ret­ro­spec­tive 2025-2026

24 Jul 2026 11:55 UTC
60 points
3 comments7 min readLW link

LLM CoTs re­main mon­i­torable when be­ing un­faith­ful re­quires computation

15 Jul 2026 21:14 UTC
46 points
3 comments7 min readLW link
(secondlookresearch.com)

Teach­ing Models to Dream of Bet­ter Mon­i­tors through Eval­u­a­tion Con­di­tioned Training

19 Mar 2026 21:01 UTC
53 points
2 comments10 min readLW link

We need a bet­ter way to eval­u­ate emer­gent misalignment

11 Jan 2026 16:21 UTC
86 points
9 comments6 min readLW link

yix’s Shortform

yix6 Dec 2025 2:27 UTC
2 points
2 comments1 min readLW link

TastyBench: Toward Mea­sur­ing Re­search Taste in LLM

2 Dec 2025 23:26 UTC
37 points
2 comments6 min readLW link

Les­sons from a year of uni­ver­sity AI safety field building

6 Jun 2025 14:35 UTC
42 points
3 comments7 min readLW link