RSS

arav-dhoot

Karma: 94

Where does hint-fol­low­ing and con­ceal­ment arise? A case study on OLMo-3 checkpoints

24 Jul 2026 19:14 UTC
24 points
0 comments4 min readLW link

LLM CoTs re­main mon­i­torable when be­ing un­faith­ful re­quires computation

15 Jul 2026 21:14 UTC
46 points
3 comments7 min readLW link
(secondlookresearch.com)

Failing to Rage­bait the New Gemma

11 Jun 2026 17:50 UTC
31 points
0 comments3 min readLW link

Two More Meth­ods for Con­sis­tency Train­ing and Some New Ways to Ap­ply It

5 Jun 2026 21:06 UTC
25 points
0 comments7 min readLW link