RSS

egan

Karma: 171

Mea­sur­ing Spu­ri­ous Cor­re­la­tions with Fea­ture Strength

egan11 Aug 2026 18:05 UTC
36 points
1 comment10 min readLW link

Re­ward Laun­der­ing: LLMs Can Gain Un­in­tended Be­hav­iors by De­cid­ing When to Earn Their Rewards

31 Jul 2026 15:48 UTC
77 points
3 comments4 min readLW link

Hint-based CoT faith­ful­ness evals still mostly work on Claude

egan30 Jul 2026 20:44 UTC
24 points
0 comments6 min readLW link

Re­search Sab­o­tage in ML Codebases

30 Apr 2026 0:26 UTC
63 points
4 comments6 min readLW link

How will we do SFT on mod­els with opaque rea­son­ing?

21 Feb 2026 0:00 UTC
32 points
17 comments7 min readLW link

Four Down­sides of Train­ing Poli­cies Online

4 Jan 2026 3:17 UTC
30 points
4 comments3 min readLW link