RSS

Josh Engels

Karma: 1,115

Research scientist on the DeepMind interp team. Thoughts are my own and do not reflect views of my employer.

Where Do LLM Values Come From?

9 Jul 2026 19:54 UTC
17 points
1 comment18 min readLW link

Data fil­ter­ing works a lot worse than you would ex­pect

7 Jul 2026 4:41 UTC
59 points
8 comments3 min readLW link

LLM-Driven Fea­ture Discovery

22 Jun 2026 22:26 UTC
36 points
1 comment5 min readLW link

How trans­par­ent is Diffu­sionGemma (and why it mat­ters)

20 Jun 2026 20:05 UTC
87 points
2 comments4 min readLW link

Why Do Naive SFT Filters For Safety Prop­er­ties Fail?

14 Jun 2026 19:45 UTC
60 points
7 comments10 min readLW link

SFT Drives Gem­ini’s Safety Properties

13 Jun 2026 15:31 UTC
90 points
4 comments1 min readLW link

Build­ing and eval­u­at­ing model diffing agents

12 Jun 2026 17:14 UTC
62 points
2 comments12 min readLW link

[pa­per] Train­ing on Doc­u­ments About Mon­i­tor­ing Leads to CoT Obfuscation

27 May 2026 9:39 UTC
32 points
1 comment4 min readLW link
(arxiv.org)

Test your best meth­ods on our hard CoT in­terp tasks

26 Mar 2026 19:24 UTC
59 points
2 comments19 min readLW link

Train­ing on Doc­u­ments About Mon­i­tor­ing Leads To CoT Obfuscation

18 Mar 2026 20:37 UTC
65 points
5 comments16 min readLW link

Thought Edit­ing: Steer­ing Models by Edit­ing Their Chain of Thought

3 Feb 2026 9:51 UTC
23 points
0 comments5 min readLW link

Brief Ex­plo­ra­tions in LLM Value Rankings

12 Jan 2026 18:16 UTC
39 points
1 comment11 min readLW link