RSS

Neel Nanda

Karma: 16,637

As­tra can do a con­cern­ing amount with no chain of thought

Neel Nanda10 Sep 2026 2:31 UTC
210 points
20 comments8 min readLW link

Does Diffu­sionGemma do la­tent rea­son­ing?

16 Aug 2026 4:22 UTC
40 points
0 comments9 min readLW link

Why do mod­els task game?

6 Aug 2026 22:16 UTC
94 points
1 comment24 min readLW link

R-lens: Mak­ing J-lens More Faith­ful on Early Layers

5 Aug 2026 20:02 UTC
83 points
5 comments7 min readLW link

The AGI Safety and Align­ment team at Google Deep­Mind is Hiring (July 2026)

31 Jul 2026 15:53 UTC
71 points
4 comments6 min readLW link
(gdmalignment.substack.com)

Towards sur­fac­ing model al­gorithms with meta-to­kens in the J-Space

20 Jul 2026 19:45 UTC
47 points
1 comment10 min readLW link

Data fil­ter­ing works a lot worse than you would ex­pect

7 Jul 2026 4:41 UTC
59 points
9 comments3 min readLW link

A Re­view of An­thropic’s Global Workspace Paper

Neel Nanda6 Jul 2026 20:59 UTC
135 points
6 comments25 min readLW link

The Case for Model Forensics

26 Jun 2026 15:09 UTC
49 points
0 comments10 min readLW link

LLM-Driven Fea­ture Discovery

22 Jun 2026 22:26 UTC
36 points
1 comment5 min readLW link

How trans­par­ent is Diffu­sionGemma (and why it mat­ters)

20 Jun 2026 20:05 UTC
87 points
2 comments4 min readLW link

Syn­thetic doc­u­ment fine­tun­ing for in­still­ing pos­i­tive traits

16 Jun 2026 0:04 UTC
63 points
1 comment10 min readLW link

Why Do Naive SFT Filters For Safety Prop­er­ties Fail?

14 Jun 2026 19:45 UTC
65 points
7 comments10 min readLW link

SFT Drives Gem­ini’s Safety Properties

13 Jun 2026 15:31 UTC
96 points
4 comments1 min readLW link

Build­ing and eval­u­at­ing model diffing agents

12 Jun 2026 17:14 UTC
62 points
2 comments12 min readLW link

Models May Be­have Worse When Eval Aware

11 Jun 2026 9:28 UTC
91 points
8 comments13 min readLW link

Build­ing Bet­ter Ac­ti­va­tion Oracles

4 Jun 2026 18:34 UTC
68 points
1 comment7 min readLW link
(arxiv.org)

Test your best meth­ods on our hard CoT in­terp tasks

26 Mar 2026 19:24 UTC
59 points
2 comments19 min readLW link

How well do mod­els fol­low their con­sti­tu­tions?

12 Mar 2026 0:07 UTC
107 points
5 comments26 min readLW link

Cen­sored LLMs as a Nat­u­ral Testbed for Se­cret Knowl­edge Elicitation

9 Mar 2026 18:50 UTC
41 points
3 comments5 min readLW link