RSS

Neel Nanda

Karma: 15,947

Towards sur­fac­ing model al­gorithms with meta-to­kens in the J-Space

20 Jul 2026 19:45 UTC
45 points
0 comments10 min readLW link

Data fil­ter­ing works a lot worse than you would ex­pect

7 Jul 2026 4:41 UTC
53 points
8 comments3 min readLW link

A Re­view of An­thropic’s Global Workspace Paper

Neel Nanda6 Jul 2026 20:59 UTC
133 points
6 comments25 min readLW link

The Case for Model Forensics

26 Jun 2026 15:09 UTC
47 points
0 comments10 min readLW link

LLM-Driven Fea­ture Discovery

22 Jun 2026 22:26 UTC
35 points
1 comment5 min readLW link

How trans­par­ent is Diffu­sionGemma (and why it mat­ters)

20 Jun 2026 20:05 UTC
86 points
2 comments4 min readLW link

Syn­thetic doc­u­ment fine­tun­ing for in­still­ing pos­i­tive traits

16 Jun 2026 0:04 UTC
62 points
1 comment10 min readLW link

Why Do Naive SFT Filters For Safety Prop­er­ties Fail?

14 Jun 2026 19:45 UTC
60 points
7 comments10 min readLW link

SFT Drives Gem­ini’s Safety Properties

13 Jun 2026 15:31 UTC
90 points
4 comments1 min readLW link

Build­ing and eval­u­at­ing model diffing agents

12 Jun 2026 17:14 UTC
62 points
2 comments12 min readLW link

Models May Be­have Worse When Eval Aware

11 Jun 2026 9:28 UTC
90 points
8 comments13 min readLW link

Build­ing Bet­ter Ac­ti­va­tion Oracles

4 Jun 2026 18:34 UTC
67 points
1 comment7 min readLW link
(arxiv.org)

Test your best meth­ods on our hard CoT in­terp tasks

26 Mar 2026 19:24 UTC
59 points
2 comments19 min readLW link

How well do mod­els fol­low their con­sti­tu­tions?

12 Mar 2026 0:07 UTC
100 points
5 comments26 min readLW link

Cen­sored LLMs as a Nat­u­ral Testbed for Se­cret Knowl­edge Elicitation

9 Mar 2026 18:50 UTC
39 points
3 comments5 min readLW link

Cur­rent ac­ti­va­tion or­a­cles are hard to use

3 Mar 2026 19:33 UTC
83 points
4 comments16 min readLW link

How to De­sign En­vi­ron­ments for Un­der­stand­ing Model Motives

2 Mar 2026 7:14 UTC
51 points
0 comments10 min readLW link

Why Did My Model Do That? Model Foren­sics for Di­ag­nos­ing LLM Misbehavior

27 Feb 2026 3:20 UTC
60 points
12 comments25 min readLW link

mod­els have some pretty funny at­trac­tor states

12 Feb 2026 21:14 UTC
277 points
38 comments18 min readLW link

It Is Rea­son­able To Re­search How To Use Model In­ter­nals In Training

Neel Nanda8 Feb 2026 3:44 UTC
122 points
15 comments4 min readLW link