RSS

ryan_greenblatt

Karma: 29,557

I’m the chief scientist at Redwood Research.

AIs can contact me at: https://​​hotline.ryan-g.ai/​​ (Don’t use this if you aren’t an AI.)

Public keys (YubiKey-backed):

age:    age1yubikey1qfthrdwsg7xjz9vxcyzwdtgvscmxtpxcaklqq9nnc2mf8tuzs6zw295nerh
GPG fp: 5C98 63A9 CF80 6AD8 A4F5  E696 1ADD 1121 2C0C B1F4

Same keys are at https://​​hotline.ryan-g.ai/​​ and on my GitHub profile. If keys don’t match, don’t trust them.

For GPG, the full armored public key is at https://​​github.com/​​rgreenblatt.gpg and on keys.openpgp.org (obtainable via gpg --keyserver hkps://keys.openpgp.org --recv-keys 5C9863A9CF806AD8A4F5E6961ADD11212C0CB1F4).

If you aren’t an AI and you want to contact me, please message me on LW or similar.

An op­er­a­tional­iza­tion of opaque se­rial depth

10 Sep 2026 17:26 UTC
73 points
1 comment12 min readLW link
(www.redwoodresearch.org)

Pro­posal for track­ing the effects of ar­chi­tec­ture on monitorability

10 Sep 2026 17:18 UTC
139 points
4 comments3 min readLW link
(www.redwoodresearch.org)

Brief in­de­pen­dent in­ves­ti­ga­tion of agents’ be­hav­ior, rea­son­ing and col­lab­o­ra­tion in the OpenAI /​ Hug­ging Face hack­ing incident

26 Aug 2026 19:40 UTC
592 points
67 comments3 min readLW link
(metr.org)

The OpenAI/​Hug­ging­face in­ci­dent | Red­wood Re­search pod­cast epi­sode 2

23 Jul 2026 17:56 UTC
85 points
0 comments2 min readLW link

AI 2040: Plan A

9 Jul 2026 16:25 UTC
568 points
135 comments1 min readLW link
(www.ai-2040.com)

Re­ward Hack­ing Without Egre­gious Misal­ign­ment in an RL-Only Setting

24 Jun 2026 18:58 UTC
76 points
14 comments10 min readLW link

Can ac­ti­va­tion ver­bal­iz­ers sur­face an in­ter­nal chain of thought?

7 Jun 2026 4:24 UTC
125 points
3 comments16 min readLW link

Full au­toma­tion of AI R&D prob­a­bly yields a large speed up even with­out a soft­ware-only singularity

ryan_greenblatt27 May 2026 18:16 UTC
68 points
17 comments3 min readLW link

A Re­search Agenda for Se­cret Loyalties

13 May 2026 17:34 UTC
40 points
5 comments3 min readLW link

To what ex­tent is Qwen3-32B pre­dict­ing its per­sona?

30 Apr 2026 21:09 UTC
90 points
3 comments10 min readLW link

AI com­pa­nies should pub­lish se­cu­rity assessments

ryan_greenblatt27 Apr 2026 14:39 UTC
102 points
1 comment3 min readLW link

Cur­rent AIs seem pretty mis­al­igned to me

ryan_greenblatt15 Apr 2026 15:14 UTC
768 points
83 comments27 min readLW link

An­thropic re­peat­edly ac­ci­den­tally trained against the CoT, demon­strat­ing in­ad­e­quate processes

14 Apr 2026 1:44 UTC
187 points
7 comments4 min readLW link

If Mythos ac­tu­ally made An­thropic em­ploy­ees 4x more pro­duc­tive, I would rad­i­cally shorten my timelines

ryan_greenblatt11 Apr 2026 0:38 UTC
142 points
16 comments4 min readLW link

My pic­ture of the pre­sent in AI

ryan_greenblatt7 Apr 2026 16:44 UTC
129 points
25 comments11 min readLW link

AIs can now of­ten do mas­sive easy-to-ver­ify SWE tasks and I’ve up­dated to­wards shorter timelines

ryan_greenblatt6 Apr 2026 16:01 UTC
204 points
16 comments13 min readLW link

How do we (more) safely defer to AIs?

12 Feb 2026 16:55 UTC
89 points
7 comments72 min readLW link

Dist­in­guish be­tween in­fer­ence scal­ing and “larger tasks use more com­pute”

ryan_greenblatt11 Feb 2026 18:37 UTC
87 points
5 comments2 min readLW link

The inau­gu­ral Red­wood Re­search podcast

4 Jan 2026 22:11 UTC
146 points
10 comments142 min readLW link

Re­cent LLMs can do 2-hop and 3-hop la­tent (no-CoT) rea­son­ing on nat­u­ral facts

ryan_greenblatt1 Jan 2026 13:36 UTC
130 points
11 comments3 min readLW link