RSS

Steven Byrnes

Karma: 30,225

I’m an AGI safety /​ AI alignment researcher in Boston with a particular focus on brain algorithms. Research Fellow at Astera. I’m also at: Substack, X/​Twitter, Bluesky, RSS, email, and more at this link. See https://​​sjbyrnes.com/​​agi.html for a summary of my research and sorted list of writing. Physicist by training. Leave me anonymous feedback here.

RL & search is a ter­rify­ing way to build AGI (an FAQ)

Steven Byrnes27 Jul 2026 14:50 UTC
117 points
13 comments14 min readLW link

LLMs are (still) mostly pow­ered by imi­ta­tive learn­ing, not RL

Steven Byrnes24 Jul 2026 14:26 UTC
167 points
28 comments9 min readLW link

Will al­most all fu­ture com­pa­nies even­tu­ally be founded and run by au­tonomous AIs?

Steven Byrnes22 Jul 2026 20:04 UTC
80 points
2 comments8 min readLW link

What do I mean by “Ar­tifi­cial Gen­eral In­tel­li­gence”?

Steven Byrnes20 Jul 2026 21:13 UTC
42 points
5 comments4 min readLW link

Notes on tech­ni­cal al­ign­ment via hu­man-like so­cial drives

Steven Byrnes8 Jul 2026 18:30 UTC
70 points
10 comments42 min readLW link

Sym­pa­thy for both sides of the egre­gious mis­al­ign­ment debate

Steven Byrnes12 Jun 2026 16:26 UTC
212 points
27 comments4 min readLW link

Em­pow­er­ment, cor­rigi­bil­ity, etc. are sim­ple ab­strac­tions (of a messed-up on­tol­ogy)

Steven Byrnes11 May 2026 17:48 UTC
190 points
74 comments16 min readLW link

Some takes on UV & cancer

Steven Byrnes10 Apr 2026 0:31 UTC
51 points
30 comments6 min readLW link

“Act-based ap­proval-di­rected agents”, for IDA skeptics

Steven Byrnes18 Mar 2026 18:47 UTC
72 points
8 comments5 min readLW link

You can’t imi­ta­tion-learn how to con­tinual-learn

Steven Byrnes16 Mar 2026 21:20 UTC
209 points
54 comments6 min readLW link

Pod­cast: Jeremy Howard is bear­ish on LLMs

Steven Byrnes6 Mar 2026 21:39 UTC
85 points
24 comments5 min readLW link
(www.youtube.com)

Why we should ex­pect ruth­less so­ciopath ASI

Steven Byrnes18 Feb 2026 17:28 UTC
161 points
66 comments8 min readLW link

The brain is a ma­chine that runs an algorithm

Steven Byrnes17 Feb 2026 19:36 UTC
116 points
18 comments4 min readLW link

In (highly con­tin­gent!) defense of in­ter­pretabil­ity-in-the-loop ML training

Steven Byrnes6 Feb 2026 16:32 UTC
85 points
11 comments3 min readLW link

The na­ture of LLM al­gorith­mic progress (v2)

Steven Byrnes5 Feb 2026 19:17 UTC
125 points
28 comments13 min readLW link

Are there les­sons from high-re­li­a­bil­ity en­g­ineer­ing for AGI safety?

Steven Byrnes2 Feb 2026 15:26 UTC
162 points
16 comments8 min readLW link

New ver­sion of “In­tro to Brain-Like-AGI Safety”

Steven Byrnes23 Jan 2026 16:21 UTC
63 points
1 comment19 min readLW link

My AGI safety re­search—2025 re­view, ’26 plans

Steven Byrnes11 Dec 2025 17:05 UTC
137 points
4 comments12 min readLW link

Re­ward Func­tion De­sign: a starter pack

Steven Byrnes8 Dec 2025 19:15 UTC
82 points
14 comments3 min readLW link

We need a field of Re­ward Func­tion Design

Steven Byrnes8 Dec 2025 19:15 UTC
118 points
12 comments5 min readLW link