ArchiveSequencesAbout
QuestionsEventsShortformAlignment ForumAF Comments
HomeFeaturedAllTagsRecent Comments
RSS
NewHotActiveOld
Page 1

Four LLM loss func­tions → four fla­vors of LLM misalignment

Steven Byrnes10 Aug 2026 16:16 UTC
275 points
25 comments6 min readLW link

FAQ: Isn’t AGI com­ing too soon for re­pro­ge­net­ics to help?

TsviBT8 Aug 2026 8:54 UTC
147 points
68 comments14 min readLW link

The Long (Self-)Correction

Wei Dai24 Jul 2026 21:01 UTC
241 points
62 comments2 min readLW link

LLMs are (still) mostly pow­ered by imi­ta­tive learn­ing, not RL

Steven Byrnes24 Jul 2026 14:26 UTC
227 points
39 comments9 min readLW link

AI 2040: Plan A

Daniel Kokotajlo, elifland, Thomas Larsen, romeo, bhalstead and ryan_greenblatt
9 Jul 2026 16:25 UTC
551 points
130 comments1 min readLW link
(www.ai-2040.com)

Sur­pris­ing facts about the slave trade

Joseph Miller26 Jun 2026 1:34 UTC
462 points
51 comments8 min readLW link

The worth­less­ness of vi­tamin D is mildly exaggerated

dynomight23 Jun 2026 19:39 UTC
190 points
28 comments20 min readLW link
(dynomight.net)

A Mechanis­tic Ex­pla­na­tion of Prompt In­jec­tion (and why you should study roles)

charlescye and Jasmine C.
22 Jun 2026 14:09 UTC
398 points
60 comments16 min readLW link

The LLM shog­goth meme is weirder than you think

HedonicEscalator19 Jun 2026 23:35 UTC
259 points
48 comments7 min readLW link
(hedonicescalator.substack.com)

Guardian An­gels: LLM Per­son­al­iza­tion for Pro­duc­tivity and Security

gwern17 Jun 2026 3:21 UTC
192 points
40 comments2 min readLW link
(gwern.net)

Es­ti­mat­ing No-CoT Task-Com­ple­tion Time Hori­zons of Fron­tier AI Models

Anders Cairns Woodruff, Francis Rhys Ward, Dewi Gould, Rauno Arike, Jason R Brown, Jo Jiao, wlanderson, ariana_azarbal, harrymayne, Patrick Leask, Twm Stone, Josh Hills, Ida Caspary, Shubhorup Biswas and Julian Stastny
10 Jun 2026 17:58 UTC
278 points
23 comments4 min readLW link

Trees are mostly made of air and a gen­er­al­iz­able les­son for AI safety

Zephaniah Roe29 May 2026 4:08 UTC
324 points
62 comments4 min readLW link

Mnemonic por­traits for 19,023 hu­man genes

Brinedew28 May 2026 22:16 UTC
361 points
29 comments15 min readLW link

Models find­ing soft­ware vuln­er­a­bil­ities is not the pri­mary source of cy­ber­se­cu­rity risk

Dean Valentine14 May 2026 3:39 UTC
315 points
25 comments2 min readLW link

Em­pow­er­ment, cor­rigi­bil­ity, etc. are sim­ple ab­strac­tions (of a messed-up on­tol­ogy)

Steven Byrnes11 May 2026 17:48 UTC
193 points
74 comments16 min readLW link

Who Got Breasts First and How We Got Them

rba11 May 2026 13:11 UTC
165 points
59 comments10 min readLW link

Bad Prob­lems Don’t Stop Be­ing Bad Be­cause Some­body’s Wrong About Fault Analysis

Linch9 May 2026 1:30 UTC
268 points
76 comments3 min readLW link

x-risk-themed

kave6 May 2026 15:16 UTC
251 points
26 comments3 min readLW link
(kaverennedy.substack.com)

Ir­re­triev­abil­ity; or, Mur­phy’s Curse of Oneshot­ness upon ASI

Eliezer Yudkowsky4 May 2026 22:11 UTC
370 points
132 comments22 min readLW link

How Go Play­ers Disem­power Them­selves to AI

Ashe Vazquez Nuñez1 May 2026 23:24 UTC
744 points
79 comments8 min readLW link
Back to topNext