RSS

De­bate Train­ing Re­duces Re­ward Hack­ing in RLAIF

19 Aug 2026 12:17 UTC
49 points
4 comments6 min readLW link
(gdmalignment.substack.com)

Does Diffu­sionGemma do la­tent rea­son­ing?

16 Aug 2026 4:22 UTC
39 points
0 comments9 min readLW link

AI swarms are start­ing to pose in­di­rect takeover risk

12 Aug 2026 5:05 UTC
137 points
5 comments10 min readLW link

An any­time al­gorithm for mix­ing the com­putable measures

Cole Wyeth12 Aug 2026 0:57 UTC
26 points
0 comments4 min readLW link

Misal­igned AIs could use kil­ler robots to take over

11 Aug 2026 19:03 UTC
131 points
9 comments5 min readLW link
(turntrout.com)

Four LLM loss func­tions → four fla­vors of LLM misalignment

Steven Byrnes10 Aug 2026 16:16 UTC
329 points
30 comments6 min readLW link

Why do mod­els task game?

6 Aug 2026 22:16 UTC
88 points
1 comment24 min readLW link

User aware­ness in fron­tier models

6 Aug 2026 20:43 UTC
59 points
1 comment12 min readLW link
(transluce.org)

R-lens: Mak­ing J-lens More Faith­ful on Early Layers

5 Aug 2026 20:02 UTC
75 points
5 comments7 min readLW link

Re­turn­ing to ARC

paulfchristiano4 Aug 2026 22:27 UTC
374 points
35 comments9 min readLW link

Con­crete Eval­u­a­tions to In­ves­ti­gate the OpenAI Model That Hacked Hug­ging Face

3 Aug 2026 9:23 UTC
133 points
6 comments37 min readLW link

Value Leak­age: An LLM’s An­swers Are Silently Shaped by Its Own Values

31 Jul 2026 16:32 UTC
75 points
7 comments17 min readLW link

AGI Safety and Align­ment at Google Deep­Mind: A Sum­mary of Re­cent Work (July 2026)

31 Jul 2026 15:57 UTC
85 points
0 comments9 min readLW link
(gdmalignment.substack.com)

The AGI Safety and Align­ment team at Google Deep­Mind is Hiring (July 2026)

31 Jul 2026 15:53 UTC
71 points
2 comments6 min readLW link
(gdmalignment.substack.com)

OpenAI has already ended an in­ter­nal pause

Charbel-Raphaël31 Jul 2026 12:03 UTC
109 points
0 comments1 min readLW link

Thou­sand-di­men­sional structure

30 Jul 2026 14:04 UTC
167 points
7 comments10 min readLW link

Im­pre­cise be­liefs: a tiny introduction

davidad29 Jul 2026 22:17 UTC
79 points
36 comments6 min readLW link

Value Gen­er­al­i­sa­tion 3: Pre-al­igned AIs

Stuart_Armstrong29 Jul 2026 15:58 UTC
16 points
0 comments4 min readLW link

Value Gen­er­al­i­sa­tion 2: The Miss­ing Hole in AIs’ abilities

Stuart_Armstrong29 Jul 2026 15:58 UTC
16 points
8 comments10 min readLW link

Value Gen­er­al­i­sa­tion 1: a Re­search and De­ploy­ment Program

Stuart_Armstrong29 Jul 2026 15:57 UTC
22 points
2 comments4 min readLW link