RSS

AI swarms are start­ing to pose in­di­rect takeover risk

12 Aug 2026 5:05 UTC
127 points
5 comments10 min readLW link

An any­time al­gorithm for mix­ing the com­putable measures

Cole Wyeth12 Aug 2026 0:57 UTC
24 points
0 comments4 min readLW link

Four LLM loss func­tions → four fla­vors of LLM misalignment

Steven Byrnes10 Aug 2026 16:16 UTC
218 points
12 comments6 min readLW link

Why do mod­els task game?

6 Aug 2026 22:16 UTC
87 points
1 comment24 min readLW link

User aware­ness in fron­tier models

6 Aug 2026 20:43 UTC
58 points
1 comment12 min readLW link
(transluce.org)

R-lens: Mak­ing J-lens More Faith­ful on Early Layers

5 Aug 2026 20:02 UTC
70 points
4 comments7 min readLW link

Re­turn­ing to ARC

paulfchristiano4 Aug 2026 22:27 UTC
364 points
32 comments9 min readLW link

Con­crete Eval­u­a­tions to In­ves­ti­gate the OpenAI Model That Hacked Hug­ging Face

3 Aug 2026 9:23 UTC
133 points
6 comments37 min readLW link

Value Leak­age: An LLM’s An­swers Are Silently Shaped by Its Own Values

31 Jul 2026 16:32 UTC
75 points
7 comments17 min readLW link

AGI Safety and Align­ment at Google Deep­Mind: A Sum­mary of Re­cent Work (July 2026)

31 Jul 2026 15:57 UTC
85 points
0 comments9 min readLW link
(gdmalignment.substack.com)

The AGI Safety and Align­ment team at Google Deep­Mind is Hiring (July 2026)

31 Jul 2026 15:53 UTC
67 points
2 comments6 min readLW link
(gdmalignment.substack.com)

OpenAI has already ended an in­ter­nal pause

Charbel-Raphaël31 Jul 2026 12:03 UTC
109 points
0 comments1 min readLW link

Thou­sand-di­men­sional structure

30 Jul 2026 14:04 UTC
161 points
6 comments10 min readLW link

Im­pre­cise be­liefs: a tiny introduction

davidad29 Jul 2026 22:17 UTC
79 points
36 comments6 min readLW link

Value Gen­er­al­i­sa­tion 3: Pre-al­igned AIs

Stuart_Armstrong29 Jul 2026 15:58 UTC
16 points
0 comments4 min readLW link

Value Gen­er­al­i­sa­tion 2: The Miss­ing Hole in AIs’ abilities

Stuart_Armstrong29 Jul 2026 15:58 UTC
16 points
8 comments10 min readLW link

Value Gen­er­al­i­sa­tion 1: a Re­search and De­ploy­ment Program

Stuart_Armstrong29 Jul 2026 15:57 UTC
22 points
2 comments4 min readLW link

Re­search di­rec­tions in con­den­sa­tion: va­ri­eties of objectivity

SamEisenstat28 Jul 2026 17:30 UTC
58 points
1 comment12 min readLW link

RL & search is a ter­rify­ing way to build AGI (an FAQ)

Steven Byrnes27 Jul 2026 14:50 UTC
139 points
15 comments14 min readLW link

The Long (Self-)Correction

Wei Dai24 Jul 2026 21:01 UTC
240 points
62 comments2 min readLW link