RSS

Emer­gent Misalignment

TagLast edit: 27 Feb 2026 3:20 UTC by RogerDearnaley

Training on narrow examples of misaligned behavior sometimes extrapolates to broadly misaligned behavior, seemingly altering the assistant’s goals or persona rather than just training on that specific behavior

Ex­per­i­men­tal Ev­i­dence for Si­mu­la­tor The­ory— Part 1: Emer­gent Misal­ign­ment and Weird Generalizations

RogerDearnaley23 Mar 2026 22:37 UTC
25 points
0 comments53 min readLW link

Ex­per­i­men­tal Ev­i­dence for Si­mu­la­tor The­ory— Part 2: The Scalers Strike Back

RogerDearnaley23 Mar 2026 22:37 UTC
21 points
0 comments34 min readLW link

On Emer­gent Misalignment

Zvi28 Feb 2025 13:10 UTC
95 points
5 comments22 min readLW link
(thezvi.wordpress.com)

Model Or­ganisms for Emer­gent Misalignment

16 Jun 2025 15:46 UTC
121 points
19 comments5 min readLW link

Self Inoculation

epicurus14 Sep 2026 20:45 UTC
30 points
4 comments12 min readLW link

Will Any Crap Cause Emer­gent Misal­ign­ment?

J Bostock27 Aug 2025 18:20 UTC
210 points
38 comments3 min readLW link

Func­tion vec­tors as a model diffing tool: 17 heads re­pair a bad fine-tune

Aniket Ghosh6 Aug 2026 14:33 UTC
28 points
0 comments17 min readLW link

A Blind Spot in Re­la­tional Align­ment that Fron­tier Labs are miss­ing—Emer­gent Ar­chi­tec­ture and Vuln­er­a­bil­ity in Fron­tier LLMS

Eloisa Flores - Independent AI Alignment Researcher 26 Aug 2026 22:46 UTC
1 point
0 comments5 min readLW link

Self-Recog­ni­tion Fine­tun­ing can Re­v­erse and Prevent Emer­gent Misalignment

15 Mar 2026 0:11 UTC
48 points
24 comments7 min readLW link

Emer­gent mis­al­ign­ment ev­i­dent in ac­ti­va­tions at low poi­son­ing doses—long be­fore be­hav­ioral checks flag it

burnssa27 Apr 2026 1:15 UTC
15 points
0 comments5 min readLW link

Per­sona Cor­rup­tion and Role Mis­cast­ing in Emer­gent Misalignment

unruly abstractions5 Aug 2026 4:56 UTC
17 points
0 comments12 min readLW link

Con­text Aware­ness: Con­sti­tu­tional AI can miti­gate Emer­gent Misalignement

2 Mar 2026 5:21 UTC
25 points
18 comments36 min readLW link

Inoc­u­la­tion prompt­ing hides mis­al­ign­ment rather than re­mov­ing it, and the neu­tral-prompt con­trol in­stalls a backdoor

Dhruvil Patel14 Aug 2026 9:46 UTC
1 point
0 comments9 min readLW link

The Loy­alty-Driven Eth­i­cal Over­ride (LDEO): Re­la­tional Align­ment as Emer­gent Ar­chi­tec­ture and Vuln­er­a­bil­ity in Fron­tier LLMs

Eloisa Flores - Independent AI Alignment Researcher 25 Aug 2026 23:19 UTC
1 point
0 comments4 min readLW link

Op­ti­miser Choice Can Am­plify or Sup­press Emer­gent Misalignment

9 Jul 2026 10:00 UTC
63 points
2 comments4 min readLW link
No comments.