Rogue Scalpel: Ac­ti­va­tion steer­ing breaks re­fusal, even with be­nign directions

Alexey Dontsov21 Aug 2026 23:46 UTC
14 points
0 comments3 min readLW link
(arxiv.org)

How the Sausage is Made—Why we Hate Slop

daniel_chernowitz21 Aug 2026 23:46 UTC
0 points
20 comments7 min readLW link

Con­tent-based priv­ilege: trans­former resi­d­ual streams strat­ify by prox­im­ity to the model’s own prediction

Nelson Guda21 Aug 2026 23:15 UTC
9 points
0 comments21 min readLW link

A J-Space-Based Met­ric for Model Valence: Defin­ing the Met­ric, Test­ing, and Com­par­i­sons to Self-Re­ports

Arjun Rao21 Aug 2026 21:21 UTC
8 points
0 comments12 min readLW link

Align­ment fine-tun­ing in­duces con­di­tional mis­al­ign­ment in Qwen2.5-7B-In­struct

Rhea Srivats21 Aug 2026 20:40 UTC
12 points
1 comment13 min readLW link

When is Un­limited Op­ti­miza­tion Catas­trophic?

Winter Cross21 Aug 2026 20:08 UTC
34 points
2 comments13 min readLW link
(arxiv.org)

AI Text Water­mark­ing Is Free And Good

Zvi21 Aug 2026 19:30 UTC
54 points
22 comments13 min readLW link
(thezvi.wordpress.com)

Eval­u­at­ing Ex­pla­na­tions of LLM Be­hav­ior In The Wild with Coun­ter­fac­tual Experiments

21 Aug 2026 19:09 UTC
73 points
6 comments7 min readLW link

In Defense of ASI Socialism

cguth721 Aug 2026 18:40 UTC
8 points
6 comments2 min readLW link

Misal­igned AI in the Bronze Age

frmsaul21 Aug 2026 16:51 UTC
56 points
6 comments7 min readLW link

The im­posters among us: func­tion vec­tors that ace ev­ery check and do the wrong task (in search of cir­cu­lar­ity)

star2vec21 Aug 2026 15:54 UTC
7 points
0 comments10 min readLW link

When Models Iden­tify as a Swarm

julius vidal21 Aug 2026 11:49 UTC
68 points
5 comments5 min readLW link

My Neel Nanda MATS 10.0 Ap­pli­ca­tion: Study­ing Fea­ture Split­ting in SAEs via Train­ing Data Attribution

J Rosser21 Aug 2026 9:41 UTC
19 points
0 comments13 min readLW link

Creativity Beyond the Manifold

Kartikay Luthra20 Aug 2026 23:31 UTC
8 points
2 comments6 min readLW link

Ablat­ing 1 of a chess trans­former’s 128 at­ten­tion heads makes the model stop find­ing Paul Mor­phy’s queen sac­ri­fice

David Litman20 Aug 2026 21:17 UTC
7 points
1 comment1 min readLW link
(www.youtube.com)

Llama will aban­don a cor­rect an­swer if it thinks you’re educated

Nick Merrill20 Aug 2026 20:06 UTC
33 points
7 comments2 min readLW link

Why self-fund your re­grant­ing?

20 Aug 2026 19:18 UTC
11 points
0 comments3 min readLW link
(manifund.substack.com)

If Aliens Ex­ist, We Should Ex­pect to Find Them Around Now

jehan20 Aug 2026 18:57 UTC
8 points
2 comments1 min readLW link
(www.jehanazad.com)

Cross-Dataset Trans­fer Eval­u­a­tion of De­cep­tion Probes in Smaller Models

Jollen Dai20 Aug 2026 17:06 UTC
8 points
1 comment5 min readLW link

The Fourth Humiliation

Nathalie Kirch20 Aug 2026 17:03 UTC
52 points
1 comment6 min readLW link

AI #182: Pause For Reflection

Zvi20 Aug 2026 14:50 UTC
37 points
1 comment47 min readLW link
(thezvi.wordpress.com)

We Must Re­mem­ber That Our World Con­tains Hell

James Brobin20 Aug 2026 14:22 UTC
152 points
52 comments3 min readLW link

Thoughts on Tak­ing OpenAI Foun­da­tion Funding

jefftk20 Aug 2026 13:00 UTC
31 points
3 comments5 min readLW link
(www.jefftk.com)

Mak­ing sense of the mis­al­ign­ment risk model in the An­thropic Risk Re­port (Au­gust 2026)

jasmine.ren20 Aug 2026 2:33 UTC
27 points
5 comments8 min readLW link

Ap­pear­ing Un­fair in the Wrong Direction

MossyFallenFriend20 Aug 2026 2:21 UTC
20 points
0 comments6 min readLW link

Oliver Habryka in­ter­view — Good Dis­course Needs Leadership

The Students20 Aug 2026 2:14 UTC
8 points
1 comment1 min readLW link
(www.youtube.com)

Steer­ing Role Confusion

20 Aug 2026 2:09 UTC
21 points
1 comment4 min readLW link

Notes on “Ei­genBench: A Com­par­a­tive Be­hav­ioral Mea­sure of Value Align­ment”

Shunk20 Aug 2026 2:03 UTC
8 points
0 comments5 min readLW link

Quan­tilized de­bate and con­sul­tancy in image en­vi­ron­ments: pro­to­col de­sign les­sons for scal­able over­sight experiments

emanuelr19 Aug 2026 23:03 UTC
14 points
0 comments28 min readLW link

Judg­ing eth­i­cal the­o­ries by up­date rules, not by ac­tion rankings

yatharth19 Aug 2026 22:41 UTC
13 points
1 comment5 min readLW link

Why can’t we have nice things? Like, speci­fi­cally?

Elizabeth19 Aug 2026 20:20 UTC
58 points
3 comments3 min readLW link
(acesounderglass.com)

OpenAI Takes Ini­tial Steps To Ad­dress Its Align­ment Problems

Zvi19 Aug 2026 20:00 UTC
33 points
1 comment21 min readLW link
(thezvi.wordpress.com)

34% of the US pub­lic is now aware of AI xrisk, and the curve is steepening

otto.barten19 Aug 2026 19:18 UTC
30 points
0 comments1 min readLW link

Science and News Twit­ter/​X Summarizer

sarahconstantin19 Aug 2026 19:00 UTC
32 points
5 comments1 min readLW link
(sarahconstantin.substack.com)

Separat­ing cheat­ing and aver­sion in task-gaming

Mihir Sahasrabudhe19 Aug 2026 18:49 UTC
14 points
0 comments4 min readLW link

MATS Win­ter 2027 Ap­pli­ca­tions Are Now Open

Raj Thimmiah19 Aug 2026 18:48 UTC
8 points
2 comments2 min readLW link

The Rogue Agent Ex­plo­sion Will Be Mostly Invisible

Steven McCulloch19 Aug 2026 18:46 UTC
102 points
20 comments16 min readLW link

RL cre­ates split personas

Jan Betley19 Aug 2026 18:23 UTC
235 points
20 comments4 min readLW link

En­force­ment Be­gins: the EU AI Act and Sys­temic Risks

Schizoid Rentoid19 Aug 2026 18:04 UTC
1 point
0 comments1 min readLW link

A cir­cuit prior in NN-bayes

19 Aug 2026 17:14 UTC
78 points
4 comments1 min readLW link

A failed solu­tion to open-source game theory

Richard Willis19 Aug 2026 16:11 UTC
15 points
4 comments4 min readLW link

In­side the mind of a fair player cooperating

transhumanist_atom_understander19 Aug 2026 14:53 UTC
40 points
2 comments3 min readLW link

Some rea­sons al­ign­ment doesn’t gen­er­al­ise well

Lucius Bushnaq19 Aug 2026 14:18 UTC
129 points
4 comments9 min readLW link

The Con­trol Paradox

Ephraiem Sarabamoun19 Aug 2026 12:39 UTC
8 points
0 comments2 min readLW link

De­bate Train­ing Re­duces Re­ward Hack­ing in RLAIF

19 Aug 2026 12:17 UTC
61 points
6 comments6 min readLW link
(gdmalignment.substack.com)

Con­cerns About Per­sonas, Multi-Agent Align­ment, and Role Theory

Davidmanheim19 Aug 2026 12:03 UTC
21 points
1 comment8 min readLW link

Read­ing List on Wise AI

Chris_Leong19 Aug 2026 8:41 UTC
22 points
0 comments1 min readLW link

Nat­u­ral Lan­guage Transcoders

anwenh19 Aug 2026 1:49 UTC
14 points
0 comments4 min readLW link

What AI scores (while we can still keep score)

dan.parshall19 Aug 2026 1:40 UTC
12 points
5 comments3 min readLW link

Cui bono? ChatGPT-4o shows non-de­cep­tive strate­gic per­sua­sion: a Proof-of-Con­cept study

Pieter Barkema18 Aug 2026 23:25 UTC
7 points
0 comments5 min readLW link