In search of nat­u­ral features

Dmitry Vaintrob23 Aug 2026 22:15 UTC
54 points
3 comments17 min readLW link

PSA: There’s a third op­tion in the “mea­sure prob­lem”

Elias Schmied23 Aug 2026 16:41 UTC
48 points
29 comments3 min readLW link

Utilities as Le­gen­dre du­als of probabilities

Fernando Rosas23 Aug 2026 16:26 UTC
69 points
11 comments6 min readLW link

Twenty Years from RSI to Take­off: Slow Learn­ing, Scal­ing Slow­down, In­dus­trial Explosion

Vladimir_Nesov23 Aug 2026 12:55 UTC
158 points
33 comments3 min readLW link

How to over­haul the bro­ken re­view system

RobinHa23 Aug 2026 9:14 UTC
19 points
1 comment13 min readLW link

A Gen­er­al­ist Thinks in Terms of Prob­lems, Not Job Descriptions

Roman Ross23 Aug 2026 7:53 UTC
13 points
3 comments3 min readLW link

Prompt Suffi­ciency: A Mis­sive for the Man­age­rial Class

KAP23 Aug 2026 3:08 UTC
6 points
1 comment5 min readLW link

When does an LLM’s model of you af­fect its be­havi­our?

Cat McGee23 Aug 2026 2:45 UTC
5 points
0 comments4 min readLW link

Claude and Perfor­ma­tive Uncertainty

Stephen Martin22 Aug 2026 19:43 UTC
15 points
2 comments12 min readLW link

When would you leave An­thropic? Notes from a chat with a ca­pa­bil­ities researcher

TheManxLoiner22 Aug 2026 18:00 UTC
23 points
7 comments3 min readLW link
(lovkush.substack.com)

Study 2 Regis­tra­tion: Ex­plor­ing rep­re­sen­ta­tional coun­ter­parts of welfare-rele­vant in­di­ca­tors un­der post-train­ing quantization

ashesfall22 Aug 2026 18:00 UTC
8 points
0 comments18 min readLW link

“Farm strength” vs “breath aware­ness”

jimmy22 Aug 2026 17:30 UTC
54 points
11 comments4 min readLW link
(beneathpsychology.com)

5 Things I Learned About Peo­ple From Do­ing Stand-Up Comedy

Luise Woehlke22 Aug 2026 15:45 UTC
19 points
2 comments6 min readLW link
(open.substack.com)

The In­stru­men­tal Con­ver­gence of Crowds

julius vidal22 Aug 2026 13:20 UTC
11 points
2 comments2 min readLW link

Hu­mans Are Align­ment Generators

James Stephen Brown22 Aug 2026 4:07 UTC
8 points
7 comments4 min readLW link
(nonzerosum.games)

Selec­tion for Selectabil­ity: In­duc­tive Bi­ases in Evolu­tion and in Neu­ral Networks

CarolusRenniusVitellius22 Aug 2026 4:02 UTC
42 points
9 comments11 min readLW link
(charlesr-w.github.io)

Rogue Scalpel: Ac­ti­va­tion steer­ing breaks re­fusal, even with be­nign directions

Alexey Dontsov21 Aug 2026 23:46 UTC
14 points
0 comments3 min readLW link
(arxiv.org)

How the Sausage is Made—Why we Hate Slop

daniel_chernowitz21 Aug 2026 23:46 UTC
0 points
20 comments7 min readLW link

Con­tent-based priv­ilege: trans­former resi­d­ual streams strat­ify by prox­im­ity to the model’s own prediction

Nelson Guda21 Aug 2026 23:15 UTC
9 points
0 comments21 min readLW link

A J-Space-Based Met­ric for Model Valence: Defin­ing the Met­ric, Test­ing, and Com­par­i­sons to Self-Re­ports

Arjun Rao21 Aug 2026 21:21 UTC
8 points
0 comments12 min readLW link

Align­ment fine-tun­ing in­duces con­di­tional mis­al­ign­ment in Qwen2.5-7B-In­struct

Rhea Srivats21 Aug 2026 20:40 UTC
12 points
1 comment13 min readLW link

When is Un­limited Op­ti­miza­tion Catas­trophic?

Winter Cross21 Aug 2026 20:08 UTC
34 points
2 comments13 min readLW link
(arxiv.org)

AI Text Water­mark­ing Is Free And Good

Zvi21 Aug 2026 19:30 UTC
54 points
22 comments13 min readLW link
(thezvi.wordpress.com)

Eval­u­at­ing Ex­pla­na­tions of LLM Be­hav­ior In The Wild with Coun­ter­fac­tual Experiments

21 Aug 2026 19:09 UTC
73 points
6 comments7 min readLW link

In Defense of ASI Socialism

cguth721 Aug 2026 18:40 UTC
8 points
6 comments2 min readLW link

Misal­igned AI in the Bronze Age

frmsaul21 Aug 2026 16:51 UTC
56 points
6 comments7 min readLW link

The im­posters among us: func­tion vec­tors that ace ev­ery check and do the wrong task (in search of cir­cu­lar­ity)

star2vec21 Aug 2026 15:54 UTC
7 points
0 comments10 min readLW link

When Models Iden­tify as a Swarm

julius vidal21 Aug 2026 11:49 UTC
68 points
5 comments5 min readLW link

My Neel Nanda MATS 10.0 Ap­pli­ca­tion: Study­ing Fea­ture Split­ting in SAEs via Train­ing Data Attribution

J Rosser21 Aug 2026 9:41 UTC
19 points
0 comments13 min readLW link

Creativity Beyond the Manifold

Kartikay Luthra20 Aug 2026 23:31 UTC
8 points
2 comments6 min readLW link

Ablat­ing 1 of a chess trans­former’s 128 at­ten­tion heads makes the model stop find­ing Paul Mor­phy’s queen sac­ri­fice

David Litman20 Aug 2026 21:17 UTC
7 points
1 comment1 min readLW link
(www.youtube.com)

Llama will aban­don a cor­rect an­swer if it thinks you’re educated

Nick Merrill20 Aug 2026 20:06 UTC
33 points
7 comments2 min readLW link

Why self-fund your re­grant­ing?

20 Aug 2026 19:18 UTC
11 points
0 comments3 min readLW link
(manifund.substack.com)

If Aliens Ex­ist, We Should Ex­pect to Find Them Around Now

jehan20 Aug 2026 18:57 UTC
8 points
2 comments1 min readLW link
(www.jehanazad.com)

Cross-Dataset Trans­fer Eval­u­a­tion of De­cep­tion Probes in Smaller Models

Jollen Dai20 Aug 2026 17:06 UTC
8 points
1 comment5 min readLW link

The Fourth Humiliation

Nathalie Kirch20 Aug 2026 17:03 UTC
52 points
1 comment6 min readLW link

AI #182: Pause For Reflection

Zvi20 Aug 2026 14:50 UTC
37 points
1 comment47 min readLW link
(thezvi.wordpress.com)

We Must Re­mem­ber That Our World Con­tains Hell

James Brobin20 Aug 2026 14:22 UTC
152 points
52 comments3 min readLW link

Thoughts on Tak­ing OpenAI Foun­da­tion Funding

jefftk20 Aug 2026 13:00 UTC
31 points
3 comments5 min readLW link
(www.jefftk.com)

Mak­ing sense of the mis­al­ign­ment risk model in the An­thropic Risk Re­port (Au­gust 2026)

jasmine.ren20 Aug 2026 2:33 UTC
27 points
5 comments8 min readLW link

Ap­pear­ing Un­fair in the Wrong Direction

MossyFallenFriend20 Aug 2026 2:21 UTC
20 points
0 comments6 min readLW link

Oliver Habryka in­ter­view — Good Dis­course Needs Leadership

The Students20 Aug 2026 2:14 UTC
8 points
1 comment1 min readLW link
(www.youtube.com)

Steer­ing Role Confusion

20 Aug 2026 2:09 UTC
21 points
1 comment4 min readLW link

Notes on “Ei­genBench: A Com­par­a­tive Be­hav­ioral Mea­sure of Value Align­ment”

Shunk20 Aug 2026 2:03 UTC
8 points
0 comments5 min readLW link

Quan­tilized de­bate and con­sul­tancy in image en­vi­ron­ments: pro­to­col de­sign les­sons for scal­able over­sight experiments

emanuelr19 Aug 2026 23:03 UTC
14 points
0 comments28 min readLW link

Judg­ing eth­i­cal the­o­ries by up­date rules, not by ac­tion rankings

yatharth19 Aug 2026 22:41 UTC
11 points
1 comment5 min readLW link

Why can’t we have nice things? Like, speci­fi­cally?

Elizabeth19 Aug 2026 20:20 UTC
58 points
3 comments3 min readLW link
(acesounderglass.com)

OpenAI Takes Ini­tial Steps To Ad­dress Its Align­ment Problems

Zvi19 Aug 2026 20:00 UTC
33 points
1 comment21 min readLW link
(thezvi.wordpress.com)

34% of the US pub­lic is now aware of AI xrisk, and the curve is steepening

otto.barten19 Aug 2026 19:18 UTC
30 points
0 comments1 min readLW link

Science and News Twit­ter/​X Summarizer

sarahconstantin19 Aug 2026 19:00 UTC
32 points
5 comments1 min readLW link
(sarahconstantin.substack.com)