RSS

Dead­lock in the Par­li­a­ment of the Self

Lorxus10 Oct 2026 18:36 UTC
11 points
0 comments13 min readLW link
(tiled-with-pentagons.blogspot.com)

exfil­tra­tion through self-distillation

jonathanbreitg10 Oct 2026 17:49 UTC
7 points
5 comments3 min readLW link

The po­ten­tially deadly threat of AI out­put-optimization

Steff10 Oct 2026 17:13 UTC
4 points
0 comments6 min readLW link

The Prob­lem With“Doomers” and “Op­ti­mists”

Olivia Scharfman10 Oct 2026 17:08 UTC
9 points
0 comments5 min readLW link

The Non-Com­pas­sion­ate Case for Model Welfare

ixotope10 Oct 2026 16:16 UTC
7 points
6 comments5 min readLW link
(ixotopic.substack.com)

Much more than you wanted to know about wombats

becausecurious10 Oct 2026 15:57 UTC
14 points
0 comments2 min readLW link

Utili­tar­i­anism and Autism

Walter Veit10 Oct 2026 14:45 UTC
2 points
2 comments5 min readLW link
(walterveit.substack.com)

Ex­am­ing Emer­gent Misal­ign­ment in a re­cur­rent LLM with a logit lens

nesiacel10 Oct 2026 14:22 UTC
8 points
0 comments4 min readLW link

In­her­i­tance of Re­fusals from Abliter­ated Models

Minh Hoang10 Oct 2026 10:57 UTC
11 points
0 comments7 min readLW link

Claude Haiku 4.5 sub­mits false po­lice tip; An­thropic takes 72 days to notice

becausecurious10 Oct 2026 4:43 UTC
13 points
0 comments3 min readLW link

An Align­ment Fo­rum for AIs? (or: Ver­ifi­ca­tion in the Age of Slop)

Raemon10 Oct 2026 2:51 UTC
81 points
21 comments4 min readLW link

Cracks in the Nar­cis­sus Mirror

Gladys Preysler10 Oct 2026 0:22 UTC
5 points
1 comment5 min readLW link
(gladyspreysler.substack.com)

[Paper] Distil­la­tion for In­crim­i­na­tion and Distil­la­tion for Capabilities

9 Oct 2026 22:14 UTC
43 points
1 comment8 min readLW link
(blog.redwoodresearch.org)

Ad­mira­tion and Align­ment

Alex Mussgnug9 Oct 2026 18:06 UTC
3 points
0 comments8 min readLW link

What can lan­guage mod­els teach us about un­der­stand­ing?

Asvin9 Oct 2026 18:02 UTC
19 points
0 comments10 min readLW link

In­for­ma­tion-Pro­lifer­at­ing Dynamics

interstice9 Oct 2026 17:06 UTC
9 points
0 comments6 min readLW link
(thermontology.com)

Clar­ify­ing types of rogue AI ac­tivity: Break­out, breakin, exfil­tra­tion, … What’s what?

Oliver Sourbut9 Oct 2026 17:03 UTC
14 points
0 comments10 min readLW link
(www.oliversourbut.net)

The Public In­tel­lec­tual Is Dead. And AI Has Killed Them.

Walter Veit9 Oct 2026 16:49 UTC
3 points
1 comment4 min readLW link
(walterveit.substack.com)

Filling in the con­vex hull

gjm9 Oct 2026 16:03 UTC
28 points
0 comments8 min readLW link

I can’t be­lieve it’s not BUTTER!

becausecurious9 Oct 2026 15:45 UTC
77 points
11 comments4 min readLW link