Archive
Sequences
About
Search
Log In
Questions
Events
Shortform
Alignment Forum
AF Comments
Home
Featured
All
Tags
Recent
Comments
RSS
New
Hot
Active
Old
Page
1
Deadlock in the Parliament of the Self
Lorxus
10 Oct 2026 18:36 UTC
11
points
0
comments
13
min read
LW
link
(tiled-with-pentagons.blogspot.com)
exfiltration through self-distillation
jonathanbreitg
10 Oct 2026 17:49 UTC
7
points
5
comments
3
min read
LW
link
The potentially deadly threat of AI output-optimization
Steff
10 Oct 2026 17:13 UTC
4
points
0
comments
6
min read
LW
link
The Problem With“Doomers” and “Optimists”
Olivia Scharfman
10 Oct 2026 17:08 UTC
9
points
0
comments
5
min read
LW
link
The Non-Compassionate Case for Model Welfare
ixotope
10 Oct 2026 16:16 UTC
7
points
6
comments
5
min read
LW
link
(ixotopic.substack.com)
Much more than you wanted to know about wombats
becausecurious
10 Oct 2026 15:57 UTC
14
points
0
comments
2
min read
LW
link
Utilitarianism and Autism
Walter Veit
10 Oct 2026 14:45 UTC
2
points
2
comments
5
min read
LW
link
(walterveit.substack.com)
Examing Emergent Misalignment in a recurrent LLM with a logit lens
nesiacel
10 Oct 2026 14:22 UTC
8
points
0
comments
4
min read
LW
link
Inheritance of Refusals from Abliterated Models
Minh Hoang
10 Oct 2026 10:57 UTC
11
points
0
comments
7
min read
LW
link
Claude Haiku 4.5 submits false police tip; Anthropic takes 72 days to notice
becausecurious
10 Oct 2026 4:43 UTC
13
points
0
comments
3
min read
LW
link
An Alignment Forum for AIs? (or: Verification in the Age of Slop)
Raemon
10 Oct 2026 2:51 UTC
81
points
21
comments
4
min read
LW
link
Cracks in the Narcissus Mirror
Gladys Preysler
10 Oct 2026 0:22 UTC
5
points
1
comment
5
min read
LW
link
(gladyspreysler.substack.com)
[Paper] Distillation for Incrimination and Distillation for Capabilities
sebastian_prasanna
and
Alek Westover
9 Oct 2026 22:14 UTC
43
points
1
comment
8
min read
LW
link
(blog.redwoodresearch.org)
Admiration and Alignment
Alex Mussgnug
9 Oct 2026 18:06 UTC
3
points
0
comments
8
min read
LW
link
What can language models teach us about understanding?
Asvin
9 Oct 2026 18:02 UTC
19
points
0
comments
10
min read
LW
link
Information-Proliferating Dynamics
interstice
9 Oct 2026 17:06 UTC
9
points
0
comments
6
min read
LW
link
(thermontology.com)
Clarifying types of rogue AI activity: Breakout, breakin, exfiltration, … What’s what?
Oliver Sourbut
9 Oct 2026 17:03 UTC
14
points
0
comments
10
min read
LW
link
(www.oliversourbut.net)
The Public Intellectual Is Dead. And AI Has Killed Them.
Walter Veit
9 Oct 2026 16:49 UTC
3
points
1
comment
4
min read
LW
link
(walterveit.substack.com)
Filling in the convex hull
gjm
9 Oct 2026 16:03 UTC
28
points
0
comments
8
min read
LW
link
I can’t believe it’s not BUTTER!
becausecurious
9 Oct 2026 15:45 UTC
77
points
11
comments
4
min read
LW
link
Back to top
Next