Archive
Sequences
About
Search
Log In
Questions
Events
Shortform
Alignment Forum
AF Comments
Home
Featured
All
Tags
Recent
Comments
RSS
New
Hot
Active
Old
Page
1
Alignment Midtraining Cracks Under Pressure
J Bostock
,
sidbaines
,
Daniel Tan
,
draganover
,
ma-rmartinez
and
David Africa
21 Sep 2026 16:55 UTC
19
points
0
comments
6
min read
LW
link
We’re not ready for the e/Acc × Longevity preference cascade
Jackson Wagner
21 Sep 2026 15:21 UTC
19
points
6
comments
11
min read
LW
link
A class of statement between conjecture and theorem
Jason Fantl
21 Sep 2026 15:18 UTC
7
points
2
comments
2
min read
LW
link
Gratitude and the End of the World
Ephraiem Sarabamoun
21 Sep 2026 15:06 UTC
4
points
0
comments
1
min read
LW
link
Weight smuggling likely defeats attempts to cap FLOPs per training run
Paul W
and
Pierre Peigné
21 Sep 2026 14:50 UTC
24
points
0
comments
1
min read
LW
link
Grantmakers: Consider sharing rough probabilities & brief feedback with grantees/applicants
david reinstein
21 Sep 2026 14:46 UTC
10
points
0
comments
4
min read
LW
link
Mech Interp is a Verifiable Task
Logan Riggs
21 Sep 2026 14:42 UTC
18
points
1
comment
5
min read
LW
link
A Brief History of Koinometry
Benjamin Schneider
21 Sep 2026 13:15 UTC
4
points
0
comments
4
min read
LW
link
(benjaminschneider.ch)
Empirical safety claims from frontier labs should be replicated, scrutinized, and open-sourced
Zephaniah Roe
and
yix
21 Sep 2026 5:58 UTC
53
points
2
comments
4
min read
LW
link
(secondlookresearch.com)
Where Do Chatbots Come From? What I Wish Everyone Knew About AI in 2026
Vaughn Papenhausen
21 Sep 2026 1:05 UTC
17
points
0
comments
18
min read
LW
link
(vaughnpapenhausen.substack.com)
Pasta Marketing, Magic Players, and Political Movements
J Bostock
20 Sep 2026 22:33 UTC
19
points
0
comments
7
min read
LW
link
(jbostock.substack.com)
Why do they even talk about x-risk?
MieszkoP
20 Sep 2026 22:23 UTC
19
points
16
comments
3
min read
LW
link
Evaluating task vectors, unlearning and inoculation
Xenomirant
20 Sep 2026 17:20 UTC
12
points
0
comments
10
min read
LW
link
Reflections on unlearning and inoculation
Xenomirant
20 Sep 2026 15:59 UTC
11
points
0
comments
11
min read
LW
link
Mistakes in time
Vincent_Bagayoko
20 Sep 2026 14:33 UTC
6
points
0
comments
1
min read
LW
link
Giving up control
Karl von Wendt
20 Sep 2026 13:41 UTC
19
points
4
comments
4
min read
LW
link
Did Someone Check if Rogue Agents are Interested in Self-Improvement?
Ephraiem Sarabamoun
20 Sep 2026 12:52 UTC
15
points
1
comment
1
min read
LW
link
What I’ve Learned About Depression so Far
teegs
20 Sep 2026 11:38 UTC
18
points
0
comments
20
min read
LW
link
Labs could soon start automated research into architectures driven by no-CoT perfomance
nanowell
20 Sep 2026 10:49 UTC
16
points
0
comments
1
min read
LW
link
We’ve saved the world before: what the ozone hole teaches us about AI
leogao
20 Sep 2026 8:23 UTC
114
points
13
comments
9
min read
LW
link
Back to top
Next