Archive
Sequences
About
Search
Log In
Questions
Events
Shortform
Alignment Forum
AF Comments
Home
Featured
All
Tags
Recent
Comments
RSS
New
Hot
Active
Old
Page
1
The Rogue Agent Explosion Will Be Mostly Invisible
Steven McCulloch
19 Aug 2026 18:46 UTC
18
points
1
comment
16
min read
LW
link
RL creates split personas
Jan Betley
19 Aug 2026 18:23 UTC
47
points
0
comments
4
min read
LW
link
A failed solution to open-source game theory
Richard Willis
19 Aug 2026 16:11 UTC
14
points
2
comments
4
min read
LW
link
Inside the mind of a fair player cooperating
transhumanist_atom_understander
19 Aug 2026 14:53 UTC
18
points
0
comments
3
min read
LW
link
Some reasons alignment doesn’t generalise well
Lucius Bushnaq
19 Aug 2026 14:18 UTC
54
points
0
comments
9
min read
LW
link
The Control Paradox
Ephraiem Sarabamoun
19 Aug 2026 12:39 UTC
5
points
0
comments
2
min read
LW
link
Concerns About Personas, Multi-Agent Alignment, and Role Theory
Davidmanheim
19 Aug 2026 12:03 UTC
18
points
0
comments
8
min read
LW
link
Reading List on Wise AI
Chris_Leong
19 Aug 2026 8:41 UTC
19
points
0
comments
1
min read
LW
link
What AI scores (while we can still keep score)
dan.parshall
19 Aug 2026 1:40 UTC
11
points
2
comments
3
min read
LW
link
Funding Formal Methods for the Cyberpocalypse
Max von Hippel
18 Aug 2026 16:45 UTC
21
points
6
comments
9
min read
LW
link
Policy career planning in the age of imminent superintelligence
Peter Wildeford
18 Aug 2026 14:33 UTC
60
points
2
comments
6
min read
LW
link
Automatic Programming Should Be More Like SQL
Adam Chlipala
18 Aug 2026 12:29 UTC
7
points
0
comments
12
min read
LW
link
AI Security is Harm Reduction
Quinn
18 Aug 2026 12:28 UTC
39
points
1
comment
1
min read
LW
link
Price recursion is the rational theory of reward
Abhimanyu Pallavi Sudhir
18 Aug 2026 3:22 UTC
6
points
0
comments
9
min read
LW
link
Misaligned Incentives in Pause Scenarios
Michael Soareverix
and
Antra Tessera
18 Aug 2026 1:55 UTC
39
points
10
comments
17
min read
LW
link
For Claude, capability and dispreferring CDT are the ~same thing. Much more so than for GPT.
Chi Nguyen
and
Emery Cooper
18 Aug 2026 1:25 UTC
47
points
11
comments
1
min read
LW
link
You Can’t Iterate to Trustworthy AI Code Without Understanding
ronbodkin
17 Aug 2026 23:00 UTC
8
points
0
comments
10
min read
LW
link
What gives you away: how LLMs form opinions of you
Cat McGee
17 Aug 2026 20:59 UTC
37
points
7
comments
6
min read
LW
link
Evaluating Chain-of-Thought Monitorability is Still an Open Problem: Comments on OpenAI’s Monitorability Evals
Connor Dilgren
and
Sarah Wiegreffe
17 Aug 2026 19:45 UTC
17
points
1
comment
18
min read
LW
link
Weird Re-Tokenization, Symmetries and Compression: Research Agenda
Xenomirant
and
Sami Wolf
17 Aug 2026 19:43 UTC
19
points
14
comments
16
min read
LW
link
Back to top
Next