Archive
Sequences
About
Search
Log In
Questions
Events
Shortform
Alignment Forum
AF Comments
Home
Featured
All
Tags
Recent
Comments
RSS
New
Hot
Active
Old
Page
1
Mitigating Reward Hacking as Institutional Design
beren
12 Sep 2026 6:17 UTC
9
points
0
comments
32
min read
LW
link
Let’s Own the Term “Elitism”
Martin Sustrik
12 Sep 2026 6:00 UTC
12
points
0
comments
3
min read
LW
link
(www.250bpm.com)
What I want you to do when I tell you to “think about your theory of change more carefully”
Roman Ross
12 Sep 2026 2:51 UTC
8
points
0
comments
5
min read
LW
link
Some ways AI could kill us all
Ruby
12 Sep 2026 1:08 UTC
56
points
3
comments
9
min read
LW
link
A normal Friday in 2042
RobinHa
11 Sep 2026 22:29 UTC
14
points
0
comments
13
min read
LW
link
My recommended resources for AI safety, alignment, and existential risks
Lysandre Terrisse
11 Sep 2026 22:19 UTC
8
points
0
comments
2
min read
LW
link
Consider positive feedback loops
Patodesu
11 Sep 2026 21:55 UTC
7
points
0
comments
3
min read
LW
link
Alignment Hierarchy
Gadersd
11 Sep 2026 21:26 UTC
2
points
0
comments
4
min read
LW
link
Astra’s no-CoT limits track speculative depth, not step count
MBaert
11 Sep 2026 18:05 UTC
77
points
2
comments
9
min read
LW
link
CoT controllability evals seem very under-elicited
Jozdien
11 Sep 2026 17:12 UTC
58
points
1
comment
4
min read
LW
link
Local Factor Graph Debate
Alexander Heckett
11 Sep 2026 16:59 UTC
21
points
0
comments
9
min read
LW
link
Post-AGI, we are all jobless aristocrats
djbinder
11 Sep 2026 16:53 UTC
34
points
6
comments
3
min read
LW
link
(defensesindepth.bio)
SFT Also Drives Safety Eval Results in Olmo 3
Finn Cairns
11 Sep 2026 16:20 UTC
26
points
0
comments
1
min read
LW
link
(secondlookresearch.com)
Generalized UDT 1.0 tiling
Roman Malov
11 Sep 2026 12:28 UTC
18
points
0
comments
6
min read
LW
link
We need good evals for activation faithfulness
Anthony Hughes
,
draganover
and
Andersehen
11 Sep 2026 8:20 UTC
23
points
0
comments
2
min read
LW
link
Questions for the “New Enlightenment” in the Age of AGI
Jordan Arel
11 Sep 2026 5:49 UTC
8
points
0
comments
4
min read
LW
link
Voluntary Gradual Disempowerment in the Judiciary/Legal System
Caleb Horn
11 Sep 2026 4:00 UTC
14
points
0
comments
1
min read
LW
link
Which character are we evaluating? Persona stability and AI welfare
Joshua Fonseca Rivera
11 Sep 2026 3:04 UTC
15
points
0
comments
5
min read
LW
link
Perspectives in favour of improving conceptual reasoning capabilities
Chi Nguyen
11 Sep 2026 1:32 UTC
16
points
2
comments
7
min read
LW
link
Subagents comply more
jacob_drori
10 Sep 2026 23:11 UTC
26
points
0
comments
3
min read
LW
link
Back to top
Next