Archive
Sequences
About
Search
Log In
Questions
Events
Shortform
Alignment Forum
AF Comments
Home
Featured
All
Tags
Recent
Comments
RSS
New
Hot
Active
Old
Page
1
V&V takes on “Pacing the frontier”
Yoav Hollander
14 Aug 2026 13:25 UTC
3
points
0
comments
9
min read
LW
link
(blog.foretellix.com)
Avoiding the corporate treacherous turn: crowd-sourcing design ideas
Stuart_Armstrong
14 Aug 2026 13:10 UTC
16
points
0
comments
1
min read
LW
link
Frontier agents don’t comply with standards, even when instructed to
Daan Henselmans
and
Arno Libert
14 Aug 2026 12:08 UTC
41
points
2
comments
7
min read
LW
link
What Mormons get right about community building
Jacob Brinton
14 Aug 2026 7:53 UTC
5
points
1
comment
4
min read
LW
link
Don’t forget why learning is important
Roman Ross
14 Aug 2026 7:11 UTC
9
points
1
comment
5
min read
LW
link
What If We Enforced AI Model Safety At the Level Of GPUs?
Mayowa Osibodu
14 Aug 2026 6:10 UTC
6
points
1
comment
4
min read
LW
link
Features that current AIs don’t have that future AIs will have
Alexander Gietelink Oldenziel
13 Aug 2026 21:49 UTC
32
points
2
comments
2
min read
LW
link
Some Ways I Think About Evaluating Grant Applications
sarahconstantin
13 Aug 2026 21:30 UTC
42
points
0
comments
10
min read
LW
link
(sarahconstantin.substack.com)
The Left Should Start Taking AI Capabilities Seriously
Alexei G
13 Aug 2026 21:03 UTC
−13
points
0
comments
9
min read
LW
link
(www.onethousandmeans.com)
Is Alignment Even Falsifiable? Middle Alignment, An Alignment Taxonomy, and Breaking The Problem Down Into Steps
Savannah Harlan
13 Aug 2026 20:57 UTC
5
points
0
comments
22
min read
LW
link
How to Answer a Question Without Answering The Question
Kabir Kumar
13 Aug 2026 17:06 UTC
31
points
1
comment
1
min read
LW
link
How My Students Think About AI
dvd
13 Aug 2026 16:56 UTC
256
points
11
comments
12
min read
LW
link
Automated alignment runs are hard to study!
Alejandro Aristizabal
,
draganover
,
Aleksandr Bowkis
and
Cameron Holmes
13 Aug 2026 15:04 UTC
57
points
3
comments
9
min read
LW
link
Longtermism Seems Like A Religion
James Brobin
13 Aug 2026 14:43 UTC
0
points
8
comments
3
min read
LW
link
Defense Against the Deceptive Arts
Kabir Kumar
13 Aug 2026 14:21 UTC
17
points
8
comments
1
min read
LW
link
Some problems in decision theory incorrectly precondition on policy.
Canaletto
13 Aug 2026 12:04 UTC
15
points
2
comments
3
min read
LW
link
Patterns and problems in emerging multiagent systems (Anthropic, Frontier Red Team)
Julian Bradshaw
13 Aug 2026 4:10 UTC
41
points
0
comments
1
min read
LW
link
(www.anthropic.com)
The Descenders and the Absorbers: Two Perspectives on Deep Learning
larry-dial
13 Aug 2026 4:02 UTC
11
points
0
comments
3
min read
LW
link
Free will is like temperature
Optimization Process
12 Aug 2026 23:32 UTC
43
points
8
comments
1
min read
LW
link
Every reward-hacked policy I tested triggered the OPE alarm — and so did my best honest one
JulesRoussel01
12 Aug 2026 22:51 UTC
−4
points
0
comments
48
min read
LW
link
Back to top
Next