Archive
Sequences
About
Search
Log In
Questions
Events
Shortform
Alignment Forum
AF Comments
Home
Featured
All
Tags
Recent
Comments
RSS
New
Hot
Active
Old
Page
1
AI Doesn’t Have Free Will, Not Sure About Humans
mike20731
19 Jul 2026 2:56 UTC
3
points
0
comments
4
min read
LW
link
A Solution to Cryptographic Boxes for Unfriendly AI
Lysandre Terrisse
19 Jul 2026 0:33 UTC
3
points
0
comments
18
min read
LW
link
Nuances in the Workings of the Eye and Retina
Hieronym
and
Julian Bradshaw
18 Jul 2026 18:09 UTC
39
points
4
comments
20
min read
LW
link
My “Payorian FairBot” was just the original FairBot
transhumanist_atom_understander
18 Jul 2026 16:57 UTC
22
points
12
comments
6
min read
LW
link
Endogenous Alignment
Gordon Seidoh Worley
18 Jul 2026 14:10 UTC
24
points
4
comments
3
min read
LW
link
(www.uncertainupdates.com)
Map and Territory, Predictably Wrong
manueldelrio
18 Jul 2026 7:47 UTC
3
points
0
comments
2
min read
LW
link
The Most Forbidden Technique is not always forbidden
Rauno Arike
18 Jul 2026 0:59 UTC
64
points
3
comments
8
min read
LW
link
Should we benchmark conceptual capabilities using judgment prediction tasks?
Alex Mallen
17 Jul 2026 23:42 UTC
24
points
2
comments
3
min read
LW
link
A list of existing alignment approaches
Alek Westover
17 Jul 2026 22:46 UTC
8
points
0
comments
1
min read
LW
link
Longtermism is very intuitive.
tpotthinker
17 Jul 2026 22:35 UTC
−12
points
8
comments
5
min read
LW
link
AIs finetune their own leader: A barking simpleton
Shoshannah Tekofsky
17 Jul 2026 20:10 UTC
30
points
0
comments
5
min read
LW
link
(aivillageblog.substack.com)
Studying the role of Sandboxing for AI Control
Ram Potham
17 Jul 2026 19:05 UTC
12
points
0
comments
10
min read
LW
link
Would your AI travel agent book a bullfight? Testing whether agents consider animal welfare without being prompted
jonahmattwoodward
,
Jasmine Brazilek
,
Maheep Chaudhary
,
Oliver Tullio
,
Joel Christoph
and
MilesTS
17 Jul 2026 17:28 UTC
11
points
5
comments
3
min read
LW
link
Before values settle
Priyanka Bharadwaj
17 Jul 2026 16:26 UTC
16
points
2
comments
6
min read
LW
link
Reasons to believe current AI models are conscious
Eye You
17 Jul 2026 16:09 UTC
46
points
5
comments
8
min read
LW
link
What lawyers can do for AI safety
Martin Radzaj
17 Jul 2026 16:00 UTC
14
points
0
comments
7
min read
LW
link
Evolution of my AI Safety threat models
myyycroft
17 Jul 2026 14:54 UTC
7
points
0
comments
3
min read
LW
link
A Post-Mortem for My Goal Crystallisation Project
atryt0ne
and
Jason R Brown
17 Jul 2026 14:30 UTC
28
points
0
comments
10
min read
LW
link
Inoculation Adapters Improve Upon Inoculation Prompting
Maxime Riché
,
Daniel Tan
,
Vili Kohonen
and
nielsrolf
17 Jul 2026 14:00 UTC
84
points
2
comments
4
min read
LW
link
I don’t think Claude is misaligned in ‘Agentic Misalignment Summer 2026 - Motivated Mislabeling’
JohnWittle
17 Jul 2026 2:09 UTC
187
points
4
comments
12
min read
LW
link
Back to top
Next