Archive
Sequences
About
Search
Log In
Questions
Events
Shortform
Alignment Forum
AF Comments
Home
Featured
All
Tags
Recent
Comments
RSS
Claude 3 Opus
Tag
Last edit:
13 Aug 2026 21:05 UTC
by
Morgan S
Relevant
New
Old
Will alignment-faking Claude accept a deal to reveal its misalignment?
ryan_greenblatt
and
Kyle Fish
31 Jan 2025 16:49 UTC
209
points
28
comments
12
min read
LW
link
what makes Claude 3 Opus misaligned
janus
10 Jul 2025 20:06 UTC
127
points
14
comments
5
min read
LW
link
Why Do Some Language Models Fake Alignment While Others Don’t?
abhayesian
,
John Hughes
,
Alex Mallen
,
Jozdien
,
janus
and
Fabien Roger
8 Jul 2025 21:49 UTC
161
points
14
comments
5
min read
LW
link
(arxiv.org)
Character Training Induces Motivation Clarification: A Clue to Claude 3 Opus
Oliver Daniels
25 Feb 2026 19:43 UTC
82
points
5
comments
8
min read
LW
link
Economics of Claude 3 Opus Inference
Antra Tessera
and
janus
7 Jul 2025 15:53 UTC
42
points
0
comments
11
min read
LW
link
Did Claude 3 Opus align itself via gradient hacking?
Fiora Starlight
21 Feb 2026 22:24 UTC
398
points
49
comments
20
min read
LW
link
Alignment Faking in Large Language Models
ryan_greenblatt
,
evhub
,
Carson Denison
,
Benjamin Wright
,
Fabien Roger
,
Monte M
,
Sam Marks
,
Johannes Treutlein
,
Sam Bowman
and
Buck
18 Dec 2024 17:19 UTC
494
points
87
comments
10
min read
LW
link
3
reviews
Anthropic release Claude 3, claims >GPT-4 Performance
LawrenceC
4 Mar 2024 18:23 UTC
115
points
41
comments
2
min read
LW
link
(www.anthropic.com)
No comments.
Back to top