Archive
Sequences
About
Search
Log In
Questions
Events
Shortform
Alignment Forum
AF Comments
Home
Featured
All
Tags
Recent
Comments
RSS
jcksanderson
Karma:
224
All
Posts
Comments
New
Top
Old
We’re talking past our models; or, How a model defined its “evil” vector as dread
jcksanderson
20 Jul 2026 7:27 UTC
45
points
5
comments
9
min read
LW
link
Interpretability is becoming increasingly uninterpretable
jcksanderson
9 Jul 2026 7:04 UTC
34
points
1
comment
4
min read
LW
link
(jcksanderson.com)
Don’t ignore the car crashes, and remember your freshman CS
jcksanderson
26 Jun 2026 7:06 UTC
38
points
0
comments
2
min read
LW
link
(jcksanderson.com)
jcksanderson’s Shortform
jcksanderson
24 Jun 2026 10:31 UTC
2
points
6
comments
1
min read
LW
link
Iterative Finetuning is Mostly Idempotent
Zephaniah Roe
,
jcksanderson
and
Julian H
11 May 2026 6:41 UTC
24
points
0
comments
5
min read
LW
link
Introducing the XLab AI Security Guide
Zephaniah Roe
,
jcksanderson
and
Julian H
27 Dec 2025 16:50 UTC
20
points
1
comment
5
min read
LW
link
Intriguing Properties of gpt-oss Jailbreaks
Zephaniah Roe
and
jcksanderson
13 Aug 2025 19:42 UTC
20
points
0
comments
10
min read
LW
link
(xlabaisecurity.com)
Back to top