Archive
Sequences
About
Search
Log In
Questions
Events
Shortform
Alignment Forum
AF Comments
Home
Featured
All
Tags
Recent
Comments
RSS
Puria
Karma:
512
I’m helping build geodesicresearch.ai
Personal writing at
puriaradmard.substack.com
All
Posts
Comments
New
Top
Old
Why study proto-training gaming as an adversarial alignment failure mode?
Puria
,
Edward James Young
and
Cam
8 Jul 2026 17:07 UTC
55
points
0
comments
7
min read
LW
link
Why study alignment interventions on pre-RL checkpoints?
Edward James Young
,
Puria
and
Cam
8 Jul 2026 17:07 UTC
66
points
2
comments
6
min read
LW
link
Announcing Geodesic Research
Puria
,
Cam
,
Alexandra Narin
,
Edward James Young
and
Kyle O’Brien
27 May 2026 16:40 UTC
85
points
2
comments
5
min read
LW
link
Learned Chain-of-Thought Obfuscation Generalises to Unseen Tasks
Nathaniel Mitrani
,
sassanb
,
Cam
and
Puria
21 May 2026 10:11 UTC
31
points
0
comments
5
min read
LW
link
(arxiv.org)
Alignment Pretraining: AI Discourse Causes Self-Fulfilling (Mis)alignment
Cam
,
Puria
,
Kyle O’Brien
,
David Africa
,
Samuel Ratnam
and
andyk
21 Dec 2025 0:53 UTC
207
points
25
comments
9
min read
LW
link
Architectures for Increased Externalisation of Reasoning
Karthik Viswanathan
,
Liza Pavlova
,
Mariia Koroliuk
,
Puria
,
Cam
and
Edward James Young
26 Nov 2025 20:24 UTC
38
points
2
comments
13
min read
LW
link
Generalisation Hacking: a first look at adversarial generalisation failures in deliberative alignment
Cam
and
Puria
17 Nov 2025 21:44 UTC
54
points
2
comments
8
min read
LW
link
I Am Large, I Contain Multitudes: Persona Transmission via Contextual Inference in LLMs
Shi
and
Puria
8 Sep 2025 13:52 UTC
33
points
0
comments
1
min read
LW
link
(www.researchgate.net)
Back to top