Archive
Sequences
About
Search
Log In
Questions
Events
Shortform
Alignment Forum
AF Comments
Home
Featured
All
Tags
Recent
Comments
RSS
Maheep Chaudhary
Karma:
52
All
Posts
Comments
New
Top
Old
Coercion and Deception in AI-to-AI Management
jonahmattwoodward
,
Jasmine Brazilek
,
MilesTS
and
Maheep Chaudhary
10 Aug 2026 16:13 UTC
12
points
0
comments
8
min read
LW
link
(compassionalignedml.substack.com)
Would your AI travel agent book a bullfight? Testing whether agents consider animal welfare without being prompted
jonahmattwoodward
,
Jasmine Brazilek
,
Maheep Chaudhary
,
Oliver Tullio
,
Joel Christoph
and
MilesTS
17 Jul 2026 17:28 UTC
13
points
14
comments
3
min read
LW
link
Awareness Jailbreaking: Revealing True Alignment in Evaluation-Aware Models
Maheep Chaudhary
29 Dec 2025 21:29 UTC
11
points
0
comments
4
min read
LW
link
Evaluation Awareness Scales Predictably in Open-Weights Large Language Models
Maheep Chaudhary
19 Dec 2025 2:47 UTC
23
points
0
comments
6
min read
LW
link
Back to top