RSS

Carson Denison

Karma: 1,542

I work on deceptive alignment and reward hacking at Anthropic