RSS

David Duvenaud

Karma: 383

Sim­ple probes can catch sleeper agents

23 Apr 2024 21:10 UTC
130 points
18 comments1 min readLW link
(www.anthropic.com)

Sleeper Agents: Train­ing De­cep­tive LLMs that Per­sist Through Safety Training

12 Jan 2024 19:51 UTC
305 points
95 comments3 min readLW link
(arxiv.org)