RSS

Johannes Treutlein

Karma: 1,845

All opinions are my own. Homepage: johannestreutlein.com

Steer­ing to­wards “au­to­mated grad­ing” de­grades alignment

3 Sep 2026 18:17 UTC
155 points
35 comments6 min readLW link

Value Leak­age: An LLM’s An­swers Are Silently Shaped by Its Own Values

31 Jul 2026 16:32 UTC
75 points
8 comments17 min readLW link

Eval­u­at­ing hon­esty and lie de­tec­tion tech­niques on a di­verse suite of dishon­est models

25 Nov 2025 19:33 UTC
41 points
0 comments4 min readLW link
(alignment.anthropic.com)

Build­ing and eval­u­at­ing al­ign­ment au­dit­ing agents

24 Jul 2025 19:22 UTC
47 points
1 comment5 min readLW link

Mod­ify­ing LLM Beliefs with Syn­thetic Doc­u­ment Finetuning

24 Apr 2025 21:15 UTC
77 points
12 comments2 min readLW link
(alignment.anthropic.com)

Au­dit­ing lan­guage mod­els for hid­den objectives

13 Mar 2025 19:18 UTC
159 points
15 comments13 min readLW link

Align­ment Fak­ing in Large Lan­guage Models

18 Dec 2024 17:19 UTC
494 points
87 comments10 min readLW link3 reviews

Con­nect­ing the Dots: LLMs can In­fer & Ver­bal­ize La­tent Struc­ture from Train­ing Data

21 Jun 2024 15:54 UTC
166 points
14 comments8 min readLW link1 review
(arxiv.org)

Re­port on mod­el­ing ev­i­den­tial co­op­er­a­tion in large worlds

Johannes Treutlein12 Jul 2023 16:37 UTC
47 points
3 comments1 min readLW link
(arxiv.org)