RSS

Jacob Pfau

Karma: 1,771

Scalable oversight at Resolution

Models don’t seem to be dishon­est in the way hu­mans are

22 Jul 2026 15:32 UTC
45 points
4 comments9 min readLW link

De­bate with Self-Play Best-of-N Optimization

9 Jul 2026 15:29 UTC
49 points
2 comments14 min readLW link

An­nounc­ing our $160M grant from Coeffi­cient Giving

9 Jul 2026 7:58 UTC
49 points
0 comments2 min readLW link

Re­search up­date: RL on De­bate Games shows Pro­posal Ac­cu­racy up­lift alongside Judge Hacking

2 Jul 2026 17:42 UTC
77 points
4 comments21 min readLW link

Re­s­olu­tion (fka Se­quent): scale and au­toma­tion for higher con­fi­dence in alignment

10 Jun 2026 15:37 UTC
297 points
5 comments11 min readLW link
(sequent.org)

Au­to­mated Align­ment is Harder Than You Think

14 May 2026 22:01 UTC
143 points
7 comments3 min readLW link
(arxiv.org)

From per­sonas to in­ten­tions: to­wards a sci­ence of mo­ti­va­tions for AI models

14 Apr 2026 12:26 UTC
81 points
6 comments7 min readLW link