RSS

Francis Rhys Ward

Karma: 883

Per­spec­tives on Con­tinual Learn­ing: Sur­vey Re­sults and Forecasts

24 Jun 2026 16:30 UTC
33 points
0 comments12 min readLW link

An­gles of at­tack for con­tinual learn­ing safety

16 Jun 2026 16:15 UTC
47 points
0 comments13 min readLW link

How might con­tinual learn­ing af­fect safety and al­ign­ment?

13 Jun 2026 17:34 UTC
60 points
2 comments16 min readLW link

What’s Con­tinual Learn­ing, and Why Might We Ex­pect To See It In Ad­vanced LLM Agents?

12 Jun 2026 18:43 UTC
32 points
2 comments17 min readLW link

Im­pli­ca­tions of Con­tinual Learn­ing for LLM Agents: Introduction

12 Jun 2026 18:36 UTC
49 points
0 comments6 min readLW link

Es­ti­mat­ing No-CoT Task-Com­ple­tion Time Hori­zons of Fron­tier AI Models

10 Jun 2026 17:58 UTC
275 points
23 comments4 min readLW link

Three types of model or­ganism

Francis Rhys Ward10 Jun 2026 8:50 UTC
58 points
8 comments2 min readLW link

[Paper] How does in­for­ma­tion ac­cess af­fect LLM mon­i­tors’ abil­ity to de­tect sab­o­tage?

11 Feb 2026 21:25 UTC
26 points
0 comments6 min readLW link

The Elic­i­ta­tion Game: Eval­u­at­ing ca­pa­bil­ity elic­i­ta­tion techniques

27 Feb 2025 20:33 UTC
15 points
1 comment2 min readLW link

Why care about AI per­son­hood?

Francis Rhys Ward26 Jan 2025 11:24 UTC
43 points
6 comments3 min readLW link

[Paper] AI Sand­bag­ging: Lan­guage Models can Strate­gi­cally Un­der­perform on Evaluations

13 Jun 2024 10:04 UTC
84 points
10 comments2 min readLW link
(arxiv.org)