RSS

Stuart_Armstrong

Karma: 18,575

Value gen­er­al­i­sa­tion The­ory of Change: putting it into practice

Stuart_Armstrong31 Aug 2026 10:51 UTC
20 points
4 comments7 min readLW link

Value gen­er­al­i­sa­tion The­ory of Change: the the­ory be­hind the approach

Stuart_Armstrong28 Aug 2026 13:30 UTC
27 points
0 comments9 min readLW link

Avoid­ing the cor­po­rate treach­er­ous turn: crowd-sourc­ing de­sign ideas

Stuart_Armstrong14 Aug 2026 13:10 UTC
26 points
13 comments1 min readLW link

Sim­plify­ing the an­thropic im­pos­si­bil­ity result

Stuart_Armstrong8 Aug 2026 8:25 UTC
24 points
13 comments3 min readLW link

III. An­thropic rea­son­ing has is­sues with in­finite wor­lds; D-SIA can fix this

Stuart_Armstrong3 Aug 2026 21:31 UTC
18 points
5 comments17 min readLW link

II. An­thropic rea­son­ing with du­pli­ca­tion is not con­sis­tent with the usual prob­a­bil­ity properties

Stuart_Armstrong3 Aug 2026 14:21 UTC
15 points
37 comments7 min readLW link

I. An­thropic rea­son­ing with­out du­pli­cates is just stan­dard Bayesian updating

Stuart_Armstrong30 Jul 2026 13:08 UTC
36 points
6 comments19 min readLW link

Value Gen­er­al­i­sa­tion 3: Pre-al­igned AIs

Stuart_Armstrong29 Jul 2026 15:58 UTC
16 points
2 comments4 min readLW link

Value Gen­er­al­i­sa­tion 2: The Miss­ing Hole in AIs’ abilities

Stuart_Armstrong29 Jul 2026 15:58 UTC
16 points
8 comments10 min readLW link

Value Gen­er­al­i­sa­tion 1: a Re­search and De­ploy­ment Program

Stuart_Armstrong29 Jul 2026 15:57 UTC
27 points
2 comments4 min readLW link

The true “test” dataset for a gen­er­al­ised task

Stuart_Armstrong27 Jul 2026 16:16 UTC
23 points
0 comments2 min readLW link

Oc­cam’s ra­zor is about us­ing the past to pre­dict the future

Stuart_Armstrong15 Jul 2026 19:35 UTC
52 points
6 comments3 min readLW link

Value gen­er­al­i­sa­tion: value correction

Stuart_Armstrong10 Jul 2026 7:56 UTC
25 points
3 comments6 min readLW link

Prag­matic FDT, and pre­dic­tors as game theory

Stuart_Armstrong3 Jul 2026 13:22 UTC
36 points
12 comments11 min readLW link

The fu­ture of al­ign­ment if LLMs are a bubble

Stuart_Armstrong23 Dec 2025 0:08 UTC
51 points
13 comments5 min readLW link

Go home GPT-4o, you’re drunk: emer­gent mis­al­ign­ment as low­ered inhibitions

18 Mar 2025 14:48 UTC
79 points
12 comments5 min readLW link

Us­ing Prompt Eval­u­a­tion to Com­bat Bio-Weapon Research

19 Feb 2025 12:39 UTC
11 points
2 comments3 min readLW link

Defense Against the Dark Prompts: Miti­gat­ing Best-of-N Jailbreak­ing with Prompt Evaluation

31 Jan 2025 15:36 UTC
16 points
2 comments2 min readLW link

Align­ment can im­prove gen­er­al­i­sa­tion through more ro­bustly do­ing what a hu­man wants—CoinRun example

Stuart_Armstrong21 Nov 2023 11:41 UTC
67 points
9 comments3 min readLW link

How toy mod­els of on­tol­ogy changes can be misleading

Stuart_Armstrong21 Oct 2023 21:13 UTC
42 points
0 comments2 min readLW link