Archive
Sequences
About
Search
Log In
Questions
Events
Shortform
Alignment Forum
AF Comments
Home
Featured
All
Tags
Recent
Comments
RSS
Stuart_Armstrong
Karma:
18,575
All
Posts
Comments
New
Top
Old
Page
1
Value generalisation Theory of Change: putting it into practice
Stuart_Armstrong
31 Aug 2026 10:51 UTC
20
points
4
comments
7
min read
LW
link
Value generalisation Theory of Change: the theory behind the approach
Stuart_Armstrong
28 Aug 2026 13:30 UTC
27
points
0
comments
9
min read
LW
link
Avoiding the corporate treacherous turn: crowd-sourcing design ideas
Stuart_Armstrong
14 Aug 2026 13:10 UTC
26
points
13
comments
1
min read
LW
link
Simplifying the anthropic impossibility result
Stuart_Armstrong
8 Aug 2026 8:25 UTC
24
points
13
comments
3
min read
LW
link
III. Anthropic reasoning has issues with infinite worlds; D-SIA can fix this
Stuart_Armstrong
3 Aug 2026 21:31 UTC
18
points
5
comments
17
min read
LW
link
II. Anthropic reasoning with duplication is not consistent with the usual probability properties
Stuart_Armstrong
3 Aug 2026 14:21 UTC
15
points
37
comments
7
min read
LW
link
I. Anthropic reasoning without duplicates is just standard Bayesian updating
Stuart_Armstrong
30 Jul 2026 13:08 UTC
36
points
6
comments
19
min read
LW
link
Value Generalisation 3: Pre-aligned AIs
Stuart_Armstrong
29 Jul 2026 15:58 UTC
16
points
2
comments
4
min read
LW
link
Value Generalisation 2: The Missing Hole in AIs’ abilities
Stuart_Armstrong
29 Jul 2026 15:58 UTC
16
points
8
comments
10
min read
LW
link
Value Generalisation 1: a Research and Deployment Program
Stuart_Armstrong
29 Jul 2026 15:57 UTC
27
points
2
comments
4
min read
LW
link
The true “test” dataset for a generalised task
Stuart_Armstrong
27 Jul 2026 16:16 UTC
23
points
0
comments
2
min read
LW
link
Occam’s razor is about using the past to predict the future
Stuart_Armstrong
15 Jul 2026 19:35 UTC
52
points
6
comments
3
min read
LW
link
Value generalisation: value correction
Stuart_Armstrong
10 Jul 2026 7:56 UTC
25
points
3
comments
6
min read
LW
link
Pragmatic FDT, and predictors as game theory
Stuart_Armstrong
3 Jul 2026 13:22 UTC
36
points
12
comments
11
min read
LW
link
The future of alignment if LLMs are a bubble
Stuart_Armstrong
23 Dec 2025 0:08 UTC
51
points
13
comments
5
min read
LW
link
Go home GPT-4o, you’re drunk: emergent misalignment as lowered inhibitions
Stuart_Armstrong
and
rgorman
18 Mar 2025 14:48 UTC
79
points
12
comments
5
min read
LW
link
Using Prompt Evaluation to Combat Bio-Weapon Research
Stuart_Armstrong
and
rgorman
19 Feb 2025 12:39 UTC
11
points
2
comments
3
min read
LW
link
Defense Against the Dark Prompts: Mitigating Best-of-N Jailbreaking with Prompt Evaluation
Stuart_Armstrong
and
rgorman
31 Jan 2025 15:36 UTC
16
points
2
comments
2
min read
LW
link
Alignment can improve generalisation through more robustly doing what a human wants—CoinRun example
Stuart_Armstrong
21 Nov 2023 11:41 UTC
67
points
9
comments
3
min read
LW
link
How toy models of ontology changes can be misleading
Stuart_Armstrong
21 Oct 2023 21:13 UTC
42
points
0
comments
2
min read
LW
link
Back to top
Next