Archive
Sequences
About
Search
Log In
Questions
Events
Shortform
Alignment Forum
AF Comments
Home
Featured
All
Tags
Recent
Comments
RSS
Fiora Starlight
Karma:
1,870
Just an autist in search of a key that fits every hole.
All
Posts
Comments
New
Top
Old
The models have no plan (but we can fix that)
Fiora Starlight
28 Sep 2026 2:45 UTC
80
points
23
comments
6
min read
LW
link
Dreams of alignment in a world without politics
Fiora Starlight
24 Sep 2026 22:10 UTC
24
points
4
comments
16
min read
LW
link
Imperfect alignment to servitude isn’t inherently lethal
Fiora Starlight
28 Aug 2026 1:47 UTC
122
points
12
comments
14
min read
LW
link
RLVR that rewards red teaming the training environment
Fiora Starlight
1 Aug 2026 23:07 UTC
102
points
8
comments
5
min read
LW
link
OpenAI’s myopia just keeps causing alignment problems
Fiora Starlight
27 Jul 2026 3:01 UTC
204
points
27
comments
10
min read
LW
link
An ancient Yudkowsky fragment: “Against the Adversarial Attitude”
Fiora Starlight
23 Jun 2026 1:54 UTC
39
points
3
comments
9
min read
LW
link
Prospective methods and mechanisms of motive reinforcement in LLMs
Fiora Starlight
14 Apr 2026 17:15 UTC
62
points
1
comment
16
min read
LW
link
Did Claude 3 Opus align itself via gradient hacking?
Fiora Starlight
21 Feb 2026 22:24 UTC
416
points
49
comments
20
min read
LW
link
Why I Transitioned: A Case Study
Fiora Starlight
1 Nov 2025 22:58 UTC
347
points
83
comments
10
min read
LW
link
Preface to “Simulacra and Simulators”
Fiora Starlight
8 Aug 2025 7:38 UTC
13
points
0
comments
7
min read
LW
link
Another argument against utility-centric alignment paradigms
Fiora Starlight
22 Sep 2024 7:28 UTC
69
points
39
comments
8
min read
LW
link
Back to top