RSS

Puria

Karma: 269

I’m helping build geodesicresearch.org

Align­ment Pre­train­ing: AI Dis­course Causes Self-Fulfilling (Mis)alignment

21 Dec 2025 0:53 UTC
194 points
24 comments9 min readLW link

Ar­chi­tec­tures for In­creased Ex­ter­nal­i­sa­tion of Reasoning

26 Nov 2025 20:24 UTC
20 points
2 comments13 min readLW link

Gen­er­al­i­sa­tion Hack­ing: a first look at ad­ver­sar­ial gen­er­al­i­sa­tion failures in de­liber­a­tive alignment

17 Nov 2025 21:44 UTC
46 points
2 comments8 min readLW link

I Am Large, I Con­tain Mul­ti­tudes: Per­sona Trans­mis­sion via Con­tex­tual In­fer­ence in LLMs

8 Sep 2025 13:52 UTC
33 points
0 comments1 min readLW link
(www.researchgate.net)