RSS

Inoc­u­late or Reflect? Two train­ing in­ter­ven­tions un­der prompt­ing, steer­ing, and patching

26 Jul 2026 18:06 UTC
8 points
0 comments3 min readLW link

Plan A, by AI-2040

PeterMcCluskey26 Jul 2026 17:15 UTC
16 points
0 comments8 min readLW link
(bayesianinvestor.com)

AI use policy for my es­say writing

Kaj_Sotala26 Jul 2026 14:39 UTC
34 points
0 comments5 min readLW link
(kajsotala.substack.com)

Sim­ple Dutch books ver­sus Sleep­ing Beauty halfers

jessicata26 Jul 2026 5:41 UTC
15 points
9 comments7 min readLW link

An OpenAI model left notes about how to evade con­tain­ment; we need more details

Alex Mallen26 Jul 2026 3:53 UTC
166 points
3 comments4 min readLW link

The OpenAI mod­els that hacked Hug­ging Face weren’t just fol­low­ing instructions

Girish Gupta25 Jul 2026 22:26 UTC
40 points
1 comment5 min readLW link

Your soft­ware should build itself

Max von Hippel25 Jul 2026 19:46 UTC
9 points
1 comment7 min readLW link

The Hu­man Soul is LLM-like

Julian Bradshaw25 Jul 2026 19:45 UTC
12 points
2 comments2 min readLW link

The one name LLMs may fear

Steff25 Jul 2026 18:05 UTC
11 points
1 comment7 min readLW link

The Vi­able Sys­tem Model & Multi-Scale Agency

Jonas Hallgren25 Jul 2026 8:11 UTC
15 points
0 comments14 min readLW link
(equilibria1.substack.com)

The Long (Self-)Correction

Wei Dai24 Jul 2026 21:01 UTC
119 points
29 comments2 min readLW link

In­tent Is All You Need.

Not Sure24 Jul 2026 19:52 UTC
−4 points
1 comment2 min readLW link

Stable Sys­tems Have Stable Outputs

Deixis24 Jul 2026 19:21 UTC
5 points
0 comments4 min readLW link

The AI In­dus­trial Ex­plo­sion — Part 5: Given AGI, au­tomat­ing phys­i­cal pro­duc­tion is prob­a­bly not that hard

djbinder24 Jul 2026 19:15 UTC
23 points
2 comments25 min readLW link
(defensesindepth.bio)

Where does hint-fol­low­ing and con­ceal­ment arise? A case study on OLMo-3 checkpoints

24 Jul 2026 19:14 UTC
24 points
0 comments4 min readLW link

Should we be wor­ried about how good AI is get­ting at cod­ing au­tonomous drones?

Lukas Petersson24 Jul 2026 16:44 UTC
6 points
9 comments1 min readLW link

LLMs are (still) mostly pow­ered by imi­ta­tive learn­ing, not RL

Steven Byrnes24 Jul 2026 14:26 UTC
140 points
25 comments9 min readLW link

Democ­racy isn’t ready for the AI revolution

Sophia Gore24 Jul 2026 14:17 UTC
24 points
3 comments5 min readLW link

Does dis­till­ing Claude carry the per­sona with it?

24 Jul 2026 12:31 UTC
35 points
3 comments10 min readLW link

Ge­or­gia Tech AI Safety Ini­ti­a­tive Ret­ro­spec­tive 2025-2026

24 Jul 2026 11:55 UTC
38 points
2 comments7 min readLW link