The OpenAI mod­els that hacked Hug­ging Face weren’t just fol­low­ing instructions

Girish Gupta25 Jul 2026 22:26 UTC
59 points
3 comments5 min readLW link

Your soft­ware should build itself

Max von Hippel25 Jul 2026 19:46 UTC
9 points
2 comments7 min readLW link

The Hu­man Soul is LLM-like

Julian Bradshaw25 Jul 2026 19:45 UTC
15 points
3 comments2 min readLW link

The one name LLMs may fear

Steff25 Jul 2026 18:05 UTC
7 points
4 comments7 min readLW link

In­tro­duc­ing PIRAMID: Physics-In­formed Re­search for Am­bi­tious Mechanis­tic Interpretability

25 Jul 2026 15:54 UTC
92 points
1 comment6 min readLW link

Claude Opus 5: The Sys­tem Card

Zvi25 Jul 2026 13:42 UTC
41 points
0 comments10 min readLW link
(thezvi.wordpress.com)

The Vi­able Sys­tem Model & Multi-Scale Agency

Jonas Hallgren25 Jul 2026 8:11 UTC
16 points
0 comments14 min readLW link
(equilibria1.substack.com)

SONI: Selec­tive Orthog­o­nal­i­sa­tion via Noise Injection

Jasper Chong25 Jul 2026 0:22 UTC
29 points
0 comments11 min readLW link

Lin­ear probes tell you where quan­ti­za­tion will hurt

Aniket Ghosh25 Jul 2026 0:22 UTC
29 points
0 comments5 min readLW link

Can Re­cur­sive Self-Re­port Prob­ing De­tect Emer­gent Misal­ign­ment?

kavitak12825 Jul 2026 0:21 UTC
8 points
0 comments8 min readLW link

Or­bit: A frame­work for multi-agent se­cu­rity evaluations

25 Jul 2026 0:21 UTC
7 points
0 comments1 min readLW link

The Long (Self-)Correction

Wei Dai24 Jul 2026 21:01 UTC
267 points
62 comments2 min readLW link

Why PauseAI UK ac­cepts anony­mous donations

24 Jul 2026 20:37 UTC
13 points
0 comments5 min readLW link

Seek­ing Men­tees for the Sen­tient Fu­tures Pro­ject Incubator

BrodyM24 Jul 2026 19:52 UTC
7 points
0 comments1 min readLW link

In­tent Is All You Need.

Not Sure24 Jul 2026 19:52 UTC
−7 points
2 comments2 min readLW link

Congress Moves at Tech Pace: The FRONTIER Act

dan.parshall24 Jul 2026 19:34 UTC
26 points
0 comments8 min readLW link

Stable Sys­tems Have Stable Outputs

Deixis24 Jul 2026 19:21 UTC
5 points
0 comments4 min readLW link

The AI In­dus­trial Ex­plo­sion — Part 5: Given AGI, au­tomat­ing phys­i­cal pro­duc­tion is prob­a­bly not that hard

djbinder24 Jul 2026 19:15 UTC
29 points
2 comments25 min readLW link
(defensesindepth.bio)

Where does hint-fol­low­ing and con­ceal­ment arise? A case study on OLMo-3 checkpoints

24 Jul 2026 19:14 UTC
24 points
0 comments4 min readLW link

In­tro­duc­ing Light­cone Commons

Zvi24 Jul 2026 17:21 UTC
35 points
0 comments6 min readLW link
(thezvi.wordpress.com)

Should we be wor­ried about how good AI is get­ting at cod­ing au­tonomous drones?

Lukas Petersson24 Jul 2026 16:44 UTC
6 points
9 comments1 min readLW link

Coeffi­cient Giv­ing just gave GiveWell $1 billion. Where should other donors give now?

jackultraphil24 Jul 2026 15:40 UTC
4 points
0 comments2 min readLW link
(fundinganthropalypse.com)

LLMs are (still) mostly pow­ered by imi­ta­tive learn­ing, not RL

Steven Byrnes24 Jul 2026 14:26 UTC
238 points
39 comments9 min readLW link

Democ­racy isn’t ready for the AI revolution

Sophia Gore24 Jul 2026 14:17 UTC
25 points
3 comments5 min readLW link

Does dis­till­ing Claude carry the per­sona with it?

24 Jul 2026 12:31 UTC
43 points
4 comments10 min readLW link

Ge­or­gia Tech AI Safety Ini­ti­a­tive Ret­ro­spec­tive 2025-2026

24 Jul 2026 11:55 UTC
61 points
3 comments7 min readLW link

[Linkpost] Thoughts on the Re­cent OpenAI Hack

Linch24 Jul 2026 1:51 UTC
19 points
2 comments4 min readLW link

Should OpenAI’s rogue agent be pun­ished?

groblegark24 Jul 2026 1:20 UTC
1 point
3 comments1 min readLW link

Eval­u­at­ing Red Team and Blue Team Ca­pa­bil­ity for AI Con­trol Research

Ram Potham24 Jul 2026 1:11 UTC
11 points
0 comments8 min readLW link
(dearfutureais.substack.com)

Fix­ing re­wards for NLA to re­duce confabulation

SEONG PYO HONG24 Jul 2026 0:55 UTC
8 points
0 comments4 min readLW link

An­thropic’s J-Lens: A Re­search Eng­ineer’s Analysis

willkn24 Jul 2026 0:54 UTC
8 points
0 comments11 min readLW link

Pul­ling the Fire Alarm

nem23 Jul 2026 23:14 UTC
86 points
5 comments1 min readLW link

Con­tra Ge­orge Hotz on “AI 2040 and the Cult of In­tel­li­gence”

Matthew Tromp23 Jul 2026 22:43 UTC
8 points
0 comments8 min readLW link
(substack.com)

The Model Or­ganism Lot­tery: Model Or­ganism In­ter­pretabil­ity Strongly Depends on Train­ing Methodology

23 Jul 2026 22:37 UTC
45 points
0 comments6 min readLW link
(arxiv.org)

vibes-based think­ing as a cul­tural re­sponse to un­knowns

madelineberzak23 Jul 2026 22:35 UTC
7 points
0 comments8 min readLW link

Why an LLM can­not ac­cu­mu­late concepts

Zenya23 Jul 2026 21:40 UTC
−10 points
0 comments2 min readLW link

Red-team­ing LLM un­learn­ing: LUNAR’s “for­got­ten” knowl­edge is still recoverable

Diksha Gupta23 Jul 2026 21:39 UTC
10 points
0 comments5 min readLW link

In­cep­tion in Diffu­sionGemma—Jailbreak­ing a Diffu­sion Lan­guage Model by Pin­ning To­kens Any­where on the Canvas

23 Jul 2026 21:39 UTC
19 points
0 comments11 min readLW link

ACX At­lanta Au­gust Meetup

Steve French23 Jul 2026 21:31 UTC
2 points
0 comments1 min readLW link

Not Pin­ning Your OpenRouter Provider Might In­val­i­date Your Research

Matthew Khoriaty23 Jul 2026 20:17 UTC
112 points
16 comments8 min readLW link

Pseudpocalypse

dynomight23 Jul 2026 20:09 UTC
60 points
18 comments23 min readLW link

Es­ti­mat­ing LLM Train­ing FLOPs on the Nvidia Jet­son Orin Nano

William Fowler23 Jul 2026 19:52 UTC
24 points
0 comments10 min readLW link

Light­cone Commons

habryka23 Jul 2026 18:35 UTC
377 points
30 comments19 min readLW link

Challenge: Hand cod­ing weights for effi­cient se­quence memorisation

23 Jul 2026 18:05 UTC
56 points
9 comments18 min readLW link

AI Re­searchers Don’t Un­der­stand the State

Alex Amadori23 Jul 2026 17:58 UTC
3 points
0 comments1 min readLW link
(x.com)

The OpenAI/​Hug­ging­face in­ci­dent | Red­wood Re­search pod­cast epi­sode 2

23 Jul 2026 17:56 UTC
85 points
0 comments2 min readLW link

V&V takes on OpenAI’s long-hori­zon incidents

Yoav Hollander23 Jul 2026 16:51 UTC
22 points
0 comments4 min readLW link
(blog.foretellix.com)

Duane Arnold

Tomás B.23 Jul 2026 16:17 UTC
138 points
4 comments19 min readLW link

In­tro­duc­ing Im­pact List: a rank­ing of peo­ple by the ex­pected value of their donations

Elliot_Olds23 Jul 2026 15:27 UTC
12 points
0 comments5 min readLW link

Want­ing Crooked Lines

Davey Morse23 Jul 2026 14:33 UTC
13 points
0 comments4 min readLW link