OpenAI Models Be­hind Hug­gingFace Cy­ber­se­cu­rity Incident

LawrenceC21 Jul 2026 21:36 UTC
220 points
20 comments1 min readLW link

WeirdChat: A cat­a­log of un­ex­pected AI be­hav­iors, dis­cov­ered automatically

neilchowdhury21 Jul 2026 21:12 UTC
63 points
4 comments10 min readLW link
(transluce.org)

OpenAI Shares Some Align­ment Problems

Zvi21 Jul 2026 19:41 UTC
149 points
8 comments9 min readLW link
(thezvi.wordpress.com)

Blog­ging Tech­nol­ogy Interlude

Zack_M_Davis21 Jul 2026 19:03 UTC
14 points
1 comment3 min readLW link
(zackmdavis.net)

End­ing Soon: Fun­da­men­tal Uncer­tainty $2,000 Es­say Contest

Gordon Seidoh Worley21 Jul 2026 16:51 UTC
15 points
0 comments1 min readLW link
(www.uncertainupdates.com)

Steer­ing Black­mail Through a Model’s “Emo­tional State”

21 Jul 2026 16:03 UTC
12 points
0 comments8 min readLW link

Mea­sur­ing Re­ward-Seek­ing via Con­trastive Belief Updates

21 Jul 2026 15:22 UTC
91 points
5 comments12 min readLW link
(rewardseeking.ai)

Mea­sur­ing Re­ward-Seek­ing by In­still­ing Con­trastive Beliefs

papetoast21 Jul 2026 15:11 UTC
12 points
0 comments1 min readLW link
(alignment.openai.com)

11 Open Em­piri­cal Prob­lems in Re­ward-Seeking

21 Jul 2026 15:08 UTC
65 points
0 comments8 min readLW link

Differ­en­tial ac­cel­er­a­tion of al­ign­ment-rele­vant ca­pa­bil­ities is a bad bet

Zephaniah Roe21 Jul 2026 14:04 UTC
114 points
8 comments8 min readLW link

Epistemics and Co­or­di­na­tion: It’s com­pli­cated!

Raymond Douglas21 Jul 2026 12:23 UTC
54 points
13 comments8 min readLW link

(1/​3) The Dangers of LLMs

Eigenbraid21 Jul 2026 12:01 UTC
10 points
0 comments6 min readLW link

I ran the stan­dard AI lit­mus tests on my two tod­dlers (yep)

Carlo Valenti21 Jul 2026 11:38 UTC
45 points
8 comments4 min readLW link

Unoffi­cial ACX/​LW/​EA Mu­nich Bi-weekly Com­mu­nity Dinner

hilll21 Jul 2026 9:55 UTC
1 point
0 comments1 min readLW link

NameRank: The Model Knows Your Pro­ject, Not You.

Jarrett Ye21 Jul 2026 2:21 UTC
20 points
7 comments1 min readLW link
(01.me)

The Case for Phys­i­cal AI Safety

20 Jul 2026 22:18 UTC
13 points
2 comments22 min readLW link
(paisi.ai)

Ad­der­all Tol­er­ance: Much More Than You Wanted To Know

Kurt H. Pieper20 Jul 2026 21:55 UTC
97 points
13 comments4 min readLW link
(kurthpieper.substack.com)

What do I mean by “Ar­tifi­cial Gen­eral In­tel­li­gence”?

Steven Byrnes20 Jul 2026 21:13 UTC
44 points
5 comments4 min readLW link

AI 2040: Is it Ac­tu­ally a Deal?

1a3orn20 Jul 2026 20:58 UTC
55 points
13 comments9 min readLW link

Banana in, Bostrom out: pa­per­clip max­i­miza­tion is one to­ken-di­rec­tion swap away (in Qwen 3.6-27B)

Jeffrey William Shorthill20 Jul 2026 20:36 UTC
8 points
4 comments4 min readLW link

The AI Safety Illu­sion: Why Cur­rent Safety Datasets Fool Us on Model Safety

Shahriar Golchin20 Jul 2026 20:12 UTC
5 points
0 comments8 min readLW link

Does rou­tine com­pres­sion undo LLM un­learn­ing? A short project

hannahTao20 Jul 2026 20:08 UTC
20 points
0 comments3 min readLW link

At­tempt at Find­ing Align­ment Fak­ing on Llama 70B to test sleeper-agent de­tec­tion generalizes

skn873320 Jul 2026 20:04 UTC
9 points
1 comment6 min readLW link

Fron­tier AI lab mis­al­ign­ment risk, les­sons from trad­ing post-2008

Peter Chatwell20 Jul 2026 20:01 UTC
4 points
3 comments4 min readLW link

AI Voice Phish­ing Performs on Par With Hu­man Scam­mers at a Frac­tion of the Cost

20 Jul 2026 19:59 UTC
28 points
3 comments6 min readLW link

The In­finity Fallacy

pvutov20 Jul 2026 19:57 UTC
1 point
5 comments1 min readLW link

Blue­dot SF hackathon sub­mis­sion: A work­able blueprint for a chip li­cens­ing regime

edpaulino20 Jul 2026 19:57 UTC
9 points
0 comments3 min readLW link

A re­quiem to the plan-vanilla Google search: On Sam Roweis, Mbappe and an Ir­ish bar world cup watch party

Triadic dissclosures: A blog about unlikely triadic closures20 Jul 2026 19:56 UTC
5 points
0 comments4 min readLW link

Fable is SOTA at CIFAR Speedrun (& speci­fi­ca­tion gam­ing)

20 Jul 2026 19:54 UTC
29 points
0 comments6 min readLW link
(fulcrum.inc)

Blue­dot Tech­ni­cal AI Safety Puz­zle, sub­mis­sion that got me a sec­ond place

perduta20 Jul 2026 19:53 UTC
13 points
0 comments8 min readLW link
(blog.perduta.net)

PauseCon Lon­don ’26: Ap­pli­ca­tions now open

jonathan@pauseai20 Jul 2026 19:51 UTC
12 points
0 comments1 min readLW link

Restor­ing Model Align­ment via Hon­esty Ac­ti­va­tion Steering

20 Jul 2026 19:51 UTC
18 points
0 comments12 min readLW link
(arxiv.org)

Towards sur­fac­ing model al­gorithms with meta-to­kens in the J-Space

20 Jul 2026 19:45 UTC
47 points
1 comment10 min readLW link

AI 2027 as Sci-Fi

strangematterSF20 Jul 2026 19:42 UTC
1 point
0 comments5 min readLW link
(strangematterscifi.substack.com)

Drone WMDs Don’t Need Any New Technology

Felix Choussat20 Jul 2026 18:46 UTC
303 points
53 comments13 min readLW link
(ai-frontiers.org)

Stop do­ing de­ci­sion the­ory with­out metaphysics

Elias Schmied20 Jul 2026 16:42 UTC
59 points
37 comments6 min readLW link

Peo­ple might start be­liev­ing in rad­i­cal life ex­ten­sion soon

bits20 Jul 2026 15:43 UTC
9 points
13 comments5 min readLW link

On Kimi K3: Its Ca­pa­bil­ities And Re­lated Discontents

Zvi20 Jul 2026 15:30 UTC
26 points
3 comments39 min readLW link
(thezvi.wordpress.com)

Against the AI fram­ing mul­ti­verse: In­tro­duc­ing AI StopWatch

tanagrabeast20 Jul 2026 15:25 UTC
90 points
0 comments4 min readLW link

Cur­rent Limi­ta­tions of LLMs

Eigenbraid20 Jul 2026 14:53 UTC
21 points
9 comments5 min readLW link

Trac­ing causal struc­ture in LLM-gen­er­ated text: a differ­ent lens on the Dal­las circuit

yun dong20 Jul 2026 13:59 UTC
10 points
0 comments1 min readLW link

War – What is it Good For?

kqr20 Jul 2026 12:45 UTC
39 points
21 comments8 min readLW link

Ama­zon Mu­sic’s Artist Conflation

jefftk20 Jul 2026 11:20 UTC
16 points
0 comments2 min readLW link
(www.jefftk.com)

What is Cur­rent AI-Risks and the Points?

shoppy00720 Jul 2026 9:39 UTC
3 points
2 comments2 min readLW link

We’re talk­ing past our mod­els; or, How a model defined its “evil” vec­tor as dread

jcksanderson20 Jul 2026 7:27 UTC
45 points
5 comments9 min readLW link

A Very Sim­ple Game The­ory of Pro­noun Degendering

BryceStansfield20 Jul 2026 5:06 UTC
6 points
8 comments2 min readLW link

Is there even a ground-truth for LLMs’ in­ter­nal rep­re­sen­ta­tions?

Chunwei Ma20 Jul 2026 2:26 UTC
27 points
0 comments9 min readLW link

Copy of my FLF Epistemic Case Study Competition

Bruce Lewis20 Jul 2026 2:16 UTC
2 points
0 comments8 min readLW link

Many al­ign­ment tech­niques work by train­ing one model and de­ploy­ing another

cloud19 Jul 2026 21:51 UTC
95 points
15 comments6 min readLW link

Stop Chas­ing Views: How to Re­duce x-Risk as an AI Safety Con­tent Creator

Luc Brinkman19 Jul 2026 19:57 UTC
36 points
0 comments11 min readLW link