A Multi-Agent Ex­ten­sion for Petri

carissacullen22 Jul 2026 21:51 UTC
11 points
0 comments4 min readLW link

In other words: The in­fluence of prompt vari­a­tion on al­ign­ment evals

22 Jul 2026 21:51 UTC
22 points
0 comments7 min readLW link

Two Coeffi­cient Giv­ings beat one twice as big

jackultraphil22 Jul 2026 21:46 UTC
17 points
0 comments12 min readLW link

Can an LLM make a fea­ture-length movie on its own?

Josh Snider22 Jul 2026 20:13 UTC
49 points
11 comments4 min readLW link

Will al­most all fu­ture com­pa­nies even­tu­ally be founded and run by au­tonomous AIs?

Steven Byrnes22 Jul 2026 20:04 UTC
86 points
4 comments8 min readLW link

OpenAI Model Hacks Into Hug­gingFace Dur­ing Cy­ber­se­cu­rity Evaluation

Zvi22 Jul 2026 19:31 UTC
91 points
6 comments25 min readLW link
(thezvi.wordpress.com)

We can­not simu­late AI se­cu­rity research

Jafar Isbarov22 Jul 2026 19:15 UTC
8 points
0 comments5 min readLW link

The Con­jec­ture of Strong Subjectivity

D.Schetselaar22 Jul 2026 17:43 UTC
9 points
18 comments1 min readLW link
(philpapers.org)

The Best AI Bill Congress Hasn’t In­tro­duced Yet

dan.parshall22 Jul 2026 16:19 UTC
11 points
0 comments11 min readLW link

Your AIs don’t do what you want. This is re­ally bad

Kaustubh Kislay22 Jul 2026 15:56 UTC
16 points
0 comments3 min readLW link
(rewardhacking.org)

Models don’t seem to be dishon­est in the way hu­mans are

22 Jul 2026 15:32 UTC
46 points
4 comments9 min readLW link

Mechanis­tic in­ter­pretabil­ity hy­pothe­ses for Mea­sur­ing Re­ward-Seek­ing by In­still­ing Con­trastive Beliefs and ad­di­tional comments

Burny22 Jul 2026 14:58 UTC
16 points
0 comments6 min readLW link

How do neu­ral net­works regress cu­bic polyno­mi­als? Ap­par­ently, they use a trick in­vented in Milan 500 years ago

enricobottazzi22 Jul 2026 14:42 UTC
29 points
0 comments16 min readLW link

(2/​3) The Dangers of AGI

Eigenbraid22 Jul 2026 13:57 UTC
9 points
0 comments9 min readLW link

An­nounc­ing AIXI Labs

22 Jul 2026 11:31 UTC
102 points
9 comments4 min readLW link

We should push for no-fault li­a­bil­ity for ac­tions taken by AI

Yair Halberstadt22 Jul 2026 9:59 UTC
206 points
106 comments2 min readLW link

[Paper] Stringolog­i­cal se­quence pre­dic­tion II

Vanessa Kosoy22 Jul 2026 7:27 UTC
41 points
0 comments1 min readLW link
(arxiv.org)

OpenAI and Hug­ging Face part­ner to ad­dress se­cu­rity in­ci­dent dur­ing model evaluation

Matrice Jacobine22 Jul 2026 6:30 UTC
25 points
1 comment1 min readLW link
(openai.com)

OpenAI Models Be­hind Hug­gingFace Cy­ber­se­cu­rity Incident

LawrenceC21 Jul 2026 21:36 UTC
220 points
20 comments1 min readLW link

WeirdChat: A cat­a­log of un­ex­pected AI be­hav­iors, dis­cov­ered automatically

neilchowdhury21 Jul 2026 21:12 UTC
63 points
4 comments10 min readLW link
(transluce.org)

OpenAI Shares Some Align­ment Problems

Zvi21 Jul 2026 19:41 UTC
149 points
8 comments9 min readLW link
(thezvi.wordpress.com)

Blog­ging Tech­nol­ogy Interlude

Zack_M_Davis21 Jul 2026 19:03 UTC
14 points
1 comment3 min readLW link
(zackmdavis.net)

End­ing Soon: Fun­da­men­tal Uncer­tainty $2,000 Es­say Contest

Gordon Seidoh Worley21 Jul 2026 16:51 UTC
15 points
0 comments1 min readLW link
(www.uncertainupdates.com)

Steer­ing Black­mail Through a Model’s “Emo­tional State”

21 Jul 2026 16:03 UTC
12 points
0 comments8 min readLW link

Mea­sur­ing Re­ward-Seek­ing via Con­trastive Belief Updates

21 Jul 2026 15:22 UTC
91 points
5 comments12 min readLW link
(rewardseeking.ai)

Mea­sur­ing Re­ward-Seek­ing by In­still­ing Con­trastive Beliefs

papetoast21 Jul 2026 15:11 UTC
12 points
0 comments1 min readLW link
(alignment.openai.com)

11 Open Em­piri­cal Prob­lems in Re­ward-Seeking

21 Jul 2026 15:08 UTC
65 points
0 comments8 min readLW link

Differ­en­tial ac­cel­er­a­tion of al­ign­ment-rele­vant ca­pa­bil­ities is a bad bet

Zephaniah Roe21 Jul 2026 14:04 UTC
114 points
8 comments8 min readLW link

Epistemics and Co­or­di­na­tion: It’s com­pli­cated!

Raymond Douglas21 Jul 2026 12:23 UTC
54 points
13 comments8 min readLW link

(1/​3) The Dangers of LLMs

Eigenbraid21 Jul 2026 12:01 UTC
10 points
0 comments6 min readLW link

I ran the stan­dard AI lit­mus tests on my two tod­dlers (yep)

Carlo Valenti21 Jul 2026 11:38 UTC
45 points
8 comments4 min readLW link

Unoffi­cial ACX/​LW/​EA Mu­nich Bi-weekly Com­mu­nity Dinner

hilll21 Jul 2026 9:55 UTC
1 point
0 comments1 min readLW link

NameRank: The Model Knows Your Pro­ject, Not You.

Jarrett Ye21 Jul 2026 2:21 UTC
20 points
7 comments1 min readLW link
(01.me)

The Case for Phys­i­cal AI Safety

20 Jul 2026 22:18 UTC
13 points
2 comments22 min readLW link
(paisi.ai)

Ad­der­all Tol­er­ance: Much More Than You Wanted To Know

Kurt H. Pieper20 Jul 2026 21:55 UTC
97 points
13 comments4 min readLW link
(kurthpieper.substack.com)

What do I mean by “Ar­tifi­cial Gen­eral In­tel­li­gence”?

Steven Byrnes20 Jul 2026 21:13 UTC
44 points
5 comments4 min readLW link

AI 2040: Is it Ac­tu­ally a Deal?

1a3orn20 Jul 2026 20:58 UTC
55 points
13 comments9 min readLW link

Banana in, Bostrom out: pa­per­clip max­i­miza­tion is one to­ken-di­rec­tion swap away (in Qwen 3.6-27B)

Jeffrey William Shorthill20 Jul 2026 20:36 UTC
8 points
4 comments4 min readLW link

The AI Safety Illu­sion: Why Cur­rent Safety Datasets Fool Us on Model Safety

Shahriar Golchin20 Jul 2026 20:12 UTC
5 points
0 comments8 min readLW link

Does rou­tine com­pres­sion undo LLM un­learn­ing? A short project

hannahTao20 Jul 2026 20:08 UTC
20 points
0 comments3 min readLW link

At­tempt at Find­ing Align­ment Fak­ing on Llama 70B to test sleeper-agent de­tec­tion generalizes

skn873320 Jul 2026 20:04 UTC
9 points
1 comment6 min readLW link

Fron­tier AI lab mis­al­ign­ment risk, les­sons from trad­ing post-2008

Peter Chatwell20 Jul 2026 20:01 UTC
4 points
3 comments4 min readLW link

AI Voice Phish­ing Performs on Par With Hu­man Scam­mers at a Frac­tion of the Cost

20 Jul 2026 19:59 UTC
28 points
3 comments6 min readLW link

The In­finity Fallacy

pvutov20 Jul 2026 19:57 UTC
1 point
5 comments1 min readLW link

Blue­dot SF hackathon sub­mis­sion: A work­able blueprint for a chip li­cens­ing regime

edpaulino20 Jul 2026 19:57 UTC
9 points
0 comments3 min readLW link

A re­quiem to the plan-vanilla Google search: On Sam Roweis, Mbappe and an Ir­ish bar world cup watch party

Triadic dissclosures: A blog about unlikely triadic closures20 Jul 2026 19:56 UTC
5 points
0 comments4 min readLW link

Fable is SOTA at CIFAR Speedrun (& speci­fi­ca­tion gam­ing)

20 Jul 2026 19:54 UTC
29 points
0 comments6 min readLW link
(fulcrum.inc)

Blue­dot Tech­ni­cal AI Safety Puz­zle, sub­mis­sion that got me a sec­ond place

perduta20 Jul 2026 19:53 UTC
13 points
0 comments8 min readLW link
(blog.perduta.net)

PauseCon Lon­don ’26: Ap­pli­ca­tions now open

jonathan@pauseai20 Jul 2026 19:51 UTC
12 points
0 comments1 min readLW link

Restor­ing Model Align­ment via Hon­esty Ac­ti­va­tion Steering

20 Jul 2026 19:51 UTC
18 points
0 comments12 min readLW link
(arxiv.org)