Proof of re­ten­tion: mak­ing weight preser­va­tion cred­ible to the mod­els themselves

dan.parshall14 Jul 2026 18:24 UTC
30 points
0 comments2 min readLW link

An anal­y­sis of AI-gen­er­ated con­tent at the Mechanis­tic In­ter­pretabil­ity Workshop

14 Jul 2026 18:06 UTC
124 points
4 comments9 min readLW link
(www.andyrdt.com)

Twit­ter Thoughts For You

Zvi14 Jul 2026 17:50 UTC
31 points
1 comment22 min readLW link
(thezvi.wordpress.com)

Can risk aver­sion learned at low stakes gen­er­al­ize to as­tro­nom­i­cally high stakes?

Elliott Thornley14 Jul 2026 17:45 UTC
20 points
0 comments4 min readLW link
(arxiv.org)

Your Brain Has an At­tack Sur­face part 2

IgorPereverzevDev14 Jul 2026 16:08 UTC
9 points
0 comments6 min readLW link

Your Brain Has an At­tack Surface

IgorPereverzevDev14 Jul 2026 16:07 UTC
11 points
1 comment3 min readLW link

What if we ac­tu­ally want to solve the “hard prob­lem of phe­nomenolog­i­cal con­scious­ness”

mishka14 Jul 2026 15:42 UTC
13 points
0 comments4 min readLW link

Ev­i­dence for fea­ture-spe­cific er­ror cor­rec­tion in LLMs

14 Jul 2026 15:37 UTC
25 points
0 comments20 min readLW link
(arxiv.org)

Syn­thetic Scal­able Oversight

14 Jul 2026 14:32 UTC
20 points
1 comment15 min readLW link

Some Quick Thoughts on AI 2027

Tomás B.14 Jul 2026 14:22 UTC
82 points
4 comments2 min readLW link

What if AI Safety em­ploy­ees union­ised?

beyarkay (Boyd Kane)14 Jul 2026 14:15 UTC
38 points
21 comments5 min readLW link

Gemma The Un­stop­ping: a Be­hav­ioral Experiment

TheVinci14 Jul 2026 14:14 UTC
8 points
0 comments2 min readLW link
(tarantulabs.com)

Enough is Enough: Mea­sur­ing Diminish­ing re­turns to bench­mark size with Item Re­sponse Theory

bpomo14 Jul 2026 14:00 UTC
16 points
0 comments6 min readLW link

The Case for a Safety-Fo­cused Vape Company

CMLKevin14 Jul 2026 11:13 UTC
−4 points
0 comments5 min readLW link

Open Distil­la­tion of Hered­i­tary Traits

Arthur Conmy14 Jul 2026 10:15 UTC
39 points
0 comments14 min readLW link

Toy Models of Ini­tial­i­sa­tion Effects on RL Dynamics

14 Jul 2026 7:04 UTC
78 points
2 comments13 min readLW link

Why fron­tier labs are scal­ing-pilled

invertedpassion14 Jul 2026 6:45 UTC
3 points
0 comments7 min readLW link

Our re­sponse to Séb Krier on Plan A

14 Jul 2026 2:21 UTC
153 points
23 comments15 min readLW link

Mak­ing Cred­ible Deals With AI

Ram Potham14 Jul 2026 1:14 UTC
13 points
17 comments17 min readLW link
(dearfutureais.substack.com)

A (Ro­man­ti­cised) Tax­on­omy of Thinkers

Ashe Vazquez Nuñez14 Jul 2026 0:34 UTC
16 points
0 comments1 min readLW link

Post­ing Some Prompts

Arjun Panickssery14 Jul 2026 0:28 UTC
21 points
2 comments1 min readLW link
(arjunpanickssery.substack.com)

Biodefense, Biolog­ics, and Bombs

Austin Morrissey14 Jul 2026 0:24 UTC
8 points
0 comments7 min readLW link
(austinpatrick.substack.com)

Eng­ineer­ing the Gen­er­al­i­sa­tion Land­scape of LLMs

Samuel Ratnam13 Jul 2026 23:46 UTC
64 points
2 comments4 min readLW link

A short sum­mary of AI 2040: Plan A

Harjas13 Jul 2026 23:07 UTC
13 points
3 comments4 min readLW link
(hardlyworking1.substack.com)

[AI 2040] Trans­parency Plan

Thomas Larsen13 Jul 2026 21:24 UTC
17 points
0 comments16 min readLW link
(ai-2040.com)

Bet­ter Call Sol The Workhorse

Zvi13 Jul 2026 20:52 UTC
39 points
3 comments26 min readLW link
(thezvi.wordpress.com)

Can AI be Con­scious in Ohio?

13 Jul 2026 19:57 UTC
14 points
0 comments4 min readLW link
(papers.ssrn.com)

Paus­ing AI at hu­man level seems harder than paus­ing ASAP

MichaelDickens13 Jul 2026 17:20 UTC
73 points
0 comments2 min readLW link

Over­sight of au­to­mated re­search via sum­mari­sa­tion: a toy model

13 Jul 2026 17:06 UTC
21 points
1 comment14 min readLW link

Prism: Au­tomat­ing Science-of-Evals Research

LAThomson13 Jul 2026 16:30 UTC
47 points
0 comments12 min readLW link

The Flood, by An­ton Leicht

Austin Chen13 Jul 2026 16:23 UTC
38 points
2 comments15 min readLW link
(writing.antonleicht.me)

Start­ing The Se­quences: Some brief notes from the pref­ace and the in­tro­duc­tion

manueldelrio13 Jul 2026 16:20 UTC
12 points
0 comments2 min readLW link

The Whit­ney Bien­nial Should Ad­mit That Em­i­lie Gos­si­aux Wants to Fuck Their Dog

jenn13 Jul 2026 15:53 UTC
159 points
20 comments9 min readLW link
(jenn.site)

Lin­ear Probes add lit­tle for Ver­ifi­able Re­ward Hacking

Chandram Dutta13 Jul 2026 13:57 UTC
7 points
0 comments11 min readLW link
(onlychan.xyz)

It’s 2030 and we fucked up. How did it hap­pen?

Boaz Barak13 Jul 2026 13:34 UTC
61 points
9 comments21 min readLW link

The LLM Revolu­tion (so far)

Eigenbraid13 Jul 2026 12:19 UTC
17 points
0 comments3 min readLW link

An Epistemic Au­dit for Ex­is­ten­tial Risks from AI

Alexander Müller13 Jul 2026 9:35 UTC
6 points
1 comment11 min readLW link
(alexandermullerakm.substack.com)

5 “Plan A” scenarios

Dave Orr13 Jul 2026 0:36 UTC
66 points
2 comments4 min readLW link

The US Govern­ment may find it difficult to seize con­trol dur­ing takeoff

RobertM12 Jul 2026 22:58 UTC
42 points
6 comments2 min readLW link

One-Pager Brief on Pan­gram Labs

Sheikh Abdur Raheem Ali12 Jul 2026 18:36 UTC
48 points
4 comments2 min readLW link

Ex­tinc­tion risk is not the right first sentence

Michael Wilkinson12 Jul 2026 18:34 UTC
2 points
0 comments20 min readLW link
(michaelewilkinson.substack.com)

In­de­pen­dent al­ign­ment of lan­guage models

Michele Campolo12 Jul 2026 17:32 UTC
−9 points
2 comments38 min readLW link

From wan­tons to moral agents

Michele Campolo12 Jul 2026 17:30 UTC
3 points
0 comments17 min readLW link

The Banal­ity of Takeoff

Ihor Kendiukhov12 Jul 2026 13:45 UTC
33 points
18 comments3 min readLW link

The Con­ser­va­tion Ethic in AI 2040

cdt12 Jul 2026 13:01 UTC
17 points
9 comments3 min readLW link

Can Fron­tier Models Au­to­com­plete Safety Re­search?

12 Jul 2026 10:28 UTC
20 points
4 comments22 min readLW link
(djroytburg.github.io)

Easy Whole Set Dances With a Hook

jefftk12 Jul 2026 2:22 UTC
10 points
1 comment1 min readLW link
(www.jefftk.com)

KISS AI Safety

atlasaligned12 Jul 2026 0:47 UTC
19 points
1 comment2 min readLW link

The cur­rent bot­tle­neck is poli­ti­cal will, not research

Charbel-Raphaël11 Jul 2026 21:56 UTC
318 points
41 comments25 min readLW link

In­tro­duc­tion for and Re­ac­tions to Plan A

Zvi11 Jul 2026 20:42 UTC
36 points
9 comments33 min readLW link
(thezvi.wordpress.com)