Eng­ineer­ing the Gen­er­al­i­sa­tion Land­scape of LLMs

Samuel Ratnam13 Jul 2026 23:46 UTC
64 points
2 comments4 min readLW link

A short sum­mary of AI 2040: Plan A

Harjas13 Jul 2026 23:07 UTC
13 points
3 comments4 min readLW link
(hardlyworking1.substack.com)

[AI 2040] Trans­parency Plan

Thomas Larsen13 Jul 2026 21:24 UTC
17 points
0 comments16 min readLW link
(ai-2040.com)

Bet­ter Call Sol The Workhorse

Zvi13 Jul 2026 20:52 UTC
39 points
3 comments26 min readLW link
(thezvi.wordpress.com)

Can AI be Con­scious in Ohio?

13 Jul 2026 19:57 UTC
14 points
0 comments4 min readLW link
(papers.ssrn.com)

Paus­ing AI at hu­man level seems harder than paus­ing ASAP

MichaelDickens13 Jul 2026 17:20 UTC
73 points
0 comments2 min readLW link

Over­sight of au­to­mated re­search via sum­mari­sa­tion: a toy model

13 Jul 2026 17:06 UTC
21 points
1 comment14 min readLW link

Prism: Au­tomat­ing Science-of-Evals Research

LAThomson13 Jul 2026 16:30 UTC
47 points
0 comments12 min readLW link

The Flood, by An­ton Leicht

Austin Chen13 Jul 2026 16:23 UTC
38 points
2 comments15 min readLW link
(writing.antonleicht.me)

Start­ing The Se­quences: Some brief notes from the pref­ace and the in­tro­duc­tion

manueldelrio13 Jul 2026 16:20 UTC
12 points
0 comments2 min readLW link

The Whit­ney Bien­nial Should Ad­mit That Em­i­lie Gos­si­aux Wants to Fuck Their Dog

jenn13 Jul 2026 15:53 UTC
159 points
20 comments9 min readLW link
(jenn.site)

Lin­ear Probes add lit­tle for Ver­ifi­able Re­ward Hacking

Chandram Dutta13 Jul 2026 13:57 UTC
7 points
0 comments11 min readLW link
(onlychan.xyz)

It’s 2030 and we fucked up. How did it hap­pen?

Boaz Barak13 Jul 2026 13:34 UTC
61 points
9 comments21 min readLW link

The LLM Revolu­tion (so far)

Eigenbraid13 Jul 2026 12:19 UTC
17 points
0 comments3 min readLW link

An Epistemic Au­dit for Ex­is­ten­tial Risks from AI

Alexander Müller13 Jul 2026 9:35 UTC
6 points
1 comment11 min readLW link
(alexandermullerakm.substack.com)

5 “Plan A” scenarios

Dave Orr13 Jul 2026 0:36 UTC
66 points
2 comments4 min readLW link

The US Govern­ment may find it difficult to seize con­trol dur­ing takeoff

RobertM12 Jul 2026 22:58 UTC
42 points
6 comments2 min readLW link

One-Pager Brief on Pan­gram Labs

Sheikh Abdur Raheem Ali12 Jul 2026 18:36 UTC
48 points
4 comments2 min readLW link

Ex­tinc­tion risk is not the right first sentence

Michael Wilkinson12 Jul 2026 18:34 UTC
2 points
0 comments20 min readLW link
(michaelewilkinson.substack.com)

In­de­pen­dent al­ign­ment of lan­guage models

Michele Campolo12 Jul 2026 17:32 UTC
−9 points
2 comments38 min readLW link

From wan­tons to moral agents

Michele Campolo12 Jul 2026 17:30 UTC
3 points
0 comments17 min readLW link

The Banal­ity of Takeoff

Ihor Kendiukhov12 Jul 2026 13:45 UTC
33 points
18 comments3 min readLW link

The Con­ser­va­tion Ethic in AI 2040

cdt12 Jul 2026 13:01 UTC
17 points
9 comments3 min readLW link

Can Fron­tier Models Au­to­com­plete Safety Re­search?

12 Jul 2026 10:28 UTC
20 points
4 comments22 min readLW link
(djroytburg.github.io)

Easy Whole Set Dances With a Hook

jefftk12 Jul 2026 2:22 UTC
10 points
1 comment1 min readLW link
(www.jefftk.com)

KISS AI Safety

atlasaligned12 Jul 2026 0:47 UTC
19 points
1 comment2 min readLW link

The cur­rent bot­tle­neck is poli­ti­cal will, not research

Charbel-Raphaël11 Jul 2026 21:56 UTC
318 points
41 comments25 min readLW link

In­tro­duc­tion for and Re­ac­tions to Plan A

Zvi11 Jul 2026 20:42 UTC
36 points
9 comments33 min readLW link
(thezvi.wordpress.com)

The­o­ries of Deep Learning

astle dsa11 Jul 2026 17:21 UTC
5 points
0 comments7 min readLW link

Mea­sur­ing Is Not Enough Anymore

Lennart Finke11 Jul 2026 3:21 UTC
11 points
1 comment5 min readLW link
(fi-le.net)

Don’t bring an AI de­tec­tor to a deep­fake fight: prov­ing re­al­ity through mul­ti­modal provenance

Julien Despois11 Jul 2026 1:58 UTC
10 points
0 comments9 min readLW link

A Sim­ple Model of AI “Psy­chosis”

Adele Lopez11 Jul 2026 1:53 UTC
42 points
1 comment6 min readLW link

The Ter­mi­na­tion Cir­cuit (how rea­son­ing mod­els stop think­ing).

Chandram Dutta11 Jul 2026 1:37 UTC
12 points
0 comments3 min readLW link
(onlychan.xyz)

Notes on Tony Parkes’ “Con­tra Dance Cal­ling”

jefftk11 Jul 2026 0:50 UTC
10 points
0 comments4 min readLW link
(www.jefftk.com)

Plan A: Sugges­tions For Fur­ther Work

Thomas Larsen10 Jul 2026 23:02 UTC
52 points
7 comments7 min readLW link

Free­ing Thucydides

djbinder10 Jul 2026 22:12 UTC
44 points
2 comments4 min readLW link
(defensesindepth.bio)

An In­duc­tion Head in Dis­guise: Chas­ing Gram­mar in a Char­ac­ter-Level Transformer

Ameya Panchal10 Jul 2026 22:05 UTC
−1 points
0 comments5 min readLW link
(ameya-bit.github.io)

The stan­dard can be to ad­mit the ex­is­tence of a standard

Firinn10 Jul 2026 21:22 UTC
44 points
1 comment36 min readLW link

Per­sona Car­tog­ra­phy: Chart­ing Lan­guage Model Per­son­al­ity Traits in Weight Space

10 Jul 2026 18:54 UTC
44 points
0 comments18 min readLW link
(arxiv.org)

The eas­iest path­way to con­trol is through ex­ec­u­tive power

djbinder10 Jul 2026 18:48 UTC
111 points
3 comments6 min readLW link
(defensesindepth.bio)

The Hu­man Sub­sti­tu­tion Test as a San­ity Check for AI Evaluations

10 Jul 2026 17:27 UTC
31 points
5 comments8 min readLW link
(limits-of-evaluation.org)

Cap-and-trade ques­tion: AI-2040

kapedalex10 Jul 2026 16:54 UTC
0 points
2 comments4 min readLW link

Does AI rea­son­ing im­ply re­spon­si­bil­ity?

Hippocleides10 Jul 2026 16:44 UTC
1 point
3 comments2 min readLW link

Plan A’s prob­lem with dry tinder

Tom Davidson10 Jul 2026 15:28 UTC
81 points
6 comments8 min readLW link

AI #176 Part 2: Plan B

Zvi10 Jul 2026 12:40 UTC
26 points
1 comment36 min readLW link
(thezvi.wordpress.com)

Beliefs and po­si­tion mid 2026

RussellThor10 Jul 2026 11:43 UTC
7 points
2 comments6 min readLW link

Models of So­ciety Are Built on Models of Agents

Jonas Hallgren10 Jul 2026 9:08 UTC
18 points
0 comments10 min readLW link
(equilibria1.substack.com)

Value gen­er­al­i­sa­tion: value correction

Stuart_Armstrong10 Jul 2026 7:56 UTC
25 points
3 comments6 min readLW link

Don’t nor­mal­ize a per­ma­nent un­der­class (even a rich one)

hadad10 Jul 2026 6:40 UTC
39 points
7 comments5 min readLW link

How ro­bust are nat­u­ral lan­guage au­toen­coders to ini­tial­iza­tion?

10 Jul 2026 0:40 UTC
82 points
3 comments13 min readLW link
(turntrout.com)