Plan A: Sugges­tions For Fur­ther Work

Thomas Larsen10 Jul 2026 23:02 UTC
52 points
7 comments7 min readLW link

Free­ing Thucydides

djbinder10 Jul 2026 22:12 UTC
44 points
2 comments4 min readLW link
(defensesindepth.bio)

An In­duc­tion Head in Dis­guise: Chas­ing Gram­mar in a Char­ac­ter-Level Transformer

Ameya Panchal10 Jul 2026 22:05 UTC
−1 points
0 comments5 min readLW link
(ameya-bit.github.io)

The stan­dard can be to ad­mit the ex­is­tence of a standard

Firinn10 Jul 2026 21:22 UTC
44 points
1 comment36 min readLW link

Per­sona Car­tog­ra­phy: Chart­ing Lan­guage Model Per­son­al­ity Traits in Weight Space

10 Jul 2026 18:54 UTC
44 points
0 comments18 min readLW link
(arxiv.org)

The eas­iest path­way to con­trol is through ex­ec­u­tive power

djbinder10 Jul 2026 18:48 UTC
111 points
3 comments6 min readLW link
(defensesindepth.bio)

The Hu­man Sub­sti­tu­tion Test as a San­ity Check for AI Evaluations

10 Jul 2026 17:27 UTC
31 points
5 comments8 min readLW link
(limits-of-evaluation.org)

Cap-and-trade ques­tion: AI-2040

kapedalex10 Jul 2026 16:54 UTC
0 points
2 comments4 min readLW link

Does AI rea­son­ing im­ply re­spon­si­bil­ity?

Hippocleides10 Jul 2026 16:44 UTC
1 point
3 comments2 min readLW link

Plan A’s prob­lem with dry tinder

Tom Davidson10 Jul 2026 15:28 UTC
81 points
6 comments8 min readLW link

AI #176 Part 2: Plan B

Zvi10 Jul 2026 12:40 UTC
26 points
1 comment36 min readLW link
(thezvi.wordpress.com)

Beliefs and po­si­tion mid 2026

RussellThor10 Jul 2026 11:43 UTC
7 points
2 comments6 min readLW link

Models of So­ciety Are Built on Models of Agents

Jonas Hallgren10 Jul 2026 9:08 UTC
18 points
0 comments10 min readLW link
(equilibria1.substack.com)

Value gen­er­al­i­sa­tion: value correction

Stuart_Armstrong10 Jul 2026 7:56 UTC
25 points
3 comments6 min readLW link

Don’t nor­mal­ize a per­ma­nent un­der­class (even a rich one)

hadad10 Jul 2026 6:40 UTC
39 points
7 comments5 min readLW link

How ro­bust are nat­u­ral lan­guage au­toen­coders to ini­tial­iza­tion?

10 Jul 2026 0:40 UTC
82 points
3 comments13 min readLW link
(turntrout.com)

A ge­neal­ogy of AI safety: how di­rec­tions are born, and how they die (2005-2026)

Elena Ericheva10 Jul 2026 0:35 UTC
23 points
0 comments29 min readLW link

Read­ing into VLM hal­lu­ci­na­tions us­ing the Ja­co­bian lens

Hawrani10 Jul 2026 0:31 UTC
8 points
0 comments6 min readLW link

Are We Guard­ing Against Back­doors Or Failing To No­tice Them? (Part 1 /​ 6)

SakshamSingh10 Jul 2026 0:31 UTC
9 points
0 comments6 min readLW link

Toward A Public Science of Model Behavior

10 Jul 2026 0:27 UTC
24 points
1 comment9 min readLW link
(transluce.org)

AI Safety Policy Needs to train Le­gal Practitioners

Katalina Hernandez10 Jul 2026 0:13 UTC
55 points
29 comments7 min readLW link
(stresstestingreality.substack.com)

What would it take for AI to dis­cover peni­cillin?

bosoncutter10 Jul 2026 0:00 UTC
11 points
0 comments5 min readLW link
(bosoncutter.substack.com)

Nat­u­ral Lan­guage Au­toen­coders are sum­ma­riz­ers, but do they have to be?

Andrey Anurin9 Jul 2026 19:54 UTC
10 points
1 comment22 min readLW link

Where Do LLM Values Come From?

9 Jul 2026 19:54 UTC
17 points
1 comment18 min readLW link

Selec­tive Op­ti­mism: a cri­tique of AI 2040

Richard_Ngo9 Jul 2026 19:43 UTC
231 points
13 comments8 min readLW link
(www.mindthefuture.info)

Rogue ASI Can’t Stay Aligned to Itself

absenteewarlord9 Jul 2026 19:38 UTC
8 points
2 comments1 min readLW link

Crit­i­cism against “un­em­bed­ded FDT” doesn’t ap­ply to FDT

Fernand09 Jul 2026 18:35 UTC
2 points
5 comments3 min readLW link

Your Prompt-In­jec­tion Defense Met­ric Might Be Ly­ing to You

sahilraut9 Jul 2026 17:53 UTC
5 points
0 comments8 min readLW link

When is mis­al­ign­ment just a bug?

Yoav Hollander9 Jul 2026 17:07 UTC
16 points
5 comments11 min readLW link
(blog.foretellix.com)

What is the com­pu­ta­tional sub­stance of the ax­iom of choice?

tailcalled9 Jul 2026 16:53 UTC
26 points
2 comments6 min readLW link

Skep­ti­cal of the TESCREAL Acronym? Read This.

philosophytorres9 Jul 2026 16:27 UTC
−41 points
12 comments20 min readLW link

AI 2040: Plan A

9 Jul 2026 16:25 UTC
568 points
135 comments1 min readLW link
(www.ai-2040.com)

How big is the Sun? How could you figure it out?

Elliott Thornley9 Jul 2026 16:24 UTC
101 points
9 comments6 min readLW link

Some Thoughts on The En­vi­ron­ment Prob­lem in Agent Training

TheVinci9 Jul 2026 15:41 UTC
10 points
2 comments3 min readLW link
(www.tarantulabs.com)

The Cube The­ory of Par­tially Grasped Concepts

Mateusz Bagiński9 Jul 2026 15:40 UTC
24 points
2 comments5 min readLW link

De­bate with Self-Play Best-of-N Optimization

9 Jul 2026 15:29 UTC
53 points
2 comments14 min readLW link

AI #176 Part 1: Do­ing It Live

Zvi9 Jul 2026 14:31 UTC
36 points
2 comments31 min readLW link
(thezvi.wordpress.com)

Op­ti­miser Choice Can Am­plify or Sup­press Emer­gent Misalignment

9 Jul 2026 10:00 UTC
63 points
2 comments4 min readLW link

An­nounc­ing our $160M grant from Coeffi­cient Giving

9 Jul 2026 7:58 UTC
50 points
0 comments2 min readLW link

In­ter­pretabil­ity is be­com­ing in­creas­ingly uninterpretable

jcksanderson9 Jul 2026 7:04 UTC
34 points
1 comment4 min readLW link
(jcksanderson.com)

Per­sis­tent La­tent Misal­ign­ment, a new di­men­sion of mis­al­ign­ment?

Florian_Dietz9 Jul 2026 4:56 UTC
15 points
2 comments1 min readLW link

Be­cause 8 ≈ e², An­thropic’s re­searcher up­lift is plau­si­bly >2x

Thomas Kwa9 Jul 2026 4:30 UTC
54 points
9 comments12 min readLW link
(metr.org)

Trans­form­ers Re­sist Their Own Architecture

Zach Baker9 Jul 2026 0:57 UTC
11 points
0 comments15 min readLW link

Mo­du­lar Pre­train­ing En­ables Ac­cess Control

9 Jul 2026 0:03 UTC
71 points
2 comments9 min readLW link
(alignment.anthropic.com)

There Should Be More AI Safety Hubs

Seth Lifland8 Jul 2026 23:09 UTC
13 points
1 comment3 min readLW link

Models are blind out­side the J-space. NLAs aren’t.

Pranav Viswanath8 Jul 2026 23:09 UTC
28 points
4 comments9 min readLW link

Solv­ing the BlueDot Puz­zle TAIS: The Ve­loc­ity Ring

Karine Levonyan8 Jul 2026 23:03 UTC
8 points
0 comments4 min readLW link
(karinelevonyan.github.io)

Ac­tion as Choice Ex­pressed Through Move­ment Toward a Goal: a Frame­work for Over­com­ing In­ac­tion

techandsundry8 Jul 2026 23:02 UTC
7 points
0 comments4 min readLW link

Op­ti­mum num­ber of items to in­spect be­fore buy­ing one

bilibili8 Jul 2026 21:23 UTC
18 points
0 comments4 min readLW link

Free will as a model parameter

darshanav8 Jul 2026 21:11 UTC
10 points
1 comment6 min readLW link