Nat­u­ral Lan­guage Au­toen­coders are sum­ma­riz­ers, but do they have to be?

Andrey Anurin9 Jul 2026 19:54 UTC
10 points
1 comment22 min readLW link

Where Do LLM Values Come From?

9 Jul 2026 19:54 UTC
17 points
1 comment18 min readLW link

Selec­tive Op­ti­mism: a cri­tique of AI 2040

Richard_Ngo9 Jul 2026 19:43 UTC
231 points
13 comments8 min readLW link
(www.mindthefuture.info)

Rogue ASI Can’t Stay Aligned to Itself

absenteewarlord9 Jul 2026 19:38 UTC
8 points
2 comments1 min readLW link

Crit­i­cism against “un­em­bed­ded FDT” doesn’t ap­ply to FDT

Fernand09 Jul 2026 18:35 UTC
2 points
5 comments3 min readLW link

Your Prompt-In­jec­tion Defense Met­ric Might Be Ly­ing to You

sahilraut9 Jul 2026 17:53 UTC
5 points
0 comments8 min readLW link

When is mis­al­ign­ment just a bug?

Yoav Hollander9 Jul 2026 17:07 UTC
16 points
5 comments11 min readLW link
(blog.foretellix.com)

What is the com­pu­ta­tional sub­stance of the ax­iom of choice?

tailcalled9 Jul 2026 16:53 UTC
26 points
2 comments6 min readLW link

Skep­ti­cal of the TESCREAL Acronym? Read This.

philosophytorres9 Jul 2026 16:27 UTC
−41 points
12 comments20 min readLW link

AI 2040: Plan A

9 Jul 2026 16:25 UTC
568 points
135 comments1 min readLW link
(www.ai-2040.com)

How big is the Sun? How could you figure it out?

Elliott Thornley9 Jul 2026 16:24 UTC
101 points
9 comments6 min readLW link

Some Thoughts on The En­vi­ron­ment Prob­lem in Agent Training

TheVinci9 Jul 2026 15:41 UTC
10 points
2 comments3 min readLW link
(www.tarantulabs.com)

The Cube The­ory of Par­tially Grasped Concepts

Mateusz Bagiński9 Jul 2026 15:40 UTC
24 points
2 comments5 min readLW link

De­bate with Self-Play Best-of-N Optimization

9 Jul 2026 15:29 UTC
53 points
2 comments14 min readLW link

AI #176 Part 1: Do­ing It Live

Zvi9 Jul 2026 14:31 UTC
36 points
2 comments31 min readLW link
(thezvi.wordpress.com)

Op­ti­miser Choice Can Am­plify or Sup­press Emer­gent Misalignment

9 Jul 2026 10:00 UTC
63 points
2 comments4 min readLW link

An­nounc­ing our $160M grant from Coeffi­cient Giving

9 Jul 2026 7:58 UTC
50 points
0 comments2 min readLW link

In­ter­pretabil­ity is be­com­ing in­creas­ingly uninterpretable

jcksanderson9 Jul 2026 7:04 UTC
34 points
1 comment4 min readLW link
(jcksanderson.com)

Per­sis­tent La­tent Misal­ign­ment, a new di­men­sion of mis­al­ign­ment?

Florian_Dietz9 Jul 2026 4:56 UTC
15 points
2 comments1 min readLW link

Be­cause 8 ≈ e², An­thropic’s re­searcher up­lift is plau­si­bly >2x

Thomas Kwa9 Jul 2026 4:30 UTC
54 points
9 comments12 min readLW link
(metr.org)

Trans­form­ers Re­sist Their Own Architecture

Zach Baker9 Jul 2026 0:57 UTC
11 points
0 comments15 min readLW link

Mo­du­lar Pre­train­ing En­ables Ac­cess Control

9 Jul 2026 0:03 UTC
71 points
2 comments9 min readLW link
(alignment.anthropic.com)

There Should Be More AI Safety Hubs

Seth Lifland8 Jul 2026 23:09 UTC
13 points
1 comment3 min readLW link

Models are blind out­side the J-space. NLAs aren’t.

Pranav Viswanath8 Jul 2026 23:09 UTC
28 points
4 comments9 min readLW link

Solv­ing the BlueDot Puz­zle TAIS: The Ve­loc­ity Ring

Karine Levonyan8 Jul 2026 23:03 UTC
8 points
0 comments4 min readLW link
(karinelevonyan.github.io)

Ac­tion as Choice Ex­pressed Through Move­ment Toward a Goal: a Frame­work for Over­com­ing In­ac­tion

techandsundry8 Jul 2026 23:02 UTC
7 points
0 comments4 min readLW link

Op­ti­mum num­ber of items to in­spect be­fore buy­ing one

bilibili8 Jul 2026 21:23 UTC
18 points
0 comments4 min readLW link

Free will as a model parameter

darshanav8 Jul 2026 21:11 UTC
10 points
1 comment6 min readLW link

Find fund­ing, fast

Austin Chen8 Jul 2026 21:10 UTC
43 points
0 comments3 min readLW link
(manifund.substack.com)

Child­hood and Ed­u­ca­tion #20: Phones and Screens

Zvi8 Jul 2026 19:41 UTC
37 points
1 comment17 min readLW link
(thezvi.wordpress.com)

AI Safety Hong Kong Read­ing Group: Ma­chine Learn­ing and Hu­man Values

Schizoid Rentoid8 Jul 2026 19:19 UTC
1 point
0 comments1 min readLW link

Can the U.S. and China Deny AI?

8 Jul 2026 18:45 UTC
20 points
0 comments14 min readLW link
(secondstrike.substack.com)

A Real-Life Ex­am­ple of an Aligned Sys­tem Killing Hun­dreds of People

London L.8 Jul 2026 18:43 UTC
3 points
8 comments4 min readLW link

Notes on tech­ni­cal al­ign­ment via hu­man-like so­cial drives

Steven Byrnes8 Jul 2026 18:30 UTC
74 points
17 comments42 min readLW link

Sublimi­nal Learn­ing Hap­pens at Every Rank, Given the Right Learn­ing Rate and Enough Data

Lawrence Feng8 Jul 2026 17:32 UTC
60 points
5 comments7 min readLW link

AI Safety at the Fron­tier: Paper High­lights of May & June 2026

gasteigerjo8 Jul 2026 17:19 UTC
16 points
0 comments10 min readLW link

Why study proto-train­ing gam­ing as an ad­ver­sar­ial al­ign­ment failure mode?

8 Jul 2026 17:07 UTC
55 points
0 comments7 min readLW link

Why study al­ign­ment in­ter­ven­tions on pre-RL check­points?

8 Jul 2026 17:07 UTC
66 points
2 comments6 min readLW link

Ar­tifi­cial Am­bas­sadors: The Three Cases for Send­ing LLMs to Deep Space

Ekaterina Matson8 Jul 2026 16:08 UTC
7 points
0 comments12 min readLW link
(medium.com)

Refram­ing LessWrong-style de­ci­sion the­ory as “com­mit­ment the­ory”

Elias Schmied8 Jul 2026 15:25 UTC
61 points
33 comments11 min readLW link

Why I’m a moral anti-re­al­ist but may be un­able to con­vince you

Nina Panickssery8 Jul 2026 15:23 UTC
23 points
12 comments7 min readLW link
(blog.ninapanickssery.com)

AI Fore­cast­ing in 2026: What 11 Analy­ses Say

Ben Wilson8 Jul 2026 14:39 UTC
25 points
0 comments18 min readLW link
(www.metaculus.com)

Why mod­els game evals might mat­ter as much as whether they do it

Igor Ivanov8 Jul 2026 14:30 UTC
30 points
8 comments4 min readLW link

The Em­pa­thetic Poor

a unemployed pastor- de S Brito Gabriel8 Jul 2026 14:21 UTC
5 points
2 comments16 min readLW link

How Valuable are Tra­di­tions?

martinkunev8 Jul 2026 13:44 UTC
10 points
5 comments5 min readLW link

How much slower does take­off go with 10× less com­pute?

brendanhalstead8 Jul 2026 10:04 UTC
46 points
4 comments4 min readLW link

Hu­man Em­pow­er­ment in an AI Society

Alvin Ånestrand8 Jul 2026 9:01 UTC
24 points
0 comments36 min readLW link

The mosquito bucket of doom works

dominicq8 Jul 2026 8:18 UTC
134 points
10 comments5 min readLW link
(blog.d11r.eu)

[Linkpost]Six Story Prompts I Want to Read

Linch8 Jul 2026 6:12 UTC
15 points
0 comments11 min readLW link
(linch.substack.com)

the al­ls­ton lab rat ex­is­ten­tial framing

somayya8 Jul 2026 2:45 UTC
−3 points
0 comments2 min readLW link