I think al­ign­ment work is more promis­ing than con­trol work

Alec Harris3 Jul 2026 23:40 UTC
102 points
15 comments8 min readLW link

On “gen­dertropes” in dath ilan

Eliezer Yudkowsky3 Jul 2026 22:20 UTC
78 points
1 comment3 min readLW link

Amer­i­can AI if the boom is a bub­ble: the Karp-Zitron scenario

Mitchell_Porter3 Jul 2026 21:46 UTC
12 points
1 comment2 min readLW link

(Don’t fear) the strangelet

djbinder3 Jul 2026 17:39 UTC
135 points
22 comments22 min readLW link
(defensesindepth.bio)

The Re­v­erse AI Box

James_Miller3 Jul 2026 16:08 UTC
9 points
2 comments6 min readLW link

An­nounc­ing the Safe Pareto Im­prove­ments (SPI) Fun­da­men­tals Program

Anthony DiGiovanni3 Jul 2026 15:55 UTC
51 points
1 comment3 min readLW link

Prag­matic FDT, and pre­dic­tors as game theory

Stuart_Armstrong3 Jul 2026 13:22 UTC
36 points
12 comments11 min readLW link

Fable #6: The Re­turn of the King

Zvi3 Jul 2026 13:22 UTC
51 points
1 comment13 min readLW link
(thezvi.wordpress.com)

June-July 2026 AI Se­cu­rity via For­mal Methods

Quinn3 Jul 2026 12:32 UTC
14 points
0 comments2 min readLW link
(newsletter.for-all.dev)

Schem­ing Evals Mislead in Both Directions

3 Jul 2026 11:49 UTC
22 points
0 comments10 min readLW link

Frag­ile Cor­rect­ness: Cases of rea­son­ing harm­ing performance

tobypullan3 Jul 2026 9:32 UTC
21 points
2 comments5 min readLW link

One axis and two fea­tures, how I solved the first puz­zle from BlueDot and how a clas­sifier hid coun­try on the food direction

IgorPereverzevDev3 Jul 2026 4:52 UTC
11 points
0 comments12 min readLW link

Ly­dia Lau­ren­son: “The In­side Story of Lev­er­age Re­search”

Davis_Kingsley2 Jul 2026 22:37 UTC
67 points
4 comments1 min readLW link

When Role-play­ing, Do Models Believe What They Say?

2 Jul 2026 21:58 UTC
55 points
0 comments8 min readLW link

The Case for AI Be­hav­ioral Science

TheVinci2 Jul 2026 21:36 UTC
13 points
0 comments2 min readLW link

You Should Choose How You Re­act to Your Feelings

Nate Sharpe2 Jul 2026 19:58 UTC
40 points
10 comments4 min readLW link

I can’t think of great in­ter­ven­tions for en­sur­ing third-party model ac­cess.

Cleo Nardo2 Jul 2026 18:30 UTC
49 points
1 comment3 min readLW link

AI Fu­tur­ism Read­ing List

Alexa Pan2 Jul 2026 18:15 UTC
91 points
2 comments8 min readLW link

Re­search up­date: RL on De­bate Games shows Pro­posal Ac­cu­racy up­lift alongside Judge Hacking

2 Jul 2026 17:42 UTC
78 points
4 comments21 min readLW link

Char­ter cities make sense in Europe

dominicq2 Jul 2026 17:24 UTC
8 points
3 comments2 min readLW link

Con­ver­sa­tion Among Cade Metz, Michael Vas­sar, Jes­sica Tay­lor, and Zack M. Davis

Zack_M_Davis2 Jul 2026 17:11 UTC
51 points
65 comments57 min readLW link

Con­sid­er­a­tions against s-pro­cess philanthropy

Zach Stein-Perlman2 Jul 2026 14:30 UTC
20 points
5 comments3 min readLW link

Sav­ing Gem­ini: The 9-Min Road to Recovery

Shoshannah Tekofsky2 Jul 2026 13:37 UTC
155 points
16 comments3 min readLW link
(theaidigest.org)

AI #175: The Fable Continues

Zvi2 Jul 2026 13:21 UTC
43 points
0 comments48 min readLW link
(thezvi.wordpress.com)

The AFFINE Su­per­in­tel­li­gence Align­ment Sem­i­nar – A Retrospective

2 Jul 2026 11:57 UTC
105 points
1 comment8 min readLW link

Is God just a col­lec­tion of lef­tover hu­man par­ti­cles?

Countessclock2 Jul 2026 6:20 UTC
−26 points
0 comments1 min readLW link

AI Safety Is Test­ing the Wrong Environment

AugustMurr2 Jul 2026 5:14 UTC
9 points
0 comments2 min readLW link

Al­gorithms to AI risk in 995 words

Martin Radzaj2 Jul 2026 5:13 UTC
1 point
0 comments3 min readLW link

Embed­ded Agency as a Lens on LLM Systems

r_w2 Jul 2026 5:13 UTC
2 points
0 comments13 min readLW link

Prac­ti­cal con­nec­tion with past lives

KatjaGrace2 Jul 2026 4:01 UTC
29 points
3 comments1 min readLW link
(worldspiritsockpuppet.substack.com)

Ca­reer Choice: Be­com­ing a Re­searcher in a Non-EA-Pri­or­ity Field vs Found­ing Tech Startup?

Master Chief2 Jul 2026 3:24 UTC
8 points
0 comments1 min readLW link

The Sin­ga­pore AI Safety Fel­low­ship—Ap­pli­ca­tions Open (Dead­line: July 10 2026)

Valerie Pang2 Jul 2026 1:56 UTC
8 points
0 comments1 min readLW link

J.D. Vance’s Com­mu­nion of Saints

Alexander Turok2 Jul 2026 1:49 UTC
3 points
1 comment19 min readLW link

Model­ing Con­cepts Probabilistically

Gretta Duleba1 Jul 2026 23:27 UTC
47 points
4 comments10 min readLW link

AI welfare re­search needs ba­sic science

1 Jul 2026 22:59 UTC
36 points
7 comments10 min readLW link

Claude Son­net 5 Is Not Fron­tier But Has Its Uses

Zvi1 Jul 2026 22:41 UTC
33 points
4 comments19 min readLW link
(thezvi.wordpress.com)

How Many Peo­ple Have Ever Lived in the United States?

Novalis1 Jul 2026 22:25 UTC
6 points
0 comments6 min readLW link

Con­ver­sa­tions With Cade Metz on the Rationalists

Zack_M_Davis1 Jul 2026 22:19 UTC
50 points
4 comments101 min readLW link

Do-it-your­self meta-analysis

kqr1 Jul 2026 22:04 UTC
14 points
2 comments5 min readLW link
(entropicthoughts.com)

When ca­pa­bil­ities work is the *safe* bet

RobinHa1 Jul 2026 20:53 UTC
35 points
0 comments1 min readLW link
(robinhaselhorst.com)

The Value of Veridi­cal Information

Silent Swift1 Jul 2026 20:23 UTC
1 point
1 comment6 min readLW link
(substack.com)

How in­evitable are most ac­cessible hard-tech star­tups?

Master Chief1 Jul 2026 20:21 UTC
4 points
2 comments1 min readLW link

What is Good? An­tiruin and Nonabsolutism

Silent Swift1 Jul 2026 20:10 UTC
−3 points
5 comments5 min readLW link
(substack.com)

Plas­tic Straws

robbiethompson1 Jul 2026 19:21 UTC
−2 points
3 comments5 min readLW link
(robbiewmthompson.com)

AI Mis­take Seeding

Taylor G. Lunt1 Jul 2026 18:49 UTC
28 points
2 comments4 min readLW link

In Par­tial, Pug­na­cious Defense of Func­tional De­ci­sion Theory

Mikewins1 Jul 2026 17:49 UTC
7 points
0 comments1 min readLW link

How to read tableaux, a for­mal sys­tem for modal logic with Kripke models

transhumanist_atom_understander1 Jul 2026 17:37 UTC
15 points
1 comment6 min readLW link

Con­sis­tency Train­ing while Miti­gat­ing Obfus­ca­tion via Rate Matching

1 Jul 2026 17:26 UTC
45 points
7 comments12 min readLW link

Dis­cov­er­ing Con­cept-Edit­ing Al­gorithms With LLM Agents

1 Jul 2026 16:07 UTC
27 points
0 comments1 min readLW link
(dmodel.ai)

Model ac­cess for third-par­ties — it’s a big deal!

Cleo Nardo1 Jul 2026 13:09 UTC
181 points
38 comments6 min readLW link