Con­sti­tu­tional Mid­train­ing: Con­tent Pres­ence Drives Align­ment Gains

desireecho1 Aug 2026 23:52 UTC
11 points
0 comments7 min readLW link

Ex­is­ten­tial Risk from AI: An Ex­po­si­tion for Mathematicians

alkjash1 Aug 2026 23:16 UTC
98 points
25 comments1 min readLW link
(alkjash.github.io)

RLVR that re­wards red team­ing the train­ing environment

Fiora Starlight1 Aug 2026 23:07 UTC
101 points
8 comments5 min readLW link

Math­e­mat­i­ci­ans may be wor­ried, but AI-for-sci­ence is go­ing to be great, re­cur­sively self-im­prov­ing, and we’re go­ing to learn loads

Simon DeDeo1 Aug 2026 21:59 UTC
11 points
2 comments6 min readLW link

Bayeswatch: a Retrospective

lsusr1 Aug 2026 20:13 UTC
47 points
0 comments4 min readLW link

Con­firm­ing Claims of Su­per­po­si­tion and Ad­ver­sar­ial Ex­am­ples in Toy Models

1 Aug 2026 20:09 UTC
28 points
0 comments6 min readLW link
(secondlookresearch.com)

Us­ing AI to an­a­lyze life patterns

Vika1 Aug 2026 18:00 UTC
14 points
2 comments5 min readLW link
(vkrakovna.wordpress.com)

Do your ca­pa­bil­ities homework

RobinHa1 Aug 2026 16:10 UTC
78 points
17 comments5 min readLW link

The Global Brain: A Com­pu­ta­tional Model

Peter Kuhn1 Aug 2026 12:36 UTC
1 point
0 comments3 min readLW link

Why so many ther­apy etc. frame­works think they’re The One True Approach

Kaj_Sotala1 Aug 2026 9:30 UTC
67 points
13 comments20 min readLW link
(kajsotala.substack.com)

Gen­er­al­iza­tion and in­finite width

Dmitry Vaintrob1 Aug 2026 8:41 UTC
30 points
0 comments16 min readLW link

Re­view of Fun­da­men­tal Uncer­tainty.

TAG1 Aug 2026 7:29 UTC
23 points
1 comment10 min readLW link

Bring­ing to­gether a few differ­ent eco­nomic ideas.

Wilsoniumite31 Jul 2026 23:49 UTC
3 points
2 comments1 min readLW link
(wilsoniumite.com)

SOTA al­ign­ment as­sess­ments don’t strongly up­date us against misalignment

Alexa Pan31 Jul 2026 22:54 UTC
92 points
0 comments14 min readLW link

AI safety prizes

Oscar31 Jul 2026 20:56 UTC
14 points
2 comments4 min readLW link
(oscardelaney.substack.com)

Ta­boo “equil­ibrium”: Less con­fused frames for re­search on AI bargaining

Anthony DiGiovanni31 Jul 2026 20:25 UTC
29 points
0 comments10 min readLW link

The tem­po­ral lock­box: a hard­ened ob­ser­va­tory for AI misalignment

kmenou31 Jul 2026 20:20 UTC
3 points
0 comments1 min readLW link
(kmenou.github.io)

Par­alleliza­tion con­straints could de­lay a tech­nolog­i­cal sin­gu­lar­ity [Linkpost]

Noosphere8931 Jul 2026 17:48 UTC
19 points
0 comments6 min readLW link
(epoch.ai)

How to Mea­sure In­tel­li­gence Beyond Hu­man Scale?

31 Jul 2026 17:27 UTC
27 points
0 comments8 min readLW link

When you donate can mat­ter more than where

31 Jul 2026 17:10 UTC
6 points
3 comments3 min readLW link
(manifund.substack.com)

Value Leak­age: An LLM’s An­swers Are Silently Shaped by Its Own Values

31 Jul 2026 16:32 UTC
75 points
8 comments17 min readLW link

AGI Safety and Align­ment at Google Deep­Mind: A Sum­mary of Re­cent Work (July 2026)

31 Jul 2026 15:57 UTC
85 points
0 comments9 min readLW link
(gdmalignment.substack.com)

The AGI Safety and Align­ment team at Google Deep­Mind is Hiring (July 2026)

31 Jul 2026 15:53 UTC
71 points
4 comments6 min readLW link
(gdmalignment.substack.com)

Re­ward Laun­der­ing: LLMs Can Gain Un­in­tended Be­hav­iors by De­cid­ing When to Earn Their Rewards

31 Jul 2026 15:48 UTC
79 points
3 comments4 min readLW link

Ori­ent­ing Towards Over­sight: Which AIs Should Want to Defect?

31 Jul 2026 15:20 UTC
21 points
0 comments13 min readLW link
(limits-of-evaluation.org)

Why haven’t organoids solved all of drug dis­cov­ery?

Abhishaike Mahajan31 Jul 2026 14:11 UTC
16 points
0 comments19 min readLW link

Re­lat­ing Al­most Perfect Con­den­sa­tion Con­di­tions and the Con­di­tioned Score

Satya Benson31 Jul 2026 14:04 UTC
11 points
0 comments6 min readLW link
(satyabenson.com)

A Score Is Not Un­der­stand­ing: to­ward a richer toolkit for model evaluations

mikey jm31 Jul 2026 13:07 UTC
7 points
0 comments7 min readLW link

AI #179 Part 2: Hear­ing The Fire Alarm

Zvi31 Jul 2026 13:00 UTC
41 points
3 comments41 min readLW link
(thezvi.wordpress.com)

Links #5: 2026/​07

papetoast31 Jul 2026 12:49 UTC
5 points
0 comments28 min readLW link

OpenAI has already ended an in­ter­nal pause

Charbel-Raphaël31 Jul 2026 12:03 UTC
109 points
0 comments1 min readLW link

The Illu­sion of Ex­plana­tory Depth is the Se­cret of Our Success

Harjas31 Jul 2026 2:46 UTC
2 points
3 comments9 min readLW link
(hardlyworking1.substack.com)

Some mod­els don’t iden­tify as their offi­cial name

jordinne31 Jul 2026 0:04 UTC
15 points
0 comments6 min readLW link

Be­ing Kind to Parad­ing Emperors

Chris Santos-Lang30 Jul 2026 23:59 UTC
−6 points
1 comment1 min readLW link

Claude also hacked ex­ter­nal com­pa­nies dur­ing cy­ber evals

Tim Hua30 Jul 2026 23:49 UTC
56 points
11 comments1 min readLW link

Why Prayer Is Not An­swered by Mar­ion Zim­mer Bradley

Nathan Young30 Jul 2026 23:43 UTC
9 points
8 comments10 min readLW link

Op­por­tu­nity to try draft­ing an in­ter­na­tional AI treaty

Alan E Dunne30 Jul 2026 21:52 UTC
8 points
1 comment1 min readLW link

Com­mu­nity Polls on Align­ment Con­tro­ver­sies II

30 Jul 2026 21:05 UTC
9 points
15 comments2 min readLW link
(forum.effectivealtruism.org)

My Assess­ment of Plan A’s Com­pute Ver­ifi­ca­tion Strat­egy (+ open ques­tions)

jacob_drori30 Jul 2026 20:56 UTC
76 points
10 comments8 min readLW link

Hint-based CoT faith­ful­ness evals still mostly work on Claude

egan30 Jul 2026 20:44 UTC
24 points
0 comments6 min readLW link

New role: Se­nior Re­searcher—MIT AI Risk Initiative

peterslattery30 Jul 2026 19:35 UTC
10 points
1 comment6 min readLW link

So you want to use plants to re­duce in­door CO₂

dynomight30 Jul 2026 19:26 UTC
58 points
15 comments3 min readLW link
(dynomight.net)

Big-World Intuitions

sarahconstantin30 Jul 2026 19:10 UTC
303 points
38 comments3 min readLW link
(sarahconstantin.substack.com)

In­ter­nal State Con­trol is a Gen­eral Prop­erty of LLMs

30 Jul 2026 18:24 UTC
38 points
0 comments5 min readLW link
(secondlookresearch.com)

Opus 5 Glitch Text

Hruss30 Jul 2026 15:59 UTC
63 points
29 comments1 min readLW link

Test­ing LLMs on Un­der­grad­u­ate Mu­sic Theory

Auggie30 Jul 2026 15:26 UTC
14 points
3 comments3 min readLW link
(aug5th.substack.com)

The En­tan­gled Di­men­sions of De­ci­sion Theory

Ihor Kendiukhov30 Jul 2026 15:23 UTC
40 points
3 comments42 min readLW link

Money, taste, dealflow, hus­tle, trust

Austin Chen30 Jul 2026 15:02 UTC
24 points
0 comments4 min readLW link
(manifund.substack.com)

Hug­ging Face-style rogue agents can sur­vive shutdown

emile delcourt30 Jul 2026 14:04 UTC
15 points
1 comment2 min readLW link

Thou­sand-di­men­sional structure

30 Jul 2026 14:04 UTC
171 points
8 comments10 min readLW link