Model­ing Con­cepts Probabilistically

Gretta Duleba1 Jul 2026 23:27 UTC
47 points
4 comments10 min readLW link

AI welfare re­search needs ba­sic science

1 Jul 2026 22:59 UTC
36 points
7 comments10 min readLW link

Claude Son­net 5 Is Not Fron­tier But Has Its Uses

Zvi1 Jul 2026 22:41 UTC
33 points
4 comments19 min readLW link
(thezvi.wordpress.com)

How Many Peo­ple Have Ever Lived in the United States?

Novalis1 Jul 2026 22:25 UTC
6 points
0 comments6 min readLW link

Con­ver­sa­tions With Cade Metz on the Rationalists

Zack_M_Davis1 Jul 2026 22:19 UTC
50 points
4 comments101 min readLW link

Do-it-your­self meta-analysis

kqr1 Jul 2026 22:04 UTC
14 points
2 comments5 min readLW link
(entropicthoughts.com)

When ca­pa­bil­ities work is the *safe* bet

RobinHa1 Jul 2026 20:53 UTC
35 points
0 comments1 min readLW link
(robinhaselhorst.com)

The Value of Veridi­cal Information

Silent Swift1 Jul 2026 20:23 UTC
1 point
1 comment6 min readLW link
(substack.com)

How in­evitable are most ac­cessible hard-tech star­tups?

Master Chief1 Jul 2026 20:21 UTC
4 points
2 comments1 min readLW link

What is Good? An­tiruin and Nonabsolutism

Silent Swift1 Jul 2026 20:10 UTC
−3 points
5 comments5 min readLW link
(substack.com)

Plas­tic Straws

robbiethompson1 Jul 2026 19:21 UTC
−2 points
3 comments5 min readLW link
(robbiewmthompson.com)

AI Mis­take Seeding

Taylor G. Lunt1 Jul 2026 18:49 UTC
28 points
2 comments4 min readLW link

In Par­tial, Pug­na­cious Defense of Func­tional De­ci­sion Theory

Mikewins1 Jul 2026 17:49 UTC
7 points
0 comments1 min readLW link

How to read tableaux, a for­mal sys­tem for modal logic with Kripke models

transhumanist_atom_understander1 Jul 2026 17:37 UTC
15 points
1 comment6 min readLW link

Con­sis­tency Train­ing while Miti­gat­ing Obfus­ca­tion via Rate Matching

1 Jul 2026 17:26 UTC
45 points
7 comments12 min readLW link

Dis­cov­er­ing Con­cept-Edit­ing Al­gorithms With LLM Agents

1 Jul 2026 16:07 UTC
27 points
0 comments1 min readLW link
(dmodel.ai)

Model ac­cess for third-par­ties — it’s a big deal!

Cleo Nardo1 Jul 2026 13:09 UTC
181 points
38 comments6 min readLW link

Most Cur­rent Model Or­ganisms Leak: Per­plex­ity Differenc­ing Often Re­veals Fine­tun­ing Objectives

1 Jul 2026 10:07 UTC
27 points
0 comments7 min readLW link

A Black Box Made Less Opaque (part 4)

Matthew McDonnell1 Jul 2026 7:30 UTC
8 points
0 comments9 min readLW link

The Once and Pre­sent Fable (Fable 5 restora­tion linkpost)

fluxxrider1 Jul 2026 7:20 UTC
15 points
5 comments1 min readLW link

When should you know the point?

KatjaGrace1 Jul 2026 6:31 UTC
33 points
3 comments1 min readLW link
(worldspiritsockpuppet.substack.com)

A CERN for AI is a dis­trac­tion; push for an IAEA instead

Charbel-Raphaël1 Jul 2026 6:30 UTC
48 points
2 comments4 min readLW link

You Should Come to The AI Protest

Ronak_Mehta1 Jul 2026 4:20 UTC
91 points
2 comments4 min readLW link

Ap­ply to the Inau­gu­ral PIBBSS Win­ter Re­search Fel­low­ship!

Ami941 Jul 2026 3:54 UTC
25 points
0 comments2 min readLW link

Why aren’t there more AlphaFolds?

nimakeivan1 Jul 2026 3:42 UTC
24 points
3 comments17 min readLW link

Please make me care about x-risk

Kate Delbeke1 Jul 2026 3:26 UTC
7 points
2 comments3 min readLW link

The Prob­lem with Chat

magfrump1 Jul 2026 3:19 UTC
6 points
0 comments1 min readLW link
(www.magfrump.net)

Green

Biff Wiff1 Jul 2026 1:10 UTC
33 points
7 comments2 min readLW link

Links #4: 2026/​06 Part 2

papetoast1 Jul 2026 0:43 UTC
8 points
1 comment30 min readLW link

‘AI alle­gory steganog­ra­phy’ in Claude short sto­ries in the Unslop con­test?

gwern1 Jul 2026 0:38 UTC
29 points
2 comments1 min readLW link
(www.hyperstitionai.com)

The con­se­quences of lock­ing in­tel­li­gence away: an in­tro­duc­tion to Claude re­lays in China

CMLKevin30 Jun 2026 22:48 UTC
112 points
6 comments2 min readLW link

“Cor­rect An­swer Fea­tures” Can­not Ex­plain Mul­ti­ple Choice Capabilities

Koby Lewis30 Jun 2026 22:46 UTC
8 points
0 comments3 min readLW link
(kobylewis.net)

The Once And Fu­ture Fable #5

Zvi30 Jun 2026 21:50 UTC
49 points
4 comments17 min readLW link
(thezvi.wordpress.com)

Clue­less­ness: Sum­mary of the ar­gu­ment, why it mat­ters, and counterarguments

Anthony DiGiovanni30 Jun 2026 20:54 UTC
25 points
6 comments9 min readLW link

Is it eth­i­cal to work on gen­eral-pur­pose robots given the risk of cy­ber­hack­ing?

Master Chief30 Jun 2026 19:27 UTC
6 points
1 comment1 min readLW link

That Which Can­not Be Poked With A Stick Is The Mind-Killer

Firinn30 Jun 2026 19:01 UTC
32 points
5 comments32 min readLW link

Con­nect to your past selves

PatrickDFarley30 Jun 2026 17:01 UTC
10 points
3 comments5 min readLW link

Un­jour­nal tool hub (find­ing & as­sess­ing high-im­pact re­search ques­tions, cruxes, pa­pers, etc.)

david reinstein30 Jun 2026 15:50 UTC
15 points
0 comments1 min readLW link

Pre­limi­nary in­ves­ti­ga­tion: KL penalties in RL can in­crease CoT unfaithfulness

30 Jun 2026 13:08 UTC
81 points
2 comments13 min readLW link

Struc­tural Proxies

Raymond Douglas30 Jun 2026 12:38 UTC
39 points
0 comments8 min readLW link

Why Pre­fer Any De­ci­sion The­ory?

J Bostock30 Jun 2026 12:06 UTC
30 points
24 comments6 min readLW link

What Ca­pable Agents Must Know: Why AI Con­scious­ness May Be an Inevitable Byproduct of Capability

Aran Nayebi30 Jun 2026 11:48 UTC
57 points
12 comments14 min readLW link

Agency is not a nat­u­ral kind (and why that might mat­ter for al­ign­ment)

SJ_Beard30 Jun 2026 8:50 UTC
36 points
11 comments4 min readLW link

In par­tial defence of p(doom)

Mikhail Samin30 Jun 2026 8:50 UTC
32 points
3 comments2 min readLW link

In­ter­per­son­al­ized recommendations

KatjaGrace30 Jun 2026 7:01 UTC
13 points
0 comments1 min readLW link
(worldspiritsockpuppet.substack.com)

How should you slow down AI progress if it be­comes nec­es­sary?

Felipe Calero-Forero30 Jun 2026 6:39 UTC
13 points
1 comment20 min readLW link
(felipecalerof.substack.com)

Sepa­ra­tion of Knowl­edge and Rea­son­ing?

HGross30 Jun 2026 6:13 UTC
7 points
0 comments1 min readLW link

More Failed Eg­gless Choux

jefftk30 Jun 2026 0:53 UTC
14 points
0 comments1 min readLW link
(www.jefftk.com)

The Slo­gan Strikes Again

mrdodson30 Jun 2026 0:34 UTC
10 points
2 comments3 min readLW link

Could AI Out­grow Con­scious­ness?

29 Jun 2026 20:02 UTC
6 points
2 comments4 min readLW link