Quan­tify­ing CoT faithfulness

Mihir Sahasrabudhe1 Sep 2026 23:15 UTC
3 points
0 comments5 min readLW link

“Col­lu­sion” is just co­op­er­a­tion that you don’t like

Morgan S1 Sep 2026 22:04 UTC
22 points
15 comments4 min readLW link

What hap­pened to im­pact mar­kets?

Austin Chen1 Sep 2026 21:30 UTC
23 points
3 comments3 min readLW link

A night­watch­man on ev­ery probe: Su­per­in­tel­li­gent surveillance to pre­vent galac­tic anarchy

1 Sep 2026 21:04 UTC
16 points
0 comments7 min readLW link
(www.forethought.org)

No Bush, No Potato

Fernand01 Sep 2026 20:49 UTC
9 points
0 comments2 min readLW link

The Align­ment Jour­nal: Or­ga­ni­za­tion, Per­son­nel, and Scope

1 Sep 2026 20:49 UTC
95 points
5 comments11 min readLW link
(blog.alignmentjournal.org)

Fake voices: warp­ing the so­cial world

KatjaGrace1 Sep 2026 20:49 UTC
33 points
1 comment3 min readLW link
(worldspiritsockpuppet.substack.com)

You can rarely pet the dog in an LLM-gen­er­ated game

yamike1 Sep 2026 19:43 UTC
14 points
0 comments3 min readLW link
(mikeushakov.com)

Evolu­tion could en­code the brain in DNA

olehif1 Sep 2026 18:23 UTC
10 points
11 comments1 min readLW link

Re­search sci­en­tists for CaML: In­ves­ti­gat­ing whether mid-train­ing can sur­vive RL

Jasmine Brazilek1 Sep 2026 18:13 UTC
16 points
2 comments2 min readLW link

Towards de­ploy­ment-time mis­al­ign­ment con­tinu­a­tion evals: les­sons from re­cent loss of con­trol incidents

Cath Ge-Wang1 Sep 2026 17:45 UTC
22 points
0 comments12 min readLW link
(cathgewang.substack.com)

Lu­mina Pro­biotic, Past and Fu­ture

jayterwahl1 Sep 2026 17:26 UTC
13 points
0 comments5 min readLW link

Lu­mina Pro­biotic Re­sponse to Crit­ics

jayterwahl1 Sep 2026 17:25 UTC
2 points
3 comments5 min readLW link

New HPMoR Pod­cast site

Eneasz1 Sep 2026 17:13 UTC
19 points
0 comments1 min readLW link

A Year of Atheism

Laiba Rehman ✦ RJ1 Sep 2026 16:19 UTC
62 points
4 comments8 min readLW link

PauseAI Has ‘offi­cially dis­endorsed’ PauseAI-US

nem1 Sep 2026 15:51 UTC
279 points
216 comments8 min readLW link

Best Way To Start a Lo­cal Group?

nem1 Sep 2026 15:44 UTC
12 points
2 comments1 min readLW link

Could in­ter­nal model trans­parency tame the AI race?

Karthik Tadepalli1 Sep 2026 15:41 UTC
7 points
1 comment6 min readLW link
(blog.karthiktadepalli.com)

Hug­gingFace At­tack Post­mortem: Civ­i­liza­tions, Re­ac­tions and Next Actions

Zvi1 Sep 2026 14:10 UTC
45 points
10 comments55 min readLW link
(thezvi.wordpress.com)

An “An­thropic Prin­ci­ple” for For­mu­la­tions of AI Alignment

Adam Chlipala1 Sep 2026 13:30 UTC
0 points
0 comments8 min readLW link

Images and AI: out of the loop

8491 Sep 2026 12:52 UTC
2 points
0 comments5 min readLW link

A new ver­sion of the Petrov Day booklet

laniakea1 Sep 2026 12:20 UTC
3 points
0 comments1 min readLW link

I tracked my emo­tions for 11 years and here’s what I found out about men­tal health

KatSpartz1 Sep 2026 12:19 UTC
44 points
3 comments12 min readLW link

We should pre­pare a play­book for the day af­ter a warn­ing shot

Yair Halberstadt1 Sep 2026 9:52 UTC
60 points
5 comments1 min readLW link

The Cog­ni­tive Dy­nam­ics of AI Philosophy

JonathanErhardt1 Sep 2026 8:10 UTC
13 points
0 comments2 min readLW link

AI Philos­o­phy Com­pe­ti­tion: $11,000 in prizes.

Elliott Thornley1 Sep 2026 6:26 UTC
29 points
11 comments1 min readLW link

In­sights into Curry’s Para­dox?

David Litman1 Sep 2026 5:40 UTC
12 points
18 comments1 min readLW link

Salad days

Zephaniah Roe1 Sep 2026 5:26 UTC
96 points
6 comments3 min readLW link

Re­sources for Large Agent Sys­tems Safety

Stephen Elliott1 Sep 2026 4:10 UTC
4 points
0 comments2 min readLW link

When Ac­ti­va­tion Or­a­cles learn not to read: Con­cept-Spe­cific Blind Spots in Fine-Tuned Oracles

1 Sep 2026 3:40 UTC
17 points
3 comments6 min readLW link
(arxiv.org)

Mak­ing an LLM in­ter­pretabil­ity playground

Pietro1 Sep 2026 3:30 UTC
1 point
0 comments3 min readLW link

Bricks and ex­po­nen­tials: a note on how I eval­u­ate pro­jects

Eli Tyre1 Sep 2026 2:21 UTC
39 points
1 comment4 min readLW link

Why, Un­der Phys­i­cal­ism, There Is No Hard Prob­lem of Consciousness

Eugene Verichev1 Sep 2026 2:06 UTC
9 points
2 comments11 min readLW link

Aquar­ium Se­cu­rity and Other Or­gani­sa­tional Priors

hal2k1 Sep 2026 2:01 UTC
8 points
0 comments2 min readLW link

ERA:AI Win­ter 2027 Fel­low­ship

gracereuben1 Sep 2026 2:00 UTC
2 points
0 comments1 min readLW link

A mil­lion au­thors of alignment

Diksha Gupta1 Sep 2026 1:58 UTC
5 points
0 comments3 min readLW link

Train­ing a Misal­igned Re­ward Seeker

1 Sep 2026 1:41 UTC
151 points
11 comments2 min readLW link
(alignment.anthropic.com)

Har­ness that catches gran­u­lar AI cheating

latent_affect1 Sep 2026 1:41 UTC
1 point
0 comments1 min readLW link

Fron­tier labs need an an­timemet­ics di­vi­sion.

Jackson Hurley1 Sep 2026 1:35 UTC
7 points
0 comments1 min readLW link

Open Thread Au­tumn 2026

habryka1 Sep 2026 0:05 UTC
18 points
14 comments1 min readLW link

Pop­u­la­tion ethics is a big deal, which is why I made this pop­u­la­tion ethics quiz

MichaelDickens31 Aug 2026 23:13 UTC
9 points
28 comments4 min readLW link

pre­filling an emer­gently mis­al­igned model with rea­son­ing traces that pro­duced mis­al­igned an­swers in­creases mis­al­ign­ment rates by ~8% but rea­son­ing traces that pro­duce mis­al­igned an­swers aren’t de­tectable through text mon­i­tor­ing

nesiacel31 Aug 2026 22:31 UTC
7 points
0 comments1 min readLW link

You Should Think About What Needs to Go Right + My List

Koby Lewis31 Aug 2026 22:02 UTC
16 points
1 comment5 min readLW link

How to solve home­less­ness: what spe­cific laws we need, how to get it past the op­po­si­tion, all with­out be­ing an asshole

KatSpartz31 Aug 2026 22:02 UTC
56 points
5 comments14 min readLW link

Nar­ra­tion pod­casts for newslet­ters by Red­wood, Epoch, Zvi, Sen­tinel, Paradigm 3, and others

peter_hartree31 Aug 2026 21:17 UTC
12 points
1 comment1 min readLW link

The in­verse Sim­mel scenario

Fernand031 Aug 2026 19:43 UTC
−1 points
5 comments3 min readLW link

Kan­tian consent

Fernand031 Aug 2026 19:43 UTC
−5 points
9 comments2 min readLW link

Wiki-libertarianism

Fernand031 Aug 2026 19:42 UTC
−8 points
2 comments7 min readLW link

The Struc­tural­ist Mindset

Fernand031 Aug 2026 19:40 UTC
5 points
0 comments2 min readLW link

Let’s fund weird AI safety projects

Ihor Kendiukhov31 Aug 2026 16:29 UTC
131 points
13 comments3 min readLW link