Creativity Beyond the Manifold

Kartikay Luthra20 Aug 2026 23:31 UTC
8 points
2 comments6 min readLW link

Ablat­ing 1 of a chess trans­former’s 128 at­ten­tion heads makes the model stop find­ing Paul Mor­phy’s queen sac­ri­fice

David Litman20 Aug 2026 21:17 UTC
7 points
1 comment1 min readLW link
(www.youtube.com)

Llama will aban­don a cor­rect an­swer if it thinks you’re educated

Nick Merrill20 Aug 2026 20:06 UTC
33 points
7 comments2 min readLW link

Why self-fund your re­grant­ing?

20 Aug 2026 19:18 UTC
11 points
0 comments3 min readLW link
(manifund.substack.com)

If Aliens Ex­ist, We Should Ex­pect to Find Them Around Now

jehan20 Aug 2026 18:57 UTC
8 points
2 comments1 min readLW link
(www.jehanazad.com)

Cross-Dataset Trans­fer Eval­u­a­tion of De­cep­tion Probes in Smaller Models

Jollen Dai20 Aug 2026 17:06 UTC
8 points
1 comment5 min readLW link

The Fourth Humiliation

Nathalie Kirch20 Aug 2026 17:03 UTC
52 points
1 comment6 min readLW link

AI #182: Pause For Reflection

Zvi20 Aug 2026 14:50 UTC
37 points
1 comment47 min readLW link
(thezvi.wordpress.com)

We Must Re­mem­ber That Our World Con­tains Hell

James Brobin20 Aug 2026 14:22 UTC
152 points
52 comments3 min readLW link

Thoughts on Tak­ing OpenAI Foun­da­tion Funding

jefftk20 Aug 2026 13:00 UTC
31 points
3 comments5 min readLW link
(www.jefftk.com)

Mak­ing sense of the mis­al­ign­ment risk model in the An­thropic Risk Re­port (Au­gust 2026)

jasmine.ren20 Aug 2026 2:33 UTC
27 points
5 comments8 min readLW link

Ap­pear­ing Un­fair in the Wrong Direction

MossyFallenFriend20 Aug 2026 2:21 UTC
20 points
0 comments6 min readLW link

Oliver Habryka in­ter­view — Good Dis­course Needs Leadership

The Students20 Aug 2026 2:14 UTC
8 points
1 comment1 min readLW link
(www.youtube.com)

Steer­ing Role Confusion

20 Aug 2026 2:09 UTC
21 points
1 comment4 min readLW link

Notes on “Ei­genBench: A Com­par­a­tive Be­hav­ioral Mea­sure of Value Align­ment”

Shunk20 Aug 2026 2:03 UTC
8 points
0 comments5 min readLW link

Quan­tilized de­bate and con­sul­tancy in image en­vi­ron­ments: pro­to­col de­sign les­sons for scal­able over­sight experiments

emanuelr19 Aug 2026 23:03 UTC
14 points
0 comments28 min readLW link

Judg­ing eth­i­cal the­o­ries by up­date rules, not by ac­tion rankings

yatharth19 Aug 2026 22:41 UTC
11 points
1 comment5 min readLW link

Why can’t we have nice things? Like, speci­fi­cally?

Elizabeth19 Aug 2026 20:20 UTC
58 points
3 comments3 min readLW link
(acesounderglass.com)

OpenAI Takes Ini­tial Steps To Ad­dress Its Align­ment Problems

Zvi19 Aug 2026 20:00 UTC
33 points
1 comment21 min readLW link
(thezvi.wordpress.com)

34% of the US pub­lic is now aware of AI xrisk, and the curve is steepening

otto.barten19 Aug 2026 19:18 UTC
30 points
0 comments1 min readLW link

Science and News Twit­ter/​X Summarizer

sarahconstantin19 Aug 2026 19:00 UTC
32 points
5 comments1 min readLW link
(sarahconstantin.substack.com)

Separat­ing cheat­ing and aver­sion in task-gaming

Mihir Sahasrabudhe19 Aug 2026 18:49 UTC
14 points
0 comments4 min readLW link

MATS Win­ter 2027 Ap­pli­ca­tions Are Now Open

Raj Thimmiah19 Aug 2026 18:48 UTC
8 points
2 comments2 min readLW link

The Rogue Agent Ex­plo­sion Will Be Mostly Invisible

Steven McCulloch19 Aug 2026 18:46 UTC
102 points
20 comments16 min readLW link

RL cre­ates split personas

Jan Betley19 Aug 2026 18:23 UTC
235 points
20 comments4 min readLW link

En­force­ment Be­gins: the EU AI Act and Sys­temic Risks

Schizoid Rentoid19 Aug 2026 18:04 UTC
1 point
0 comments1 min readLW link

A cir­cuit prior in NN-bayes

19 Aug 2026 17:14 UTC
78 points
4 comments1 min readLW link

A failed solu­tion to open-source game theory

Richard Willis19 Aug 2026 16:11 UTC
15 points
4 comments4 min readLW link

In­side the mind of a fair player cooperating

transhumanist_atom_understander19 Aug 2026 14:53 UTC
40 points
2 comments3 min readLW link

Some rea­sons al­ign­ment doesn’t gen­er­al­ise well

Lucius Bushnaq19 Aug 2026 14:18 UTC
128 points
3 comments9 min readLW link

The Con­trol Paradox

Ephraiem Sarabamoun19 Aug 2026 12:39 UTC
8 points
0 comments2 min readLW link

De­bate Train­ing Re­duces Re­ward Hack­ing in RLAIF

19 Aug 2026 12:17 UTC
61 points
6 comments6 min readLW link
(gdmalignment.substack.com)

Con­cerns About Per­sonas, Multi-Agent Align­ment, and Role Theory

Davidmanheim19 Aug 2026 12:03 UTC
21 points
1 comment8 min readLW link

Read­ing List on Wise AI

Chris_Leong19 Aug 2026 8:41 UTC
22 points
0 comments1 min readLW link

Nat­u­ral Lan­guage Transcoders

anwenh19 Aug 2026 1:49 UTC
14 points
0 comments4 min readLW link

What AI scores (while we can still keep score)

dan.parshall19 Aug 2026 1:40 UTC
12 points
5 comments3 min readLW link

Cui bono? ChatGPT-4o shows non-de­cep­tive strate­gic per­sua­sion: a Proof-of-Con­cept study

Pieter Barkema18 Aug 2026 23:25 UTC
7 points
0 comments5 min readLW link

Whack-a-mole with a bro­ken ham­mer: does a model in­ter­nally track its au­toma­ton state?

star2vec18 Aug 2026 23:25 UTC
7 points
0 comments8 min readLW link

GPT-2′s IOI be­hav­ior is defined where the pa­per’s al­gorithm isn’t

Zach Allen18 Aug 2026 23:23 UTC
7 points
0 comments2 min readLW link

Put­ting your mask on first

emiliob18 Aug 2026 23:23 UTC
1 point
0 comments1 min readLW link
(emiliobarkett.github.io)

Tran­shu­man­ism & Hu­man En­hance­ment Meetup (On­line)

nxzt18 Aug 2026 23:22 UTC
1 point
0 comments1 min readLW link

Goal-di­rected op­ti­mi­sa­tion re­quires in­for­ma­tion about the world

dubinsschwarz18 Aug 2026 23:22 UTC
13 points
0 comments12 min readLW link

AI Sand­bag­ging (w/​ In­spect)

Leo H18 Aug 2026 23:20 UTC
8 points
0 comments5 min readLW link

An­thropic Risk Re­port: Au­gust 2026

Zvi18 Aug 2026 20:40 UTC
48 points
3 comments42 min readLW link
(thezvi.wordpress.com)

Fund­ing For­mal Meth­ods for the Cyberpocalypse

Max von Hippel18 Aug 2026 16:45 UTC
22 points
8 comments9 min readLW link

Policy ca­reer plan­ning in the age of im­mi­nent superintelligence

Peter Wildeford18 Aug 2026 14:33 UTC
70 points
5 comments6 min readLW link

Au­to­matic Pro­gram­ming Should Be More Like SQL

Adam Chlipala18 Aug 2026 12:29 UTC
7 points
0 comments12 min readLW link

AI Se­cu­rity is Harm Reduction

Quinn18 Aug 2026 12:28 UTC
39 points
3 comments1 min readLW link

LASR Labs Win­ter 2027 ap­pli­ca­tions are open!

18 Aug 2026 10:12 UTC
37 points
1 comment4 min readLW link

Price re­cur­sion is the ra­tio­nal the­ory of reward

Abhimanyu Pallavi Sudhir18 Aug 2026 3:22 UTC
6 points
0 comments9 min readLW link