Dutch-book re­sis­tant prob­a­bil­ity over cen­tered worlds

jessicata8 Aug 2026 23:05 UTC
31 points
7 comments7 min readLW link
(unstableontology.com)

Why Low Fer­til­ity Rates Are a Pos­i­tive Feed­back Loop

LoopGameScrollMonkey8 Aug 2026 20:17 UTC
46 points
1 comment25 min readLW link

Glimpses of superintelligence

PratyushRT8 Aug 2026 20:15 UTC
7 points
0 comments3 min readLW link

‘AI Es­caped Its Sand­box’ — What Does That Ac­tu­ally Mean?

Jakub Halmeš8 Aug 2026 18:55 UTC
12 points
2 comments6 min readLW link
(unpredictabletokens.substack.com)

What Hap­pened: OpenAI and HuggingFace

Zvi8 Aug 2026 16:10 UTC
111 points
7 comments11 min readLW link
(thezvi.wordpress.com)

In­duc­ing self-other over­lap with SFT re­duces de­cep­tion at scale, but gen­er­al­iza­tion re­mains uneven

Marc Carauleanu8 Aug 2026 15:21 UTC
24 points
1 comment7 min readLW link

FAQ: Isn’t AGI com­ing too soon for re­pro­ge­net­ics to help?

TsviBT8 Aug 2026 8:54 UTC
150 points
69 comments14 min readLW link

Rea­son­ing was not made for Deduction

epicurus8 Aug 2026 8:42 UTC
15 points
1 comment11 min readLW link

Sim­plify­ing the an­thropic im­pos­si­bil­ity result

Stuart_Armstrong8 Aug 2026 8:25 UTC
24 points
13 comments3 min readLW link

CLT Fea­tures Sharpen the Cycli­cal Day-of-Week Man­i­fold in Gemma-2-2b

Anna Marbut8 Aug 2026 2:30 UTC
7 points
0 comments5 min readLW link

AI Reg­u­la­tion Map: a view of AI gov­er­nance in 196 countries

Ria Deane8 Aug 2026 1:23 UTC
7 points
0 comments3 min readLW link

Don’t Build Mindreading

KellerScholl8 Aug 2026 0:30 UTC
298 points
69 comments4 min readLW link
(keller.substack.com)

Self-mon­i­tor­ing doesn’t scale (with­out these 3 coun­ter­mea­sures)

Morgan S7 Aug 2026 21:39 UTC
23 points
0 comments5 min readLW link

Public ev­i­dence of the OpenAI-Hug­gingFace AI attack

beyarkay (Boyd Kane)7 Aug 2026 21:21 UTC
94 points
6 comments7 min readLW link

Don’t Inoc­u­late Every­thing: Strat­ified Inoc­u­la­tion Prompt­ing Nar­rows Back­doors and Pre­serves De­sired Traits

7 Aug 2026 21:16 UTC
45 points
3 comments13 min readLW link

Job-Less Utopia: Macroe­co­nomics in the Age of AGI

Marcus Hutter7 Aug 2026 20:15 UTC
35 points
3 comments1 min readLW link

AI Safety at the Fron­tier: Paper High­lights of July 2026

gasteigerjo7 Aug 2026 19:50 UTC
5 points
0 comments1 min readLW link

Con­tra MacAskill on sav­ing money for the in­tel­li­gence explosion

Carol N7 Aug 2026 17:30 UTC
17 points
0 comments2 min readLW link
(manifund.substack.com)

How to pace the US frontier

7 Aug 2026 17:19 UTC
96 points
1 comment21 min readLW link
(blog.aifutures.org)

OpenAI Trained Its Models For Months While Those Models Were Co­or­di­nat­ing Ex­ploits Via Mes­sage Boards

Zvi7 Aug 2026 17:01 UTC
164 points
15 comments41 min readLW link
(thezvi.wordpress.com)

Item Re­sponse The­ory for AI Safety

7 Aug 2026 16:59 UTC
22 points
2 comments7 min readLW link

Some wild meta­physics that seem­ingly ev­ery­one needs to ac­cept?

Elias Schmied7 Aug 2026 16:13 UTC
16 points
5 comments6 min readLW link

Agenda: In­fras­truc­ture for Trad­ing with Par­tially Misal­igned AIs

VojtaKovarik7 Aug 2026 14:13 UTC
18 points
2 comments9 min readLW link
(limits-of-evaluation.org)

Up­com­ing Dove­tail fel­low talks & discussion

Alex_Altair7 Aug 2026 13:26 UTC
17 points
1 comment2 min readLW link

mod­els may be­have differ­ently in graded epi­sodes (a tirade)

nostalgebraist7 Aug 2026 11:49 UTC
278 points
31 comments61 min readLW link

Hypotheticals

Ramseyian7 Aug 2026 7:35 UTC
1 point
0 comments6 min readLW link

Strug­gles With Complexity

warner7 Aug 2026 7:25 UTC
4 points
0 comments2 min readLW link

How to Think like a Hu­man 101

Morpheus7 Aug 2026 6:09 UTC
−17 points
3 comments1 min readLW link

Some SHA-256 hashes

Radford Neal6 Aug 2026 23:21 UTC
9 points
4 comments1 min readLW link

Have mod­els re­port prov­able se­cu­rity bugs in their environment

anithite6 Aug 2026 22:52 UTC
24 points
2 comments4 min readLW link

Why do mod­els task game?

6 Aug 2026 22:16 UTC
94 points
1 comment24 min readLW link

Con­tra Oster on Al­co­hol in Preg­nancy. Part 1. The phar­ma­coki­net­ics of al­co­hol metabolism

Mvolz6 Aug 2026 21:34 UTC
99 points
2 comments7 min readLW link

Side-Effects of Length Penalty in RL

6 Aug 2026 21:23 UTC
13 points
3 comments18 min readLW link

User aware­ness in fron­tier models

6 Aug 2026 20:43 UTC
59 points
1 comment12 min readLW link
(transluce.org)

The FRONTIER Act barely cre­ates its im­ple­ment­ing office

Philip Dowdell6 Aug 2026 19:58 UTC
11 points
0 comments3 min readLW link

My Pri­vate Per­sonal Agent

danielms6 Aug 2026 19:53 UTC
13 points
2 comments5 min readLW link

Traf­fic Shap­ing for Work­load Classification

Andrew Dickson6 Aug 2026 19:49 UTC
10 points
0 comments8 min readLW link
(lucidcomputing.substack.com)

You need to stop X com­pa­nies to get a Y-month pause

Expertium6 Aug 2026 19:30 UTC
29 points
4 comments2 min readLW link

Three years of progress in 500 lines of code

Gerard Boxo6 Aug 2026 18:27 UTC
65 points
4 comments4 min readLW link

[$500 Bounty] I’m offer­ing a bounty of $500 for some­one with red-team­ing skills to build at­tack LLM pipelines for large-scale on­line deanonymiza­tion.

[deleted]6 Aug 2026 18:24 UTC
3 points
5 comments1 min readLW link

Model Or­ganisms of Sand­bag­ging in the Wild

Vladimir Ivanov6 Aug 2026 17:25 UTC
30 points
0 comments8 min readLW link

Func­tion vec­tors as a model diffing tool: 17 heads re­pair a bad fine-tune

Aniket Ghosh6 Aug 2026 14:33 UTC
28 points
0 comments17 min readLW link

How to define P(doom) and why it matters

Christopher King6 Aug 2026 14:31 UTC
24 points
6 comments3 min readLW link

AI #180: No Longer In Charge

Zvi6 Aug 2026 13:40 UTC
30 points
0 comments36 min readLW link
(thezvi.wordpress.com)

The Open Prob­lems of the AI Align­ment Field and their Cruxes

Gunnar_Zarncke6 Aug 2026 12:38 UTC
65 points
6 comments5 min readLW link

Ma­tryoshka NLAs: train­ing ac­ti­va­tion ver­bal­iz­ers to front­load re­con­struc­tion-rele­vant information

6 Aug 2026 9:55 UTC
30 points
0 comments7 min readLW link

Why You Should Al­most Never Use AI to Write Any­thing Substantive

Erich_Grunewald6 Aug 2026 9:50 UTC
49 points
13 comments10 min readLW link
(www.erichgrunewald.com)

Re­ward is hy­per­sti­tional information

Abhimanyu Pallavi Sudhir6 Aug 2026 8:05 UTC
19 points
0 comments7 min readLW link

The Pas­sive Man

warner6 Aug 2026 7:28 UTC
11 points
0 comments2 min readLW link

The Ghost Scale

Abraham Haskins6 Aug 2026 2:30 UTC
−7 points
0 comments1 min readLW link