Some SHA-256 hashes

Radford Neal6 Aug 2026 23:21 UTC
9 points
4 comments1 min readLW link

Have mod­els re­port prov­able se­cu­rity bugs in their environment

anithite6 Aug 2026 22:52 UTC
24 points
2 comments4 min readLW link

Why do mod­els task game?

6 Aug 2026 22:16 UTC
94 points
1 comment24 min readLW link

Con­tra Oster on Al­co­hol in Preg­nancy. Part 1. The phar­ma­coki­net­ics of al­co­hol metabolism

Mvolz6 Aug 2026 21:34 UTC
99 points
2 comments7 min readLW link

Side-Effects of Length Penalty in RL

6 Aug 2026 21:23 UTC
13 points
3 comments18 min readLW link

User aware­ness in fron­tier models

6 Aug 2026 20:43 UTC
59 points
1 comment12 min readLW link
(transluce.org)

The FRONTIER Act barely cre­ates its im­ple­ment­ing office

Philip Dowdell6 Aug 2026 19:58 UTC
11 points
0 comments3 min readLW link

My Pri­vate Per­sonal Agent

danielms6 Aug 2026 19:53 UTC
13 points
2 comments5 min readLW link

Traf­fic Shap­ing for Work­load Classification

Andrew Dickson6 Aug 2026 19:49 UTC
10 points
0 comments8 min readLW link
(lucidcomputing.substack.com)

You need to stop X com­pa­nies to get a Y-month pause

Expertium6 Aug 2026 19:30 UTC
29 points
4 comments2 min readLW link

Three years of progress in 500 lines of code

Gerard Boxo6 Aug 2026 18:27 UTC
65 points
4 comments4 min readLW link

[$500 Bounty] I’m offer­ing a bounty of $500 for some­one with red-team­ing skills to build at­tack LLM pipelines for large-scale on­line deanonymiza­tion.

pinto6 Aug 2026 18:24 UTC
3 points
5 comments1 min readLW link

Model Or­ganisms of Sand­bag­ging in the Wild

Vladimir Ivanov6 Aug 2026 17:25 UTC
30 points
0 comments8 min readLW link

Func­tion vec­tors as a model diffing tool: 17 heads re­pair a bad fine-tune

Aniket Ghosh6 Aug 2026 14:33 UTC
28 points
0 comments17 min readLW link

How to define P(doom) and why it matters

Christopher King6 Aug 2026 14:31 UTC
24 points
6 comments3 min readLW link

AI #180: No Longer In Charge

Zvi6 Aug 2026 13:40 UTC
30 points
0 comments36 min readLW link
(thezvi.wordpress.com)

The Open Prob­lems of the AI Align­ment Field and their Cruxes

Gunnar_Zarncke6 Aug 2026 12:38 UTC
65 points
6 comments5 min readLW link

Ma­tryoshka NLAs: train­ing ac­ti­va­tion ver­bal­iz­ers to front­load re­con­struc­tion-rele­vant information

6 Aug 2026 9:55 UTC
30 points
0 comments7 min readLW link

Why You Should Al­most Never Use AI to Write Any­thing Substantive

Erich_Grunewald6 Aug 2026 9:50 UTC
49 points
13 comments10 min readLW link
(www.erichgrunewald.com)

Re­ward is hy­per­sti­tional information

Abhimanyu Pallavi Sudhir6 Aug 2026 8:05 UTC
19 points
0 comments7 min readLW link

The Pas­sive Man

warner6 Aug 2026 7:28 UTC
11 points
0 comments2 min readLW link

The Ghost Scale

Abraham Haskins6 Aug 2026 2:30 UTC
−7 points
0 comments1 min readLW link

Five coun­ter­in­tu­itive in­sights from Plan A

romeo6 Aug 2026 1:19 UTC
35 points
1 comment10 min readLW link

Fun­da­men­tal Uncer­tainty Es­say Contestants

Gordon Seidoh Worley5 Aug 2026 23:01 UTC
15 points
0 comments1 min readLW link
(www.uncertainupdates.com)

Alex Turner on Leav­ing Google Deep­Mind and Disagree­ments with Yudkowsky

Liron5 Aug 2026 22:45 UTC
59 points
1 comment40 min readLW link

Thomas Schel­ling’s No­bel Prize Speech: An As­ton­ish­ing Sixty Years: The Le­gacy of Hiroshima

Nathan Young5 Aug 2026 21:30 UTC
26 points
0 comments19 min readLW link

Ter­mi­nal-Bench Leader­board Rank­ings: Luck or Skill?

nateg5510155 Aug 2026 21:20 UTC
7 points
0 comments2 min readLW link

An ar­gu­ment of Parfit’s re­con­sid­ered with log­i­cal de­ci­sion theory

transhumanist_atom_understander5 Aug 2026 20:47 UTC
12 points
14 comments5 min readLW link

R-lens: Mak­ing J-lens More Faith­ful on Early Layers

5 Aug 2026 20:02 UTC
83 points
5 comments7 min readLW link

Mea­sur­ing cod­ing agent mis­al­ign­ment in the wild

snaz5 Aug 2026 20:00 UTC
40 points
2 comments7 min readLW link

An In­ter­na­tional AI Slow­down Is Ready When­ever Poli­ti­ci­ans Are

5 Aug 2026 19:15 UTC
56 points
4 comments9 min readLW link
(newsletter.ai-frontiers.org)

Ar­gu­ments for P

Cleo Nardo5 Aug 2026 19:00 UTC
549 points
66 comments2 min readLW link

Gen­er­al­ized athe­ism rules out “in­ac­cu­rate simu­la­tion”-ism.

Eliezer Yudkowsky5 Aug 2026 18:35 UTC
165 points
56 comments9 min readLW link

Berkeley Ge­nomics Pro­ject seek­ing hires (and col­labs)

TsviBT5 Aug 2026 18:32 UTC
41 points
9 comments9 min readLW link

Charles Good­hart Ele­men­tary School

Zack_M_Davis5 Aug 2026 17:06 UTC
45 points
3 comments2 min readLW link
(zackmdavis.net)

The Three AI Pills

Zvi5 Aug 2026 16:10 UTC
75 points
22 comments15 min readLW link
(thezvi.wordpress.com)

Notes on the pos­si­bil­ity of moral progress

MichaelDickens5 Aug 2026 16:05 UTC
14 points
10 comments7 min readLW link

Re­grant­ing in 2026: now more than ever

5 Aug 2026 13:59 UTC
2 points
0 comments5 min readLW link
(manifund.substack.com)

Ran­dom­ness and Dooms­day for Weak agents

Ben5 Aug 2026 12:24 UTC
32 points
6 comments7 min readLW link

Ver­ti­cal Tabs in Chrome

jefftk5 Aug 2026 12:10 UTC
48 points
3 comments1 min readLW link
(www.jefftk.com)

The goal­posts are shrouded, not moving

philh5 Aug 2026 11:40 UTC
134 points
6 comments1 min readLW link
(reasonableapproximation.net)

Help (re)start AI Safety at UPenn!

5 Aug 2026 9:42 UTC
15 points
0 comments1 min readLW link

Per­sona Cor­rup­tion and Role Mis­cast­ing in Emer­gent Misalignment

unruly abstractions5 Aug 2026 4:56 UTC
17 points
0 comments12 min readLW link

Warn­ing Shots for AI Ex­is­ten­tial Risk: Will so­ciety an­swer the wake-up calls it re­ceives?

5 Aug 2026 4:30 UTC
31 points
0 comments13 min readLW link
(techgov.intelligence.org)

AISafety.com Hackathon 2026

Bryce Robertson5 Aug 2026 4:15 UTC
4 points
0 comments1 min readLW link

See all up­com­ing AI safety events and train­ing programs

5 Aug 2026 3:28 UTC
19 points
0 comments1 min readLW link

Var­i­ous Ways to En­able People

warner5 Aug 2026 0:03 UTC
12 points
0 comments2 min readLW link

The AI Race is Not a Pri­soner’s Dilemma

Vaughn Papenhausen4 Aug 2026 23:45 UTC
67 points
6 comments8 min readLW link

Norms Are Bro­ken Laws

sirawit4 Aug 2026 23:42 UTC
−14 points
5 comments4 min readLW link

Why I think we live in a “simu­la­tion”

Eye You4 Aug 2026 22:43 UTC
4 points
6 comments5 min readLW link