Toy Model of Ac­ti­va­tion Obfuscation

Jesse Li14 Aug 2026 23:25 UTC
16 points
0 comments7 min readLW link
(jesseli2002.github.io)

Your Agents Are Not Time Aware

14 Aug 2026 23:17 UTC
40 points
4 comments11 min readLW link

An­nounc­ing: Iliad’s New 2026 Fellowships

14 Aug 2026 22:41 UTC
37 points
0 comments1 min readLW link

Train­ing a Con­cep­tual Rea­son­ing Judge

14 Aug 2026 20:33 UTC
17 points
1 comment6 min readLW link

Why I’m Skep­ti­cal of Longtermism

James Brobin14 Aug 2026 20:17 UTC
1 point
13 comments3 min readLW link

Empty Cham­bers, Miss­ing Stairs

finitude14 Aug 2026 19:37 UTC
8 points
0 comments5 min readLW link

Do It Like Darwin

derelict543214 Aug 2026 19:11 UTC
14 points
5 comments7 min readLW link

The Day Hu­man­ity Died (Par­ody of Amer­i­can Pie)

Bentham's Bulldog14 Aug 2026 18:27 UTC
8 points
1 comment3 min readLW link

Scry­ing, Model­ing, and Nerdsnipe

Cole Wyeth14 Aug 2026 17:59 UTC
76 points
2 comments5 min readLW link

Who we’d like as regrantors

14 Aug 2026 17:40 UTC
8 points
0 comments4 min readLW link
(manifund.substack.com)

How the Amer­i­can Ex­ec­u­tive Could Con­trol AI Companies

14 Aug 2026 15:43 UTC
74 points
0 comments13 min readLW link

Open Prob­lems in Mechanis­tic In­tepretabil­ity of Biolog­i­cal AIs

Ihor Kendiukhov14 Aug 2026 15:35 UTC
21 points
0 comments2 min readLW link

V&V takes on “Pac­ing the fron­tier”

Yoav Hollander14 Aug 2026 13:25 UTC
4 points
2 comments9 min readLW link
(blog.foretellix.com)

Avoid­ing the cor­po­rate treach­er­ous turn: crowd-sourc­ing de­sign ideas

Stuart_Armstrong14 Aug 2026 13:10 UTC
26 points
14 comments1 min readLW link

Fron­tier agents don’t com­ply with stan­dards, even when in­structed to

14 Aug 2026 12:08 UTC
42 points
5 comments7 min readLW link

What Mor­mons get right about com­mu­nity building

Jacob Brinton14 Aug 2026 7:53 UTC
37 points
2 comments4 min readLW link

Don’t for­get why learn­ing is important

Roman Ross14 Aug 2026 7:11 UTC
10 points
1 comment5 min readLW link

What If We En­forced AI Model Safety At the Level Of GPUs?

Mayowa Osibodu14 Aug 2026 6:10 UTC
1 point
5 comments4 min readLW link

Chat­ting With AIs: A Breakdown

Aditya14 Aug 2026 5:21 UTC
3 points
0 comments1 min readLW link

Fea­tures that cur­rent AIs don’t have that fu­ture AIs will have

Alexander Gietelink Oldenziel13 Aug 2026 21:49 UTC
57 points
9 comments2 min readLW link

Some Ways I Think About Eval­u­at­ing Grant Applications

sarahconstantin13 Aug 2026 21:30 UTC
61 points
0 comments10 min readLW link
(sarahconstantin.substack.com)

The Left Should Start Tak­ing AI Ca­pa­bil­ities Seriously

Alexei G13 Aug 2026 21:03 UTC
−13 points
0 comments9 min readLW link
(www.onethousandmeans.com)

Is Align­ment Even Falsifi­able? Mid­dle Align­ment, An Align­ment Tax­on­omy, and Break­ing The Prob­lem Down Into Steps

Savannah Harlan13 Aug 2026 20:57 UTC
5 points
0 comments22 min readLW link

Con­crete Gen­er­al­ist Pro­jects in AI Safety (and how to do them)

13 Aug 2026 19:26 UTC
25 points
2 comments6 min readLW link
(forum.effectivealtruism.org)

Com­par­ing Congress’s Two AI Emer­gency Shut­down Mechanisms

Philip Dowdell13 Aug 2026 18:53 UTC
13 points
0 comments8 min readLW link

Is biose­cu­rity over­sat­u­rated?

Master Chief13 Aug 2026 18:13 UTC
12 points
1 comment1 min readLW link

How to An­swer a Ques­tion Without An­swer­ing The Question

Kabir Kumar13 Aug 2026 17:06 UTC
37 points
3 comments1 min readLW link

How My Stu­dents Think About AI

dvd13 Aug 2026 16:56 UTC
779 points
111 comments12 min readLW link

AI #181: As­tra Goes Cy­ber Critical

Zvi13 Aug 2026 15:20 UTC
39 points
1 comment52 min readLW link
(thezvi.wordpress.com)

Au­to­mated al­ign­ment runs are hard to study!

13 Aug 2026 15:04 UTC
68 points
4 comments9 min readLW link

Longter­mism Seems Like A Religion

James Brobin13 Aug 2026 14:43 UTC
2 points
8 comments3 min readLW link

Defense Against the De­cep­tive Arts

Kabir Kumar13 Aug 2026 14:21 UTC
11 points
11 comments1 min readLW link

Some prob­lems in de­ci­sion the­ory in­cor­rectly pre­con­di­tion on policy.

Canaletto13 Aug 2026 12:04 UTC
10 points
4 comments3 min readLW link

Pat­terns and prob­lems in emerg­ing mul­ti­a­gent sys­tems (An­thropic, Fron­tier Red Team)

Julian Bradshaw13 Aug 2026 4:10 UTC
42 points
3 comments1 min readLW link
(www.anthropic.com)

The Descen­ders and the Ab­sorbers: Two Per­spec­tives on Deep Learning

larry-dial13 Aug 2026 4:02 UTC
11 points
0 comments3 min readLW link

OC ACXLW Meetup #118 — The 40% Prob­lem & The Field That Re­fused to Die

Michael Michalchik13 Aug 2026 3:25 UTC
1 point
0 comments9 min readLW link

Mea­sur­ing Ac­ti­va­tion Con­trol in LLMs

13 Aug 2026 2:04 UTC
50 points
8 comments7 min readLW link

How Valuable are BOTECs?

Roman Ross13 Aug 2026 0:31 UTC
9 points
0 comments8 min readLW link

Mea­sur­ing Eval Aware­ness: The Real­ism Win Rate is Fragile

13 Aug 2026 0:26 UTC
14 points
3 comments4 min readLW link

What hap­pened when I tried to be vegan

finitude13 Aug 2026 0:21 UTC
44 points
2 comments5 min readLW link

LLMs have the ca­pac­ity for self-im­posed steganography

Edward Cant13 Aug 2026 0:16 UTC
7 points
0 comments7 min readLW link

Free will is like temperature

Optimization Process12 Aug 2026 23:32 UTC
47 points
10 comments1 min readLW link

Im­pact mar­kets made concrete

Carol N12 Aug 2026 22:45 UTC
6 points
0 comments6 min readLW link
(manifund.substack.com)

Un­block­ing AI’s Con­tinual Learn­ing: Hints From How Hu­mans Learn

nimakeivan12 Aug 2026 18:28 UTC
3 points
23 comments10 min readLW link

Monthly Roundup #45: Au­gust 2026

Zvi12 Aug 2026 17:50 UTC
24 points
1 comment20 min readLW link
(thezvi.wordpress.com)

Find­ing the Seams of Perception

jimmy12 Aug 2026 17:33 UTC
30 points
0 comments2 min readLW link
(beneathpsychology.com)

In­tro­duc­ing the Con­cep­tual Rea­son­ing Index

12 Aug 2026 17:08 UTC
83 points
16 comments5 min readLW link
(alignment.anthropic.com)

One at­ten­tion head car­ries knight forks in a chess trans­former, and here’s a new toolkit that found it.

David Litman12 Aug 2026 16:48 UTC
8 points
0 comments1 min readLW link
(github.com)

The Psy­chol­ogy of Cope: Ra­tion­al­ity Is Not Re­v­ersed Irrationality

Chris_Leong12 Aug 2026 12:20 UTC
27 points
12 comments3 min readLW link

De­mon Safety

Ben Pace12 Aug 2026 11:55 UTC
53 points
5 comments1 min readLW link