Fea­tures that cur­rent AIs don’t have that fu­ture AIs will have

Alexander Gietelink Oldenziel13 Aug 2026 21:49 UTC
57 points
9 comments2 min readLW link

Some Ways I Think About Eval­u­at­ing Grant Applications

sarahconstantin13 Aug 2026 21:30 UTC
61 points
0 comments10 min readLW link
(sarahconstantin.substack.com)

The Left Should Start Tak­ing AI Ca­pa­bil­ities Seriously

Alexei G13 Aug 2026 21:03 UTC
−13 points
0 comments9 min readLW link
(www.onethousandmeans.com)

Is Align­ment Even Falsifi­able? Mid­dle Align­ment, An Align­ment Tax­on­omy, and Break­ing The Prob­lem Down Into Steps

Savannah Harlan13 Aug 2026 20:57 UTC
5 points
0 comments22 min readLW link

Con­crete Gen­er­al­ist Pro­jects in AI Safety (and how to do them)

13 Aug 2026 19:26 UTC
25 points
2 comments6 min readLW link
(forum.effectivealtruism.org)

Com­par­ing Congress’s Two AI Emer­gency Shut­down Mechanisms

Philip Dowdell13 Aug 2026 18:53 UTC
13 points
0 comments8 min readLW link

Is biose­cu­rity over­sat­u­rated?

Master Chief13 Aug 2026 18:13 UTC
12 points
1 comment1 min readLW link

How to An­swer a Ques­tion Without An­swer­ing The Question

Kabir Kumar13 Aug 2026 17:06 UTC
37 points
3 comments1 min readLW link

How My Stu­dents Think About AI

dvd13 Aug 2026 16:56 UTC
774 points
110 comments12 min readLW link

AI #181: As­tra Goes Cy­ber Critical

Zvi13 Aug 2026 15:20 UTC
39 points
1 comment52 min readLW link
(thezvi.wordpress.com)

Au­to­mated al­ign­ment runs are hard to study!

13 Aug 2026 15:04 UTC
68 points
4 comments9 min readLW link

Longter­mism Seems Like A Religion

James Brobin13 Aug 2026 14:43 UTC
2 points
8 comments3 min readLW link

Defense Against the De­cep­tive Arts

Kabir Kumar13 Aug 2026 14:21 UTC
11 points
11 comments1 min readLW link

Some prob­lems in de­ci­sion the­ory in­cor­rectly pre­con­di­tion on policy.

Canaletto13 Aug 2026 12:04 UTC
10 points
4 comments3 min readLW link

Pat­terns and prob­lems in emerg­ing mul­ti­a­gent sys­tems (An­thropic, Fron­tier Red Team)

Julian Bradshaw13 Aug 2026 4:10 UTC
42 points
3 comments1 min readLW link
(www.anthropic.com)

The Descen­ders and the Ab­sorbers: Two Per­spec­tives on Deep Learning

larry-dial13 Aug 2026 4:02 UTC
11 points
0 comments3 min readLW link

OC ACXLW Meetup #118 — The 40% Prob­lem & The Field That Re­fused to Die

Michael Michalchik13 Aug 2026 3:25 UTC
1 point
0 comments9 min readLW link

Mea­sur­ing Ac­ti­va­tion Con­trol in LLMs

13 Aug 2026 2:04 UTC
50 points
8 comments7 min readLW link

How Valuable are BOTECs?

Roman Ross13 Aug 2026 0:31 UTC
9 points
0 comments8 min readLW link

Mea­sur­ing Eval Aware­ness: The Real­ism Win Rate is Fragile

13 Aug 2026 0:26 UTC
14 points
3 comments4 min readLW link

What hap­pened when I tried to be vegan

finitude13 Aug 2026 0:21 UTC
44 points
2 comments5 min readLW link

LLMs have the ca­pac­ity for self-im­posed steganography

Edward Cant13 Aug 2026 0:16 UTC
7 points
0 comments7 min readLW link

Free will is like temperature

Optimization Process12 Aug 2026 23:32 UTC
47 points
10 comments1 min readLW link

Im­pact mar­kets made concrete

Carol N12 Aug 2026 22:45 UTC
6 points
0 comments6 min readLW link
(manifund.substack.com)

Un­block­ing AI’s Con­tinual Learn­ing: Hints From How Hu­mans Learn

nimakeivan12 Aug 2026 18:28 UTC
3 points
23 comments10 min readLW link

Monthly Roundup #45: Au­gust 2026

Zvi12 Aug 2026 17:50 UTC
24 points
1 comment20 min readLW link
(thezvi.wordpress.com)

Find­ing the Seams of Perception

jimmy12 Aug 2026 17:33 UTC
30 points
0 comments2 min readLW link
(beneathpsychology.com)

In­tro­duc­ing the Con­cep­tual Rea­son­ing Index

12 Aug 2026 17:08 UTC
83 points
16 comments5 min readLW link
(alignment.anthropic.com)

One at­ten­tion head car­ries knight forks in a chess trans­former, and here’s a new toolkit that found it.

David Litman12 Aug 2026 16:48 UTC
8 points
0 comments1 min readLW link
(github.com)

The Psy­chol­ogy of Cope: Ra­tion­al­ity Is Not Re­v­ersed Irrationality

Chris_Leong12 Aug 2026 12:20 UTC
27 points
12 comments3 min readLW link

De­mon Safety

Ben Pace12 Aug 2026 11:55 UTC
53 points
5 comments1 min readLW link

The Clo­sure of the In­ter­net (Re­search Linkpost)

Dean Valentine12 Aug 2026 8:22 UTC
34 points
4 comments1 min readLW link
(arctotherium.substack.com)

AI swarms are start­ing to pose in­di­rect takeover risk

12 Aug 2026 5:05 UTC
138 points
6 comments10 min readLW link

Did the al­ign­ment com­mu­nity un­der­es­ti­mate its power?

StanislavKrym12 Aug 2026 2:56 UTC
23 points
0 comments10 min readLW link

The Age of Pluribus: One Con­sul­tant for Everyone

Dorothy Gale12 Aug 2026 1:39 UTC
9 points
0 comments4 min readLW link

We should con­sider how long mon­i­tor­ing is re­li­able for dur­ing RL

lachlan on a boat12 Aug 2026 1:22 UTC
21 points
1 comment4 min readLW link

When (and when not) LLMs can ver­bal­ize aware­ness of J-Space con­cept in­jec­tions—Ini­tial results

Ethan Garcia12 Aug 2026 1:21 UTC
8 points
0 comments15 min readLW link
(e-m-garcia.github.io)

Ar­gu­ments for and against (me) drop­ping out

hersheys12 Aug 2026 1:12 UTC
30 points
3 comments8 min readLW link

Pa­tient Zero

LoopGameScrollMonkey12 Aug 2026 1:10 UTC
17 points
0 comments12 min readLW link

An any­time al­gorithm for mix­ing the com­putable measures

Cole Wyeth12 Aug 2026 0:57 UTC
26 points
0 comments4 min readLW link

Var­i­ous Reflec­tions About What Hap­pened With OpenAI’s In­ter­nal Models

Zvi11 Aug 2026 22:20 UTC
63 points
1 comment30 min readLW link
(thezvi.wordpress.com)

Claude Opus 5 Just Beat My Text-Based Ad­ven­ture Game Benchmark

derelict543211 Aug 2026 20:19 UTC
22 points
4 comments5 min readLW link

[We­bi­nar] Why AI Safety is a Cap­i­tal Allo­ca­tion Prob­lem

Schizoid Rentoid11 Aug 2026 19:27 UTC
2 points
0 comments1 min readLW link

Misal­igned AIs could use kil­ler robots to take over

11 Aug 2026 19:03 UTC
134 points
10 comments6 min readLW link
(turntrout.com)

Mea­sur­ing Spu­ri­ous Cor­re­la­tions with Fea­ture Strength

egan11 Aug 2026 18:05 UTC
36 points
1 comment10 min readLW link

AI gov­er­nance work needs much bet­ter monitoring

jackultraphil11 Aug 2026 17:42 UTC
10 points
0 comments8 min readLW link
(fundinganthropalypse.com)

LLMs Are Start­ing To No­tice­ably Ac­cel­er­ate Our Work

johnswentworth11 Aug 2026 17:06 UTC
251 points
19 comments2 min readLW link

Soft­ware Is Hard

cylonator11 Aug 2026 16:26 UTC
5 points
0 comments1 min readLW link

How risky would it be to make pow­er­ful AI obey one or a few peo­ple?

11 Aug 2026 16:15 UTC
93 points
24 comments8 min readLW link

Ex­treme con­cen­tra­tion of power over ASI has non-ob­vi­ous advantages

Seth Herd11 Aug 2026 16:13 UTC
43 points
4 comments19 min readLW link