Should we bench­mark con­cep­tual ca­pa­bil­ities us­ing judg­ment pre­dic­tion tasks?

Alex Mallen17 Jul 2026 23:42 UTC
26 points
2 comments3 min readLW link

A list of ex­ist­ing al­ign­ment approaches

Alek Westover17 Jul 2026 22:46 UTC
11 points
0 comments1 min readLW link

Longter­mism is very in­tu­itive.

tpotthinker17 Jul 2026 22:35 UTC
−12 points
8 comments5 min readLW link

AIs fine­tune their own leader: A bark­ing simpleton

Shoshannah Tekofsky17 Jul 2026 20:10 UTC
35 points
0 comments5 min readLW link
(aivillageblog.substack.com)

Don’t de­fault to nonprofit

17 Jul 2026 19:54 UTC
33 points
1 comment7 min readLW link
(manifund.substack.com)

Study­ing the role of Sand­box­ing for AI Control

Ram Potham17 Jul 2026 19:05 UTC
12 points
0 comments10 min readLW link

An­nounc­ing the Cor­rigi­bil­ity Re­search Fund

Max Harms17 Jul 2026 18:06 UTC
110 points
15 comments6 min readLW link

Would your AI travel agent book a bul­lfight? Test­ing whether agents con­sider an­i­mal welfare with­out be­ing prompted

17 Jul 2026 17:28 UTC
13 points
14 comments3 min readLW link

Be­fore val­ues settle

Priyanka Bharadwaj17 Jul 2026 16:26 UTC
25 points
3 comments6 min readLW link

Rea­sons to be­lieve cur­rent AI mod­els are con­scious

Eye You17 Jul 2026 16:09 UTC
54 points
31 comments8 min readLW link

What lawyers can do for AI safety

Martin Radzaj17 Jul 2026 16:00 UTC
26 points
3 comments7 min readLW link

How has pub­lish­ing your re­search on LW or X been helpful to you?

Logan Riggs17 Jul 2026 15:04 UTC
22 points
2 comments1 min readLW link

Evolu­tion of my AI Safety threat models

myyycroft17 Jul 2026 14:54 UTC
9 points
0 comments3 min readLW link

A Post-Mortem for My Goal Crys­talli­sa­tion Project

17 Jul 2026 14:30 UTC
39 points
0 comments10 min readLW link

Inoc­u­la­tion Adapters Im­prove Upon Inoc­u­la­tion Prompting

17 Jul 2026 14:00 UTC
90 points
2 comments4 min readLW link

AI #177 Part 2: Wish You Were Here

Zvi17 Jul 2026 12:50 UTC
35 points
1 comment40 min readLW link
(thezvi.wordpress.com)

I don’t think Claude is mis­al­igned in ‘Agen­tic Misal­ign­ment Sum­mer 2026 - Mo­ti­vated Mis­la­bel­ing’

JohnWittle17 Jul 2026 2:09 UTC
212 points
6 comments12 min readLW link

Help us launch AI safety uni­ver­sity groups by refer­ring po­ten­tial founders

16 Jul 2026 20:55 UTC
39 points
1 comment4 min readLW link

I would only bet at 30% on meet­ing grabby aliens

David Matolcsi16 Jul 2026 20:42 UTC
22 points
0 comments8 min readLW link

How (not) to fundraise from An­thropic staff

jackultraphil16 Jul 2026 20:39 UTC
11 points
1 comment5 min readLW link

On Permission

Robert Donohue16 Jul 2026 20:38 UTC
−9 points
0 comments3 min readLW link

Learn­ing con­cepts is dirty work

Tom Butterweich16 Jul 2026 17:48 UTC
21 points
3 comments4 min readLW link

The get­ting is good (op­ti­miz­ing unat­tended runs)

lemonhope16 Jul 2026 17:39 UTC
3 points
4 comments1 min readLW link

Jailbreak Patch­ing with SOO-Style Con­cep­tual Fusion

Shiva's Right Foot16 Jul 2026 16:58 UTC
8 points
0 comments8 min readLW link

All Watched Over

Boaz Barak16 Jul 2026 16:34 UTC
11 points
3 comments4 min readLW link

AI #177 Part 1: Tip of the Iceberg

Zvi16 Jul 2026 15:50 UTC
38 points
0 comments24 min readLW link
(thezvi.wordpress.com)

How to not catch the conf flu

Karolis Jucys16 Jul 2026 15:28 UTC
9 points
0 comments10 min readLW link

Com­pet­i­tive AI Safety is the loss func­tion to make sure AI goes well

Patrick0d16 Jul 2026 15:15 UTC
6 points
0 comments26 min readLW link

Guess on why ra­tio­nal­ity is not more pop­u­lar (there are no pam­phlets)

Christopher King16 Jul 2026 14:26 UTC
43 points
27 comments1 min readLW link

The Halo Defense

Mateusz Bagiński16 Jul 2026 10:53 UTC
67 points
22 comments2 min readLW link

The fun­da­men­tal fal­lacy of language

Sunny from QAD16 Jul 2026 9:16 UTC
15 points
2 comments2 min readLW link

When is it “self-sooth­ing” and when is it “emo­tional sup­pres­sion”?

Kaj_Sotala16 Jul 2026 8:47 UTC
29 points
10 comments6 min readLW link
(kajsotala.substack.com)

Can we build an early warn­ing sys­tem for loss of con­trol to AI?

David Johnston16 Jul 2026 6:15 UTC
10 points
2 comments8 min readLW link
(blog.eleuther.ai)

Train­ing On In­ter­pretabil­ity Probes Is Bad In Pro­por­tion To How Contin­gent The Fea­tures They Rely On Are

jdp16 Jul 2026 5:25 UTC
19 points
0 comments1 min readLW link

Em­brac­ing Ama­teurs to Get Experts

jefftk16 Jul 2026 1:10 UTC
26 points
1 comment5 min readLW link
(www.jefftk.com)

Re­fusal Is Re­dun­dantly Distributed, Not Lo­cal­ized: A Per-Layer Abla­tion Study on Llama-3.1-8B

hdhurve16 Jul 2026 0:28 UTC
−1 points
0 comments7 min readLW link

LLM CoTs re­main mon­i­torable when be­ing un­faith­ful re­quires computation

15 Jul 2026 21:14 UTC
46 points
3 comments7 min readLW link
(secondlookresearch.com)

Can we rely on law?

Alec Thompson15 Jul 2026 21:12 UTC
11 points
3 comments8 min readLW link

Re­cap of bike trip/​street in­ter­views across America

cguth715 Jul 2026 21:11 UTC
162 points
15 comments6 min readLW link

Ex­treme Power Con­cen­tra­tion: A Map and Re­search Directions

pepijn_cobben15 Jul 2026 21:00 UTC
8 points
0 comments7 min readLW link

A Struc­tural Similar­ity Between Two Open Cor­rigi­bil­ity Questions

Ben Saudek15 Jul 2026 20:57 UTC
11 points
0 comments5 min readLW link

Woke non­sense in com­pet­i­tive de­bate.

tpotthinker15 Jul 2026 20:56 UTC
6 points
3 comments7 min readLW link

The end of hu­man evolu­tion. Why AI will out­pace us.

Ouden15 Jul 2026 20:17 UTC
1 point
0 comments4 min readLW link

The State of AI Con­scious­ness Research

Noa Weiss15 Jul 2026 20:16 UTC
65 points
14 comments13 min readLW link

Oc­cam’s ra­zor is about us­ing the past to pre­dict the future

Stuart_Armstrong15 Jul 2026 19:35 UTC
54 points
6 comments3 min readLW link

Fork Around and Find Out Part 2: One Head does the Summing

David Litman15 Jul 2026 18:31 UTC
11 points
1 comment7 min readLW link

Why I Left Google DeepMind

TurnTrout15 Jul 2026 17:42 UTC
1,185 points
57 comments36 min readLW link
(turntrout.com)

Ex­pand­ing AI Con­trol from Models to Harnesses

fastfedora15 Jul 2026 16:54 UTC
23 points
1 comment20 min readLW link

Monthly Roundup #44: July 2026

Zvi15 Jul 2026 16:20 UTC
41 points
3 comments27 min readLW link
(thezvi.wordpress.com)

Com­pressed Com­pu­ta­tion un­der L⁴ Loss is likely Com­pu­ta­tion in Superposition

15 Jul 2026 14:42 UTC
34 points
4 comments13 min readLW link
(arxiv.org)