RSS

We can­not simu­late AI se­cu­rity research

Jafar Isbarov22 Jul 2026 19:15 UTC
1 point
0 comments5 min readLW link

The Con­jec­ture of Strong Subjectivity

D.Schetselaar22 Jul 2026 17:43 UTC
3 points
0 comments1 min readLW link
(philpapers.org)

Your AIs don’t do what you want. This is re­ally bad

Kaustubh Kislay22 Jul 2026 15:56 UTC
9 points
0 comments3 min readLW link
(rewardhacking.org)

Models don’t seem to be dishon­est in the way hu­mans are

22 Jul 2026 15:32 UTC
27 points
0 comments9 min readLW link

Com­ment on Mea­sur­ing Re­ward-Seek­ing by In­still­ing Con­trastive Beliefs pa­per from mechanis­tic in­ter­pretabil­ity perspective

Burny22 Jul 2026 14:58 UTC
6 points
0 comments6 min readLW link

How do neu­ral net­works regress cu­bic polyno­mi­als? Ap­par­ently, they use a trick in­vented in Milan 500 years ago

enricobottazzi22 Jul 2026 14:42 UTC
10 points
0 comments16 min readLW link

(2/​3) The Dangers of AGI

Eigenbraid22 Jul 2026 13:57 UTC
8 points
0 comments9 min readLW link

An­nounc­ing AIXI Labs

22 Jul 2026 11:31 UTC
61 points
6 comments4 min readLW link

We should push for no-fault li­a­bil­ity for ac­tions taken by AI

Yair Halberstadt22 Jul 2026 9:59 UTC
80 points
33 comments2 min readLW link

OpenAI and Hug­ging Face part­ner to ad­dress se­cu­rity in­ci­dent dur­ing model evaluation

Matrice Jacobine22 Jul 2026 6:30 UTC
25 points
1 comment1 min readLW link
(openai.com)

WeirdChat: A cat­a­log of un­ex­pected AI be­hav­iors, dis­cov­ered automatically

neilchowdhury21 Jul 2026 21:12 UTC
45 points
0 comments10 min readLW link
(transluce.org)

Steer­ing Black­mail Through a Model’s “Emo­tional State”

21 Jul 2026 16:03 UTC
8 points
0 comments8 min readLW link

Mea­sur­ing Re­ward-Seek­ing via Con­trastive Belief Updates

21 Jul 2026 15:22 UTC
74 points
4 comments12 min readLW link
(rewardseeking.ai)

Mea­sur­ing Re­ward-Seek­ing by In­still­ing Con­trastive Beliefs

papetoast21 Jul 2026 15:11 UTC
12 points
0 comments1 min readLW link
(alignment.openai.com)

11 Open Em­piri­cal Prob­lems in Re­ward-Seeking

21 Jul 2026 15:08 UTC
52 points
0 comments8 min readLW link

Differ­en­tial ac­cel­er­a­tion of al­ign­ment-rele­vant ca­pa­bil­ities is a bad bet

Zephaniah Roe21 Jul 2026 14:04 UTC
72 points
0 comments8 min readLW link

Epistemics and Co­or­di­na­tion: It’s com­pli­cated!

Raymond Douglas21 Jul 2026 12:23 UTC
35 points
13 comments8 min readLW link

(1/​3) The Dangers of LLMs

Eigenbraid21 Jul 2026 12:01 UTC
8 points
0 comments6 min readLW link

I ran the stan­dard AI lit­mus tests on my two tod­dlers (yep)

Carlo Valenti21 Jul 2026 11:38 UTC
36 points
5 comments4 min readLW link

Ad­der­all Tol­er­ance: Much More Than You Wanted To Know

Kurt H. Pieper20 Jul 2026 21:55 UTC
76 points
5 comments4 min readLW link
(kurthpieper.substack.com)