RSS

Challenge: Hand cod­ing weights for effi­cient se­quence memorisation

23 Jul 2026 18:05 UTC
30 points
0 comments18 min readLW link

AI Re­searchers Don’t Un­der­stand the State

Alex Amadori23 Jul 2026 17:58 UTC
6 points
0 comments1 min readLW link
(x.com)

The OpenAI/​Hug­ging­face in­ci­dent | Red­wood Re­search pod­cast epi­sode 2

23 Jul 2026 17:56 UTC
42 points
0 comments2 min readLW link

V&V takes on OpenAI’s long-hori­zon incidents

Yoav Hollander23 Jul 2026 16:51 UTC
19 points
0 comments4 min readLW link
(blog.foretellix.com)

Want­ing Crooked Lines

Davey Morse23 Jul 2026 14:33 UTC
12 points
0 comments4 min readLW link

(3/​3) The Dangers of ASI

Eigenbraid23 Jul 2026 13:32 UTC
10 points
0 comments3 min readLW link

Sleep­ing Beauty as a Mind Killer

avturchin23 Jul 2026 11:04 UTC
9 points
9 comments7 min readLW link

Get You a Model: If You’re Not Pay­ing, You’re Miss­ing 90% Of Improvements

Czynski23 Jul 2026 4:16 UTC
12 points
2 comments1 min readLW link
(dangeroussincerity.substack.com)

Are we ex­is­ten­tially threat­ened by the type of AI mis­al­ign­ment seen in the OpenAI Hug­ging Face at­tack?

23 Jul 2026 3:40 UTC
165 points
6 comments5 min readLW link

LLMOSES

Davey Morse23 Jul 2026 2:35 UTC
−5 points
0 comments3 min readLW link

Can an LLM make a fea­ture-length movie on its own?

Josh Snider22 Jul 2026 20:13 UTC
34 points
5 comments4 min readLW link

Will al­most all fu­ture com­pa­nies even­tu­ally be founded and run by au­tonomous AIs?

Steven Byrnes22 Jul 2026 20:04 UTC
56 points
1 comment8 min readLW link

We can­not simu­late AI se­cu­rity research

Jafar Isbarov22 Jul 2026 19:15 UTC
8 points
0 comments5 min readLW link

The Con­jec­ture of Strong Subjectivity

D.Schetselaar22 Jul 2026 17:43 UTC
10 points
12 comments1 min readLW link
(philpapers.org)

The Best AI Bill Congress Hasn’t In­tro­duced Yet

dan.parshall22 Jul 2026 16:19 UTC
11 points
0 comments11 min readLW link

Your AIs don’t do what you want. This is re­ally bad

Kaustubh Kislay22 Jul 2026 15:56 UTC
16 points
0 comments3 min readLW link
(rewardhacking.org)

Models don’t seem to be dishon­est in the way hu­mans are

22 Jul 2026 15:32 UTC
45 points
3 comments9 min readLW link

Com­ment on Mea­sur­ing Re­ward-Seek­ing by In­still­ing Con­trastive Beliefs pa­per from mechanis­tic in­ter­pretabil­ity perspective

Burny22 Jul 2026 14:58 UTC
16 points
0 comments6 min readLW link

How do neu­ral net­works regress cu­bic polyno­mi­als? Ap­par­ently, they use a trick in­vented in Milan 500 years ago

enricobottazzi22 Jul 2026 14:42 UTC
28 points
0 comments16 min readLW link

(2/​3) The Dangers of AGI

Eigenbraid22 Jul 2026 13:57 UTC
8 points
0 comments9 min readLW link