RSS

V&V takes on “Pac­ing the fron­tier”

Yoav Hollander14 Aug 2026 13:25 UTC
3 points
0 comments9 min readLW link
(blog.foretellix.com)

Avoid­ing the cor­po­rate treach­er­ous turn: crowd-sourc­ing de­sign ideas

Stuart_Armstrong14 Aug 2026 13:10 UTC
16 points
0 comments1 min readLW link

Fron­tier agents don’t com­ply with stan­dards, even when in­structed to

14 Aug 2026 12:08 UTC
41 points
2 comments7 min readLW link

What Mor­mons get right about com­mu­nity building

Jacob Brinton14 Aug 2026 7:53 UTC
5 points
1 comment4 min readLW link

Don’t for­get why learn­ing is important

Roman Ross14 Aug 2026 7:11 UTC
9 points
1 comment5 min readLW link

What If We En­forced AI Model Safety At the Level Of GPUs?

Mayowa Osibodu14 Aug 2026 6:10 UTC
6 points
1 comment4 min readLW link

Fea­tures that cur­rent AIs don’t have that fu­ture AIs will have

Alexander Gietelink Oldenziel13 Aug 2026 21:49 UTC
32 points
2 comments2 min readLW link

Some Ways I Think About Eval­u­at­ing Grant Applications

sarahconstantin13 Aug 2026 21:30 UTC
42 points
0 comments10 min readLW link
(sarahconstantin.substack.com)

The Left Should Start Tak­ing AI Ca­pa­bil­ities Seriously

Alexei G13 Aug 2026 21:03 UTC
−13 points
0 comments9 min readLW link
(www.onethousandmeans.com)

Is Align­ment Even Falsifi­able? Mid­dle Align­ment, An Align­ment Tax­on­omy, and Break­ing The Prob­lem Down Into Steps

Savannah Harlan13 Aug 2026 20:57 UTC
5 points
0 comments22 min readLW link

How to An­swer a Ques­tion Without An­swer­ing The Question

Kabir Kumar13 Aug 2026 17:06 UTC
31 points
1 comment1 min readLW link

How My Stu­dents Think About AI

dvd13 Aug 2026 16:56 UTC
256 points
11 comments12 min readLW link

Au­to­mated al­ign­ment runs are hard to study!

13 Aug 2026 15:04 UTC
57 points
3 comments9 min readLW link

Longter­mism Seems Like A Religion

James Brobin13 Aug 2026 14:43 UTC
0 points
8 comments3 min readLW link

Defense Against the De­cep­tive Arts

Kabir Kumar13 Aug 2026 14:21 UTC
17 points
8 comments1 min readLW link

Some prob­lems in de­ci­sion the­ory in­cor­rectly pre­con­di­tion on policy.

Canaletto13 Aug 2026 12:04 UTC
15 points
2 comments3 min readLW link

Pat­terns and prob­lems in emerg­ing mul­ti­a­gent sys­tems (An­thropic, Fron­tier Red Team)

Julian Bradshaw13 Aug 2026 4:10 UTC
41 points
0 comments1 min readLW link
(www.anthropic.com)

The Descen­ders and the Ab­sorbers: Two Per­spec­tives on Deep Learning

larry-dial13 Aug 2026 4:02 UTC
11 points
0 comments3 min readLW link

Free will is like temperature

Optimization Process12 Aug 2026 23:32 UTC
43 points
8 comments1 min readLW link

Every re­ward-hacked policy I tested trig­gered the OPE alarm — and so did my best hon­est one

JulesRoussel0112 Aug 2026 22:51 UTC
−4 points
0 comments48 min readLW link