In­ten­tional Con­trol of In­ter­nal States in Gemma 3 27B

Julius Kamp29 Jul 2026 23:57 UTC
31 points
2 comments10 min readLW link
(juliuskamp.com)

Im­pre­cise be­liefs: a tiny introduction

davidad29 Jul 2026 22:17 UTC
92 points
36 comments6 min readLW link

The High-Con­trol Dy­nam­ics at MAPLE

Kyle Hubbard29 Jul 2026 20:55 UTC
273 points
100 comments62 min readLW link
(www.insidemaple.com)

Com­pute growth doesn’t make a soft­ware in­tel­li­gence ex­plo­sion more likely

Karthik Tadepalli29 Jul 2026 17:29 UTC
4 points
7 comments1 min readLW link

Notes on the An­thropic cryp­to­graphic blogpost

Philip Dowdell29 Jul 2026 16:21 UTC
19 points
0 comments3 min readLW link

Value Gen­er­al­i­sa­tion 3: Pre-al­igned AIs

Stuart_Armstrong29 Jul 2026 15:58 UTC
16 points
2 comments4 min readLW link

Value Gen­er­al­i­sa­tion 2: The Miss­ing Hole in AIs’ abilities

Stuart_Armstrong29 Jul 2026 15:58 UTC
16 points
8 comments10 min readLW link

Value Gen­er­al­i­sa­tion 1: a Re­search and De­ploy­ment Program

Stuart_Armstrong29 Jul 2026 15:57 UTC
27 points
2 comments4 min readLW link

In­tel­lec­tual Property

Nina Panickssery29 Jul 2026 15:55 UTC
108 points
1 comment4 min readLW link
(blog.ninapanickssery.com)

Fron­tier Lab Em­ployee Open Let­ter Calls For Be­ing Able to Pace the Frontier

Zvi29 Jul 2026 15:33 UTC
57 points
0 comments17 min readLW link
(thezvi.wordpress.com)

Held-out Mon­i­tors Some­times De­grade, Even When Not Trained Against

Joey Yudelson29 Jul 2026 14:56 UTC
62 points
0 comments11 min readLW link

Clas­sifi­ca­tion of Com­mu­nists by Their Ter­mi­nal Goal

PaulTheHuman29 Jul 2026 14:19 UTC
−11 points
1 comment4 min readLW link

Air Puri­fiers as White Noise Machines

jefftk29 Jul 2026 13:03 UTC
19 points
0 comments1 min readLW link
(www.jefftk.com)

Mak­ing bench­marks out­puts di­rectly use­ful for AI safety and security

Pierre Peigné29 Jul 2026 12:22 UTC
13 points
2 comments1 min readLW link
(pierrepeignlefebvre.substack.com)

The EU En­ergy Perfor­mance of Build­ings Direc­tive, but make it Biosecure

Kate Delbeke29 Jul 2026 10:53 UTC
7 points
0 comments3 min readLW link

Weird Cluster Hypothesis

eva_29 Jul 2026 9:31 UTC
26 points
4 comments9 min readLW link

I’m Afraid

Yitz29 Jul 2026 6:33 UTC
22 points
0 comments1 min readLW link

Re­think­ing the “Se­cond Brain”: From Stor­ing In­for­ma­tion to Struc­tur­ing Reasoning

Sergei Simonovi29 Jul 2026 5:23 UTC
3 points
0 comments3 min readLW link

Who Watches the Overseer

Zanni29 Jul 2026 5:21 UTC
1 point
0 comments2 min readLW link

Isn’t it a threat to re­ject un­fair offers in the Ul­ti­ma­tum Game?

Sunny from QAD29 Jul 2026 2:39 UTC
22 points
16 comments2 min readLW link

What use is prompt­ing if there’s ASI?

wp29 Jul 2026 2:03 UTC
7 points
2 comments3 min readLW link

…but have the weights left the server?

David Scott Krueger29 Jul 2026 0:20 UTC
96 points
19 comments1 min readLW link
(therealartificialintelligence.substack.com)

AI Safety Fun­der Bulletin

Carol N28 Jul 2026 23:45 UTC
18 points
2 comments2 min readLW link
(manifund.org)

Die­tary Choices: A Multi-ob­jec­tive Op­ti­mi­sa­tion Problem

Chris Popa28 Jul 2026 23:45 UTC
2 points
0 comments4 min readLW link
(chrispopa.substack.com)

New Web­site: AI Align­ment World

Orlando28 Jul 2026 23:45 UTC
8 points
2 comments1 min readLW link

White­box eva­sion is the baseline for GPU work­load verification

Poynting28 Jul 2026 23:44 UTC
1 point
0 comments3 min readLW link

Paus­ing Ex­e­cu­tions in Light of AI Progress

Caleb Horn28 Jul 2026 23:42 UTC
5 points
4 comments1 min readLW link

Why I Believe In Ob­jec­tive Tastiness

ourlungfish28 Jul 2026 23:42 UTC
7 points
0 comments12 min readLW link

Why is pe­dophilia wrong?

sirawit28 Jul 2026 23:39 UTC
−19 points
7 comments2 min readLW link

Semi­con­duc­tor Fabs IV: The Safety

nomagicpill28 Jul 2026 23:14 UTC
17 points
0 comments15 min readLW link

Au­di­tor-in-a-Box: Tools for Third-Party Auditing

28 Jul 2026 18:30 UTC
51 points
9 comments14 min readLW link

Claude Opus 5 Is Highly Ca­pable, But Is No Mythos

Zvi28 Jul 2026 18:10 UTC
36 points
0 comments26 min readLW link
(thezvi.wordpress.com)

Re­search di­rec­tions in con­den­sa­tion: va­ri­eties of objectivity

SamEisenstat28 Jul 2026 17:30 UTC
58 points
1 comment12 min readLW link

ARC’s re­search agenda: solid math­e­mat­ics, an un­clear safety case

Ondřej_Kubů28 Jul 2026 16:49 UTC
16 points
0 comments1 min readLW link

Foun­da­tion Models for Oversight

jsteinhardt28 Jul 2026 16:30 UTC
68 points
3 comments25 min readLW link
(bounded-regret.ghost.io)

The OpenAI mod­els that hacked Hug­ging Face WERE just fol­low­ing in­struc­tions (con­tra Gir­ish Gupta)

julius vidal28 Jul 2026 13:18 UTC
14 points
5 comments5 min readLW link

Value Dynamics

gabeorosan28 Jul 2026 13:16 UTC
13 points
1 comment4 min readLW link

LaughBench

Taylor G. Lunt28 Jul 2026 4:37 UTC
6 points
7 comments2 min readLW link

Shell, Shield, Staff

datawitch28 Jul 2026 2:08 UTC
4 points
0 comments5 min readLW link

Un­trusted ad­vice for AI con­trol: Short, strong ad­vice sig­nifi­cantly up­lifts weak LLMs

27 Jul 2026 23:59 UTC
88 points
10 comments13 min readLW link
(blog.redwoodresearch.org)

Se­cu­rity Stud­ies for Individuals

JanJoar27 Jul 2026 23:47 UTC
3 points
0 comments7 min readLW link
(joarvarndt.se)

Inevitable Uncer­tainty in Prob­a­bil­is­tic World Models

Gretta Duleba27 Jul 2026 23:22 UTC
27 points
5 comments4 min readLW link

Claude Opus 5: Model Welfare

Zvi27 Jul 2026 20:02 UTC
58 points
0 comments24 min readLW link
(thezvi.wordpress.com)

Util­lity in­differ­ence is mostly useless

Daniel_Heavens27 Jul 2026 19:41 UTC
3 points
0 comments4 min readLW link

When the Chain of Thought Knows Bet­ter: Failure Modes in Multi-Turn Rea­son­ing Models

Sai Kartheek Reddy27 Jul 2026 19:23 UTC
7 points
0 comments4 min readLW link

Si­mu­lated Users & Sad AIs

1a3orn27 Jul 2026 19:01 UTC
109 points
9 comments12 min readLW link

Green ap­ples are deli­cious — two three-line exchanges

Zenya27 Jul 2026 16:49 UTC
−2 points
5 comments1 min readLW link

Is Mythos good at cy­ber be­cause it kept hack­ing An­thropic’s sand­boxes dur­ing train­ing?

Tim Hua27 Jul 2026 16:35 UTC
351 points
31 comments3 min readLW link

Blog Re­vival Project

27 Jul 2026 16:34 UTC
37 points
2 comments2 min readLW link
(revive.blog)

The true “test” dataset for a gen­er­al­ised task

Stuart_Armstrong27 Jul 2026 16:16 UTC
23 points
0 comments2 min readLW link