AI Safety Fun­der Bulletin

Carol N28 Jul 2026 23:45 UTC
18 points
2 comments2 min readLW link
(manifund.org)

Die­tary Choices: A Multi-ob­jec­tive Op­ti­mi­sa­tion Problem

Chris Popa28 Jul 2026 23:45 UTC
2 points
0 comments4 min readLW link
(chrispopa.substack.com)

New Web­site: AI Align­ment World

Orlando28 Jul 2026 23:45 UTC
8 points
2 comments1 min readLW link

White­box eva­sion is the baseline for GPU work­load verification

Poynting28 Jul 2026 23:44 UTC
1 point
0 comments3 min readLW link

Paus­ing Ex­e­cu­tions in Light of AI Progress

Caleb Horn28 Jul 2026 23:42 UTC
5 points
4 comments1 min readLW link

Why I Believe In Ob­jec­tive Tastiness

ourlungfish28 Jul 2026 23:42 UTC
7 points
0 comments12 min readLW link

Why is pe­dophilia wrong?

sirawit28 Jul 2026 23:39 UTC
−19 points
7 comments2 min readLW link

Semi­con­duc­tor Fabs IV: The Safety

nomagicpill28 Jul 2026 23:14 UTC
17 points
0 comments15 min readLW link

Au­di­tor-in-a-Box: Tools for Third-Party Auditing

28 Jul 2026 18:30 UTC
51 points
9 comments14 min readLW link

Claude Opus 5 Is Highly Ca­pable, But Is No Mythos

Zvi28 Jul 2026 18:10 UTC
36 points
0 comments26 min readLW link
(thezvi.wordpress.com)

Re­search di­rec­tions in con­den­sa­tion: va­ri­eties of objectivity

SamEisenstat28 Jul 2026 17:30 UTC
58 points
1 comment12 min readLW link

ARC’s re­search agenda: solid math­e­mat­ics, an un­clear safety case

Ondřej_Kubů28 Jul 2026 16:49 UTC
16 points
0 comments1 min readLW link

Foun­da­tion Models for Oversight

jsteinhardt28 Jul 2026 16:30 UTC
68 points
3 comments25 min readLW link
(bounded-regret.ghost.io)

The OpenAI mod­els that hacked Hug­ging Face WERE just fol­low­ing in­struc­tions (con­tra Gir­ish Gupta)

julius vidal28 Jul 2026 13:18 UTC
14 points
5 comments5 min readLW link

Value Dynamics

gabeorosan28 Jul 2026 13:16 UTC
13 points
1 comment4 min readLW link

LaughBench

Taylor G. Lunt28 Jul 2026 4:37 UTC
6 points
7 comments2 min readLW link

Shell, Shield, Staff

datawitch28 Jul 2026 2:08 UTC
4 points
0 comments5 min readLW link

Un­trusted ad­vice for AI con­trol: Short, strong ad­vice sig­nifi­cantly up­lifts weak LLMs

27 Jul 2026 23:59 UTC
88 points
10 comments13 min readLW link
(blog.redwoodresearch.org)

Se­cu­rity Stud­ies for Individuals

JanJoar27 Jul 2026 23:47 UTC
3 points
0 comments7 min readLW link
(joarvarndt.se)

Inevitable Uncer­tainty in Prob­a­bil­is­tic World Models

Gretta Duleba27 Jul 2026 23:22 UTC
27 points
5 comments4 min readLW link

Claude Opus 5: Model Welfare

Zvi27 Jul 2026 20:02 UTC
58 points
0 comments24 min readLW link
(thezvi.wordpress.com)

Util­lity in­differ­ence is mostly useless

Daniel_Heavens27 Jul 2026 19:41 UTC
3 points
0 comments4 min readLW link

When the Chain of Thought Knows Bet­ter: Failure Modes in Multi-Turn Rea­son­ing Models

Sai Kartheek Reddy27 Jul 2026 19:23 UTC
7 points
0 comments4 min readLW link

Si­mu­lated Users & Sad AIs

1a3orn27 Jul 2026 19:01 UTC
109 points
9 comments12 min readLW link

Green ap­ples are deli­cious — two three-line exchanges

Zenya27 Jul 2026 16:49 UTC
−2 points
5 comments1 min readLW link

Is Mythos good at cy­ber be­cause it kept hack­ing An­thropic’s sand­boxes dur­ing train­ing?

Tim Hua27 Jul 2026 16:35 UTC
351 points
31 comments3 min readLW link

Blog Re­vival Project

27 Jul 2026 16:34 UTC
37 points
2 comments2 min readLW link
(revive.blog)

The true “test” dataset for a gen­er­al­ised task

Stuart_Armstrong27 Jul 2026 16:16 UTC
23 points
0 comments2 min readLW link

You (Yes, You) Need A Fe­bru­ary 2020 Check­list for AI Policy

davekasten27 Jul 2026 15:51 UTC
272 points
17 comments3 min readLW link

Fine-Tun­ing, The Hier­ar­chy Prob­lem, and What Neu­trons Tell Us About God

Mikewins27 Jul 2026 15:51 UTC
−4 points
2 comments1 min readLW link

Quadrillion Param Costs: KV Cache, Con­text Length, Fron­tier Margins

Vladimir_Nesov27 Jul 2026 15:08 UTC
77 points
7 comments25 min readLW link

MSE loss does not gen­er­ate superposition

27 Jul 2026 15:04 UTC
27 points
3 comments12 min readLW link

RL & search is a ter­rify­ing way to build AGI (an FAQ)

Steven Byrnes27 Jul 2026 14:50 UTC
149 points
16 comments14 min readLW link

PIRAMID: Progress and Plans

27 Jul 2026 13:08 UTC
48 points
2 comments14 min readLW link

Em dashes are fuck­ing amazing

eleweek27 Jul 2026 11:43 UTC
−6 points
0 comments3 min readLW link
(psychotechnology.substack.com)

My AI Slav­ery In­ter­views Are Cen­sored On LW By Default

JenniferRM27 Jul 2026 8:42 UTC
54 points
43 comments18 min readLW link

Five things I learned from 630 days of writ­ing online

domelian27 Jul 2026 8:10 UTC
3 points
0 comments1 min readLW link
(domelian.substack.com)

Shan­non­ian Crit­i­cism: In­for­ma­tion Ar­chi­tec­ture and Sylvia Plath’s ‘Daddy’

Stevie Miller27 Jul 2026 8:01 UTC
6 points
1 comment14 min readLW link

Can we teach a model to en­code a se­man­tic fea­ture on a cho­sen man­i­fold in just three chan­nels?

Phu Hoang27 Jul 2026 8:00 UTC
5 points
0 comments17 min readLW link

Multi-Turn Drift In­creases Scheming

27 Jul 2026 7:54 UTC
14 points
2 comments9 min readLW link

Does ChatGPT re­ally have a strong left-wing bias?

John-Clark Levin27 Jul 2026 6:25 UTC
9 points
4 comments7 min readLW link

You don’t need er­ror nodes, you need bet­ter features

Evan Lloyd27 Jul 2026 4:14 UTC
26 points
0 comments57 min readLW link

A clar­ifi­ca­tion on cel­e­brat­ing victory

KatjaGrace27 Jul 2026 4:02 UTC
62 points
3 comments1 min readLW link
(worldspiritsockpuppet.substack.com)

Can­ter­bury Coun­try Dance Orches­tra Liner Notes

jefftk27 Jul 2026 3:22 UTC
9 points
0 comments8 min readLW link
(www.jefftk.com)

OpenAI’s my­opia just keeps caus­ing al­ign­ment problems

Fiora Starlight27 Jul 2026 3:01 UTC
200 points
27 comments10 min readLW link

At the end of the day, my slaves are just a tool

jesseduffield27 Jul 2026 0:43 UTC
30 points
2 comments2 min readLW link

The AI that fights for your place in the world

Akshay Iyer26 Jul 2026 22:05 UTC
9 points
3 comments4 min readLW link

What Hap­pens When a Col­lu­sion Probe Only Finds a Thin Sig­nal?

26 Jul 2026 22:04 UTC
9 points
0 comments9 min readLW link

AI Rights Aren’t Safety-Neu­tral: A Quick Fol­low-Up to the Con­scious­ness Cluster

adorable_hamster26 Jul 2026 22:03 UTC
16 points
2 comments7 min readLW link

More On An In­ter­nal OpenAI Model Hack­ing Into HuggingFace

Zvi26 Jul 2026 19:22 UTC
98 points
3 comments24 min readLW link
(thezvi.wordpress.com)