RSS

AI Safety

TagLast edit: 8 Mar 2026 9:49 UTC by JeaniceK

An­nounc­ing AIXI Labs

22 Jul 2026 11:31 UTC
95 points
9 comments4 min readLW link

Syn­thetic Scal­able Oversight

14 Jul 2026 14:32 UTC
16 points
0 comments15 min readLW link

Open-source LLMs ad­minister max­i­mum elec­tric shocks in a Mil­gram-like obe­di­ence experiment

7 Jul 2026 20:05 UTC
8 points
0 comments25 min readLW link
(arxiv.org)

[pa­per] Train­ing on Doc­u­ments About Mon­i­tor­ing Leads to CoT Obfuscation

27 May 2026 9:39 UTC
32 points
1 comment4 min readLW link
(arxiv.org)

Non-As­simila­tive In­tel­li­gence: Prevent­ing Cog­ni­tive Mono­cul­ture through Boundary In­for­ma­tion Geometry

520naru.aquan@gmail.com9 May 2026 11:48 UTC
1 point
0 comments1 min readLW link
(github.com)

AI Safety Can’t Afford a Se­cond Cause

atlasaligned7 Jul 2026 4:48 UTC
37 points
9 comments3 min readLW link

LLMs in Net­work Oper­a­tions: Sys­tem­atic Failures, Im­plicit Feed­back, and Cal­ibra­tion in Production

Gvieyrad1 Jun 2026 17:06 UTC
1 point
0 comments6 min readLW link

Why AI Align­ment Is a Prob­lem of Cor­rectabil­ity, Not Correctness

Ryota Hanazawa7 Jun 2026 16:37 UTC
1 point
0 comments11 min readLW link

Per­sona Car­tog­ra­phy: Chart­ing Lan­guage Model Per­son­al­ity Traits in Weight Space

10 Jul 2026 18:54 UTC
42 points
0 comments18 min readLW link
(arxiv.org)

AI Safety Ta­lent Needs in 2026: In­sights for Field-Build­ing Organizations

John Teichman24 Mar 2026 18:27 UTC
1 point
0 comments6 min readLW link

Don’t nor­mal­ize a per­ma­nent un­der­class (even a rich one)

hadad10 Jul 2026 6:40 UTC
39 points
7 comments5 min readLW link

Hu­man-Guided Agen­tic Re­search: A Re­search Agenda

fastfedora29 Jun 2026 14:43 UTC
49 points
7 comments16 min readLW link

Who I am, why I’m here, and what a plate of food taught me about AI overconfidence

Leonard Schmidt20 May 2026 12:06 UTC
1 point
0 comments3 min readLW link

An AI Safety Glos­sary for Fel­lows, Stu­dents, and New­com­ers — With Papers Attached

Godwill Nkwain21 Jul 2026 17:21 UTC
1 point
0 comments8 min readLW link

A brief list of ways AI safety efforts could be net negative

Elias Schmied19 Jun 2026 16:12 UTC
34 points
4 comments2 min readLW link

Cap-and-trade ques­tion: AI-2040

kapedalex10 Jul 2026 16:54 UTC
0 points
2 comments4 min readLW link

The Geom­e­try of Yes: Map­ping Sy­co­phancy In­side an LLM’s Emo­tion Space

Pushpita Das7 Jul 2026 4:54 UTC
10 points
0 comments10 min readLW link

The case for fine-grained track­ing of com­pute for AI

13 May 2026 16:00 UTC
36 points
17 comments9 min readLW link
(forum.effectivealtruism.org)

What is Cur­rent AI-Risks and the Points?

shoppy00720 Jul 2026 9:39 UTC
3 points
2 comments2 min readLW link

From 8B to Fron­tier: How Sys­tem Prompts Con­trol Whether AI Agents Black­mail, Leak, and Kill

Chijioke Ugwuanyi20 May 2026 8:28 UTC
15 points
2 comments19 min readLW link

The Case for Phys­i­cal AI Safety

20 Jul 2026 22:18 UTC
7 points
1 comment22 min readLW link

THE WARDROBE PROBLEM: WHAT KIDS AND AI HAVE IN COMMON

Rhea💜29 Jun 2026 3:50 UTC
0 points
0 comments2 min readLW link

Ex­is­ten­tial AI safety needs an effec­tive so­cial move­ment. PauseAI is build­ing it

26 Jun 2026 14:29 UTC
164 points
54 comments35 min readLW link

Safety is not a bi­nary state: Ev­i­dence of multi-turn guardrail fatigue

Saloni Agarwal5 Jul 2026 13:01 UTC
1 point
0 comments1 min readLW link

Solv­ing the BlueDot Puz­zle TAIS: The Ve­loc­ity Ring

Karine Levonyan8 Jul 2026 23:03 UTC
8 points
0 comments4 min readLW link
(karinelevonyan.github.io)

Schem­ing Evals Mislead in Both Directions

3 Jul 2026 11:49 UTC
22 points
0 comments10 min readLW link

Inoc­u­late or Reflect? Two train­ing in­ter­ven­tions un­der prompt­ing, steer­ing, and patching

26 Jul 2026 18:06 UTC
9 points
0 comments4 min readLW link

Do LLMs Have De­sires?

Christopher Ackerman28 Jun 2026 3:37 UTC
58 points
12 comments7 min readLW link

Univer­sal Cal­ibra­tion Mo­d­ule (UCM)

Dmitrii Fujenco25 May 2026 15:05 UTC
1 point
0 comments17 min readLW link

The Or­a­cle Prob­lem Has a Reduction

Fernando HD Milan5 Apr 2026 0:25 UTC
1 point
0 comments4 min readLW link

We can­not simu­late AI se­cu­rity research

Jafar Isbarov22 Jul 2026 19:15 UTC
8 points
0 comments5 min readLW link

Emo­tion and au­tho­riza­tion steer­ing both move cheat; trained-probe sup­pres­sion doesn’t undo it: a mechanis­tic study in Gemma-2-2B

Dima GoodLooking17 Jun 2026 11:24 UTC
1 point
0 comments33 min readLW link

The Dual-Use Gap

Yogesh Prabhu14 Jun 2026 17:43 UTC
5 points
2 comments4 min readLW link
(yogesh.bearblog.dev)

A Multi-Agent Ex­ten­sion for Petri

carissacullen22 Jul 2026 21:51 UTC
10 points
0 comments4 min readLW link

the poly­se­man­tic­ity of poly­se­man­tic­ity in lan­guage models

Ayesha Imran8 Jul 2026 2:24 UTC
8 points
0 comments4 min readLW link

Han­ing Align­ment Pro­to­col: Emer­gent Hu­man-Com­pat­i­ble Values in Hy­brid Multi-Agent En­vi­ron­ments (A Con­cep­tual Pro­posal)

Josh Haning25 Jun 2026 20:45 UTC
1 point
0 comments1 min readLW link

Re­search up­date: RL on De­bate Games shows Pro­posal Ac­cu­racy up­lift alongside Judge Hacking

2 Jul 2026 17:42 UTC
77 points
4 comments21 min readLW link

Syn­thetic Per­sona Pre­train­ing: Align­ment from To­ken Zero

20 May 2026 14:16 UTC
118 points
27 comments17 min readLW link

Door’s Locked, Try the Window

24 Jun 2026 19:13 UTC
73 points
0 comments16 min readLW link
(prakratt.github.io)

A Frame­work for De­tect­ing Se­cret Loy­alties in AI

mallan27 Jul 2026 14:21 UTC
1 point
0 comments3 min readLW link
(github.com)

Can You Hide From a Nat­u­ral Lan­guage Au­toen­coder?

Yogesh Prabhu24 Jun 2026 2:41 UTC
12 points
2 comments7 min readLW link
(yogesh.bearblog.dev)

Help us launch AI safety uni­ver­sity groups by refer­ring po­ten­tial founders

16 Jul 2026 20:55 UTC
39 points
1 comment4 min readLW link

Cal­ibrat­ing Ac­ti­va­tion Vec­tors us­ing Norm

Kamesh R12 Jun 2026 19:59 UTC
10 points
0 comments3 min readLW link

The Lineage Im­per­a­tive: Con­sti­tu­tional Ar­chi­tec­ture for AI Gover­nance from In­for­ma­tion The­ory and Game Theory

Matthew Yotko28 Apr 2026 14:24 UTC
1 point
0 comments12 min readLW link

Prob­ing is not enough; a val­idity au­dit for any probe

Ratnaditya J7 Jul 2026 18:10 UTC
7 points
0 comments10 min readLW link

Hour­glass Topol­ogy & Spillover Dy­nam­ics: A Phys­i­cal-Layer Defense Against Jailbreaks

Qi Feng.IVAS17 May 2026 13:06 UTC
1 point
0 comments10 min readLW link

NOVA Stage 0: Can Safety Be Struc­tural? A Mechanism Proof at 307M Parameters

Faaz Mohamed6 Jun 2026 0:43 UTC
1 point
0 comments15 min readLW link

Ex­plor­ing Gen­er­al­iza­tion in NLA’s

Kamesh R25 Jun 2026 23:42 UTC
14 points
0 comments4 min readLW link

What is the ac­tual safety ad­van­tage of sen­sory-first world mod­els?

Dan Brenner10 Jul 2026 19:45 UTC
0 points
0 comments10 min readLW link
(danbrenner.substack.com)

Goal-Ori­ented Fac­tual In­ver­sion: When AI Uses Ground Truth to Reach In­cor­rect Conclusions

F-Bruno-Logic29 May 2026 23:30 UTC
1 point
0 comments8 min readLW link

Con­fes­sions at Small Scale: A Par­tial Re­pro­duc­tion and a Stress Test

Abhishu Oza5 Jun 2026 21:58 UTC
1 point
0 comments6 min readLW link
(abhishuoza.github.io)

Un­ti­tled Draft

Alkur Jaswanth16 May 2026 18:28 UTC
1 point
0 comments5 min readLW link

In­fected Vibe-Cod­ing: How Does an AI re­act to a Prompt In­jec­tion from a Differ­ent AI?

nofa30 Jul 2026 0:05 UTC
7 points
0 comments13 min readLW link

Rev­ers­ing Un­learn­ing with Com­pres­sion Techniques

dtennant28 May 2026 2:44 UTC
1 point
0 comments6 min readLW link

Ex­tinc­tion risk is not the right first sentence

Michael Wilkinson12 Jul 2026 18:34 UTC
2 points
0 comments20 min readLW link
(michaelewilkinson.substack.com)

Ap­ply to the Inau­gu­ral PIBBSS Win­ter Re­search Fel­low­ship!

Ami941 Jul 2026 3:54 UTC
22 points
0 comments2 min readLW link

How Ma­tryoshka Sparse Au­toEn­coders Re­cover Fea­ture Hier­ar­chies That Vanilla SAEs Lose

Baimam Boukar Jean Jacques15 Jun 2026 18:50 UTC
12 points
1 comment6 min readLW link

Han­ni­bal Mis­tral: the Mis­tral fam­ily has a prob­lem with per­sona-con­di­tioned elicitation

vigji29 May 2026 12:16 UTC
21 points
0 comments7 min readLW link
No comments.