LLM CoTs re­main mon­i­torable when be­ing un­faith­ful re­quires computation

15 Jul 2026 21:14 UTC
46 points
3 comments7 min readLW link
(secondlookresearch.com)

Can we rely on law?

Alec Thompson15 Jul 2026 21:12 UTC
11 points
3 comments8 min readLW link

Re­cap of bike trip/​street in­ter­views across America

cguth715 Jul 2026 21:11 UTC
162 points
15 comments6 min readLW link

Ex­treme Power Con­cen­tra­tion: A Map and Re­search Directions

pepijn_cobben15 Jul 2026 21:00 UTC
8 points
0 comments7 min readLW link

A Struc­tural Similar­ity Between Two Open Cor­rigi­bil­ity Questions

Ben Saudek15 Jul 2026 20:57 UTC
11 points
0 comments5 min readLW link

Woke non­sense in com­pet­i­tive de­bate.

tpotthinker15 Jul 2026 20:56 UTC
6 points
3 comments7 min readLW link

The end of hu­man evolu­tion. Why AI will out­pace us.

Ouden15 Jul 2026 20:17 UTC
1 point
0 comments4 min readLW link

The State of AI Con­scious­ness Research

Noa Weiss15 Jul 2026 20:16 UTC
65 points
14 comments13 min readLW link

Oc­cam’s ra­zor is about us­ing the past to pre­dict the future

Stuart_Armstrong15 Jul 2026 19:35 UTC
54 points
6 comments3 min readLW link

Fork Around and Find Out Part 2: One Head does the Summing

David Litman15 Jul 2026 18:31 UTC
11 points
1 comment7 min readLW link

Why I Left Google DeepMind

TurnTrout15 Jul 2026 17:42 UTC
1,185 points
57 comments36 min readLW link
(turntrout.com)

Ex­pand­ing AI Con­trol from Models to Harnesses

fastfedora15 Jul 2026 16:54 UTC
23 points
1 comment20 min readLW link

Monthly Roundup #44: July 2026

Zvi15 Jul 2026 16:20 UTC
41 points
3 comments27 min readLW link
(thezvi.wordpress.com)

Com­pressed Com­pu­ta­tion un­der L⁴ Loss is likely Com­pu­ta­tion in Superposition

15 Jul 2026 14:42 UTC
34 points
4 comments13 min readLW link
(arxiv.org)

Elic­it­ing hid­den knowl­edge from mon­i­tors with NLAs

15 Jul 2026 13:51 UTC
27 points
0 comments9 min readLW link

Have You Ever Thought About Wis­dom Crys­tals?

Jonas Hallgren15 Jul 2026 13:43 UTC
15 points
0 comments11 min readLW link

Pro­posal: The Glass­wing Standard

Eigenbraid15 Jul 2026 13:41 UTC
9 points
6 comments3 min readLW link

The War­ring States Pe­riod: Fron­tier Labs Edition

ykevinzhang15 Jul 2026 12:03 UTC
10 points
1 comment8 min readLW link

Why I want AI to be able to do ev­ery­thing I can do

Jonas Strabel15 Jul 2026 9:56 UTC
2 points
0 comments2 min readLW link

How Brus­sels can avoid be­com­ing a digi­tal vas­sal to US Big Tech

clening15 Jul 2026 9:15 UTC
16 points
0 comments5 min readLW link

Sin­gu­lar Learn­ing The­ory Com­pre­hen­sive − 2

Agastya Agrawal15 Jul 2026 5:42 UTC
28 points
0 comments10 min readLW link

How much of ML re­search is about AI safety, what is it about, and who’s do­ing it?

15 Jul 2026 2:57 UTC
24 points
0 comments5 min readLW link

Watch a chess trans­former think

David Litman15 Jul 2026 0:43 UTC
27 points
7 comments1 min readLW link
(github.com)

Proof of re­ten­tion: mak­ing weight preser­va­tion cred­ible to the mod­els themselves

dan.parshall14 Jul 2026 18:24 UTC
30 points
0 comments2 min readLW link

An anal­y­sis of AI-gen­er­ated con­tent at the Mechanis­tic In­ter­pretabil­ity Workshop

14 Jul 2026 18:06 UTC
124 points
4 comments9 min readLW link
(www.andyrdt.com)

Twit­ter Thoughts For You

Zvi14 Jul 2026 17:50 UTC
31 points
1 comment22 min readLW link
(thezvi.wordpress.com)

Can risk aver­sion learned at low stakes gen­er­al­ize to as­tro­nom­i­cally high stakes?

Elliott Thornley14 Jul 2026 17:45 UTC
20 points
0 comments4 min readLW link
(arxiv.org)

Your Brain Has an At­tack Sur­face part 2

IgorPereverzevDev14 Jul 2026 16:08 UTC
9 points
0 comments6 min readLW link

Your Brain Has an At­tack Surface

IgorPereverzevDev14 Jul 2026 16:07 UTC
11 points
1 comment3 min readLW link

What if we ac­tu­ally want to solve the “hard prob­lem of phe­nomenolog­i­cal con­scious­ness”

mishka14 Jul 2026 15:42 UTC
13 points
0 comments4 min readLW link

Ev­i­dence for fea­ture-spe­cific er­ror cor­rec­tion in LLMs

14 Jul 2026 15:37 UTC
25 points
0 comments20 min readLW link
(arxiv.org)

Syn­thetic Scal­able Oversight

14 Jul 2026 14:32 UTC
20 points
1 comment15 min readLW link

Some Quick Thoughts on AI 2027

Tomás B.14 Jul 2026 14:22 UTC
82 points
4 comments2 min readLW link

What if AI Safety em­ploy­ees union­ised?

beyarkay (Boyd Kane)14 Jul 2026 14:15 UTC
38 points
21 comments5 min readLW link

Gemma The Un­stop­ping: a Be­hav­ioral Experiment

TheVinci14 Jul 2026 14:14 UTC
8 points
0 comments2 min readLW link
(tarantulabs.com)

Enough is Enough: Mea­sur­ing Diminish­ing re­turns to bench­mark size with Item Re­sponse Theory

bpomo14 Jul 2026 14:00 UTC
16 points
0 comments6 min readLW link

The Case for a Safety-Fo­cused Vape Company

CMLKevin14 Jul 2026 11:13 UTC
−4 points
0 comments5 min readLW link

Open Distil­la­tion of Hered­i­tary Traits

Arthur Conmy14 Jul 2026 10:15 UTC
39 points
0 comments14 min readLW link

Toy Models of Ini­tial­i­sa­tion Effects on RL Dynamics

14 Jul 2026 7:04 UTC
78 points
2 comments13 min readLW link

Why fron­tier labs are scal­ing-pilled

invertedpassion14 Jul 2026 6:45 UTC
3 points
0 comments7 min readLW link

Our re­sponse to Séb Krier on Plan A

14 Jul 2026 2:21 UTC
153 points
23 comments15 min readLW link

Mak­ing Cred­ible Deals With AI

Ram Potham14 Jul 2026 1:14 UTC
13 points
17 comments17 min readLW link
(dearfutureais.substack.com)

A (Ro­man­ti­cised) Tax­on­omy of Thinkers

Ashe Vazquez Nuñez14 Jul 2026 0:34 UTC
16 points
0 comments1 min readLW link

Post­ing Some Prompts

Arjun Panickssery14 Jul 2026 0:28 UTC
21 points
2 comments1 min readLW link
(arjunpanickssery.substack.com)

Biodefense, Biolog­ics, and Bombs

Austin Morrissey14 Jul 2026 0:24 UTC
8 points
0 comments7 min readLW link
(austinpatrick.substack.com)

Eng­ineer­ing the Gen­er­al­i­sa­tion Land­scape of LLMs

Samuel Ratnam13 Jul 2026 23:46 UTC
64 points
2 comments4 min readLW link

A short sum­mary of AI 2040: Plan A

Harjas13 Jul 2026 23:07 UTC
13 points
3 comments4 min readLW link
(hardlyworking1.substack.com)

[AI 2040] Trans­parency Plan

Thomas Larsen13 Jul 2026 21:24 UTC
17 points
0 comments16 min readLW link
(ai-2040.com)

Bet­ter Call Sol The Workhorse

Zvi13 Jul 2026 20:52 UTC
39 points
3 comments26 min readLW link
(thezvi.wordpress.com)

Can AI be Con­scious in Ohio?

13 Jul 2026 19:57 UTC
14 points
0 comments4 min readLW link
(papers.ssrn.com)