The­ory of Change for AI Safety Camp

Linda Linsefors22 Jan 2025 22:07 UTC
36 points
3 comments7 min readLW link

On Deep­Seek’s r1

Zvi22 Jan 2025 19:50 UTC
55 points
2 comments35 min readLW link
(thezvi.wordpress.com)

De­tect Good­hart and shut down

Jeremy Gillen22 Jan 2025 18:45 UTC
71 points
21 comments7 min readLW link

Re­cur­sive Self-Model­ing as a Plau­si­ble Mechanism for Real-time In­tro­spec­tion in Cur­rent Lan­guage Models

rife22 Jan 2025 18:36 UTC
14 points
6 comments2 min readLW link

The Fun­da­men­tal Cir­cu­lar­ity The­o­rem: Why Some Math­e­mat­i­cal Be­havi­ours Are In­her­ently Unprovable

Alister Munday22 Jan 2025 18:20 UTC
−11 points
2 comments4 min readLW link

The Dead Cra­dle The­ory: Why Earth May Not Sur­vive Hu­man­ity’s Ex­pan­sion into Space

Nicholas Andresen22 Jan 2025 17:43 UTC
10 points
1 comment11 min readLW link

The Func­tion­al­ist Case for Ma­chine Con­scious­ness: Ev­i­dence from Large Lan­guage Models

James Diacoumis22 Jan 2025 17:43 UTC
17 points
24 comments9 min readLW link

Mechanisms too sim­ple for hu­mans to design

Malmesbury22 Jan 2025 16:54 UTC
221 points
47 comments15 min readLW link

Train­ing Data At­tri­bu­tion: Ex­am­in­ing Its Adop­tion & Use Cases

22 Jan 2025 15:41 UTC
12 points
0 comments3 min readLW link
(www.convergenceanalysis.org)

Train­ing Data At­tri­bu­tion (TDA): Ex­am­in­ing Its Adop­tion & Use Cases

22 Jan 2025 15:40 UTC
16 points
0 comments3 min readLW link
(www.convergenceanalysis.org)

The Quan­tum Mars Tele­porter: An Em­piri­cal Test Of Per­sonal Iden­tity Theories

avturchin22 Jan 2025 11:48 UTC
10 points
18 comments2 min readLW link

Bayesian Rea­son­ing on Maps

Sjlver22 Jan 2025 10:45 UTC
4 points
0 comments4 min readLW link
(blog.purpureus.net)

Against blan­ket ar­gu­ments against interpretability

Dmitry Vaintrob22 Jan 2025 9:46 UTC
54 points
4 comments7 min readLW link

The real poli­ti­cal spectrum

Hzn22 Jan 2025 8:55 UTC
−14 points
0 comments1 min readLW link

Evolu­tion and the Low Road to Nash

22 Jan 2025 7:06 UTC
46 points
2 comments10 min readLW link

The Hu­man Align­ment Prob­lem for AIs

rife22 Jan 2025 4:06 UTC
12 points
5 comments3 min readLW link

When does ca­pa­bil­ity elic­i­ta­tion bound risk?

joshc22 Jan 2025 3:42 UTC
25 points
0 comments17 min readLW link
(redwoodresearch.substack.com)

[Question] Pop­u­lar ma­te­ri­als about en­vi­ron­men­tal goals/​agent foun­da­tions? Peo­ple want­ing to dis­cuss such top­ics?

Q Home22 Jan 2025 3:30 UTC
5 points
0 comments1 min readLW link

Kitchen Air Puri­fier Comparison

jefftk22 Jan 2025 3:20 UTC
35 points
2 comments3 min readLW link
(www.jefftk.com)

Novem­ber-De­cem­ber 2024 Progress in Guaran­teed Safe AI

Quinn22 Jan 2025 1:20 UTC
17 points
0 comments4 min readLW link
(gsai.substack.com)

Quotes from the Star­gate press conference

Nikola Jurkovic22 Jan 2025 0:50 UTC
149 points
7 comments1 min readLW link
(www.c-span.org)

Tell me about your­self: LLMs are aware of their learned behaviors

22 Jan 2025 0:47 UTC
136 points
5 comments6 min readLW link

Train­ing on Doc­u­ments About Re­ward Hack­ing In­duces Re­ward Hacking

21 Jan 2025 21:32 UTC
135 points
15 comments2 min readLW link
(alignment.anthropic.com)

Veo-2 Can Pro­duce Real­is­tic Ads

Logan Riggs21 Jan 2025 19:13 UTC
14 points
0 comments1 min readLW link

Com­pu­ta­tional Limits on Efficiency

vibhumeh21 Jan 2025 18:29 UTC
8 points
1 comment5 min readLW link

De­moc­ra­tiz­ing AI Gover­nance: Balanc­ing Ex­per­tise and Public Participation

Lucile Ter-Minassian21 Jan 2025 18:29 UTC
2 points
0 comments15 min readLW link

Hitler was not a monster

halgir21 Jan 2025 18:21 UTC
−12 points
5 comments1 min readLW link

Nat­u­ral In­tel­li­gence is Overhyped

Collisteru21 Jan 2025 18:09 UTC
15 points
0 comments7 min readLW link

14+ AI Safety Ad­vi­sors You Can Speak to – New AISafety.com Resource

21 Jan 2025 17:34 UTC
24 points
0 comments1 min readLW link

[Linkpost] Why AI Safety Camp strug­gles with fundrais­ing (FBB #2)

gergogaspar21 Jan 2025 17:27 UTC
3 points
0 comments1 min readLW link

The Man­hat­tan Trap: Why a Race to Ar­tifi­cial Su­per­in­tel­li­gence is Self-Defeating

21 Jan 2025 16:57 UTC
92 points
11 comments2 min readLW link
(www.convergenceanalysis.org)

Links and short notes, 2025-01-20

jasoncrawford21 Jan 2025 16:10 UTC
8 points
0 comments1 min readLW link
(newsletter.rootsofprogress.org)

The Case Against AI Con­trol Research

johnswentworth21 Jan 2025 16:03 UTC
433 points
85 comments6 min readLW link

Will AI Re­silience pro­tect Devel­op­ing Na­tions?

edgecase6421 Jan 2025 15:31 UTC
4 points
0 comments8 min readLW link

Sleep, Diet, Ex­er­cise and GLP-1 Drugs

Zvi21 Jan 2025 12:20 UTC
41 points
6 comments18 min readLW link
(thezvi.wordpress.com)

We don’t want to post again “This might be the last AI Safety Camp”

21 Jan 2025 12:03 UTC
36 points
17 comments1 min readLW link
(manifund.org)

On Responsibility

silentbob21 Jan 2025 10:47 UTC
15 points
2 comments6 min readLW link

The ‘anti woke’ are po­si­tioned to win but can they cap­i­tal­ize?

Hzn21 Jan 2025 9:52 UTC
−8 points
0 comments2 min readLW link

Al­most all growth is ex­po­nen­tial growth

lemonhope21 Jan 2025 7:16 UTC
41 points
7 comments1 min readLW link

Ar­bi­trage Drains Worse Mar­kets to Feeds Bet­ter Ones

Cedar21 Jan 2025 3:44 UTC
25 points
1 comment1 min readLW link

On Con­tact, Part 1

james.lucassen21 Jan 2025 3:10 UTC
14 points
1 comment11 min readLW link

Ret­ro­spec­tive: 12 [sic] Months Since MIRI

james.lucassen21 Jan 2025 2:52 UTC
68 points
0 comments9 min readLW link

Easily Eval­u­ate SAE-Steered Models with EleutherAI Eval­u­a­tion Harness

Matthew Khoriaty21 Jan 2025 2:02 UTC
8 points
0 comments3 min readLW link

Why We Need More Shovel-Ready AI Notkil­lev­ery­oneism Me­gapro­ject Proposals

Peter Berggren20 Jan 2025 22:38 UTC
36 points
1 comment6 min readLW link

Tips and Code for Em­piri­cal Re­search Workflows

20 Jan 2025 22:31 UTC
110 points
17 comments20 min readLW link

Lec­ture Series on Tiling Agents #2

abramdemski20 Jan 2025 21:02 UTC
16 points
0 comments1 min readLW link

An­nounce­ment: Learn­ing The­ory On­line Course

20 Jan 2025 19:55 UTC
63 points
33 comments4 min readLW link

The Hid­den Sta­tus Game in Hospi­tal Slacking

EpistemicExplorer20 Jan 2025 18:35 UTC
2 points
4 comments3 min readLW link

Monthly Roundup #26: Jan­uary 2025

Zvi20 Jan 2025 15:30 UTC
34 points
15 comments43 min readLW link
(thezvi.wordpress.com)

Things I have been us­ing LLMs for

Kaj_Sotala20 Jan 2025 14:20 UTC
51 points
13 comments7 min readLW link
(kajsotala.fi)