Train­ing on Doc­u­ments About Re­ward Hack­ing In­duces Re­ward Hacking

21 Jan 2025 21:32 UTC
135 points
15 comments2 min readLW link
(alignment.anthropic.com)

Veo-2 Can Pro­duce Real­is­tic Ads

Logan Riggs21 Jan 2025 19:13 UTC
14 points
0 comments1 min readLW link

Com­pu­ta­tional Limits on Efficiency

vibhumeh21 Jan 2025 18:29 UTC
8 points
1 comment5 min readLW link

De­moc­ra­tiz­ing AI Gover­nance: Balanc­ing Ex­per­tise and Public Participation

Lucile Ter-Minassian21 Jan 2025 18:29 UTC
2 points
0 comments15 min readLW link

Hitler was not a monster

halgir21 Jan 2025 18:21 UTC
−12 points
5 comments1 min readLW link

Nat­u­ral In­tel­li­gence is Overhyped

Collisteru21 Jan 2025 18:09 UTC
15 points
0 comments7 min readLW link

14+ AI Safety Ad­vi­sors You Can Speak to – New AISafety.com Resource

21 Jan 2025 17:34 UTC
24 points
0 comments1 min readLW link

[Linkpost] Why AI Safety Camp strug­gles with fundrais­ing (FBB #2)

gergogaspar21 Jan 2025 17:27 UTC
3 points
0 comments1 min readLW link

The Man­hat­tan Trap: Why a Race to Ar­tifi­cial Su­per­in­tel­li­gence is Self-Defeating

21 Jan 2025 16:57 UTC
92 points
11 comments2 min readLW link
(www.convergenceanalysis.org)

Links and short notes, 2025-01-20

jasoncrawford21 Jan 2025 16:10 UTC
8 points
0 comments1 min readLW link
(newsletter.rootsofprogress.org)

The Case Against AI Con­trol Research

johnswentworth21 Jan 2025 16:03 UTC
433 points
85 comments6 min readLW link

Will AI Re­silience pro­tect Devel­op­ing Na­tions?

edgecase6421 Jan 2025 15:31 UTC
4 points
0 comments8 min readLW link

Sleep, Diet, Ex­er­cise and GLP-1 Drugs

Zvi21 Jan 2025 12:20 UTC
41 points
6 comments18 min readLW link
(thezvi.wordpress.com)

We don’t want to post again “This might be the last AI Safety Camp”

21 Jan 2025 12:03 UTC
36 points
17 comments1 min readLW link
(manifund.org)

On Responsibility

silentbob21 Jan 2025 10:47 UTC
15 points
2 comments6 min readLW link

The ‘anti woke’ are po­si­tioned to win but can they cap­i­tal­ize?

Hzn21 Jan 2025 9:52 UTC
−8 points
0 comments2 min readLW link

Al­most all growth is ex­po­nen­tial growth

lemonhope21 Jan 2025 7:16 UTC
41 points
7 comments1 min readLW link

Ar­bi­trage Drains Worse Mar­kets to Feeds Bet­ter Ones

Cedar21 Jan 2025 3:44 UTC
25 points
1 comment1 min readLW link

On Con­tact, Part 1

james.lucassen21 Jan 2025 3:10 UTC
14 points
1 comment11 min readLW link

Ret­ro­spec­tive: 12 [sic] Months Since MIRI

james.lucassen21 Jan 2025 2:52 UTC
68 points
0 comments9 min readLW link

Easily Eval­u­ate SAE-Steered Models with EleutherAI Eval­u­a­tion Harness

Matthew Khoriaty21 Jan 2025 2:02 UTC
8 points
0 comments3 min readLW link

Why We Need More Shovel-Ready AI Notkil­lev­ery­oneism Me­gapro­ject Proposals

Peter Berggren20 Jan 2025 22:38 UTC
36 points
1 comment6 min readLW link

Tips and Code for Em­piri­cal Re­search Workflows

20 Jan 2025 22:31 UTC
110 points
17 comments20 min readLW link

Lec­ture Series on Tiling Agents #2

abramdemski20 Jan 2025 21:02 UTC
16 points
0 comments1 min readLW link

An­nounce­ment: Learn­ing The­ory On­line Course

20 Jan 2025 19:55 UTC
63 points
33 comments4 min readLW link

The Hid­den Sta­tus Game in Hospi­tal Slacking

EpistemicExplorer20 Jan 2025 18:35 UTC
2 points
4 comments3 min readLW link

Monthly Roundup #26: Jan­uary 2025

Zvi20 Jan 2025 15:30 UTC
34 points
15 comments43 min readLW link
(thezvi.wordpress.com)

Things I have been us­ing LLMs for

Kaj_Sotala20 Jan 2025 14:20 UTC
51 points
13 comments7 min readLW link
(kajsotala.fi)

[Question] What are the chances that Su­per­hu­man Agents are already be­ing tested on the in­ter­net?

artemium20 Jan 2025 11:09 UTC
3 points
1 comment1 min readLW link

Detroit Lions—over con­fi­dence is over rated?

Hzn20 Jan 2025 10:53 UTC
6 points
0 comments1 min readLW link

Log­its, log-odds, and loss for par­allel circuits

Dmitry Vaintrob20 Jan 2025 9:56 UTC
57 points
4 comments11 min readLW link

Wor­ries about la­tent rea­son­ing in LLMs

Caleb Biddulph20 Jan 2025 9:09 UTC
48 points
11 comments7 min readLW link

SIGMI Cer­tifi­ca­tion Criteria

a littoral wizard20 Jan 2025 2:41 UTC
6 points
0 comments1 min readLW link

AXRP Epi­sode 38.5 - Adrià Gar­riga-Alonso on De­tect­ing AI Scheming

DanielFilan20 Jan 2025 0:40 UTC
9 points
0 comments16 min readLW link

The Mon­ster in Our Heads

testingthewaters19 Jan 2025 23:58 UTC
41 points
4 comments5 min readLW link

AI: How We Got Here—A Neu­ro­science Perspective

Mordechai Rorvig19 Jan 2025 23:51 UTC
5 points
0 comments2 min readLW link
(www.kickstarter.com)

Agent Foun­da­tions 2025 at CMU

19 Jan 2025 23:48 UTC
90 points
10 comments1 min readLW link

Who is mar­ket­ing AI al­ign­ment?

MrThink19 Jan 2025 21:37 UTC
23 points
4 comments1 min readLW link

Some les­sons from the OpenAI-Fron­tierMath debacle

7vik19 Jan 2025 21:09 UTC
71 points
9 comments4 min readLW link

Max­i­mally Eggy Crepes

jefftk19 Jan 2025 20:40 UTC
12 points
0 comments1 min readLW link
(www.jefftk.com)

The sec­ond bit­ter les­son — there’s a fun­da­men­tal prob­lem with al­ign­ing dis­tributed AI

aelwood19 Jan 2025 19:00 UTC
−5 points
0 comments5 min readLW link
(pursuingreality.substack.com)

The Gen­tle Romance

Richard_Ngo19 Jan 2025 18:29 UTC
243 points
46 comments15 min readLW link
(www.asimov.press)

Is the­ory good or bad for AI safety?

Dmitry Vaintrob19 Jan 2025 10:32 UTC
29 points
1 comment5 min readLW link

[Question] What’s the Right Way to think about In­for­ma­tion The­o­retic quan­tities in Neu­ral Net­works?

Dalcy19 Jan 2025 8:04 UTC
45 points
13 comments3 min readLW link

Per Trib­al­is­mum ad Astra

Martin Sustrik19 Jan 2025 6:50 UTC
30 points
5 comments2 min readLW link
(250bpm.substack.com)

Five Re­cent AI Tu­tor­ing Studies

Arjun Panickssery19 Jan 2025 3:53 UTC
94 points
0 comments2 min readLW link
(arjunpanickssery.substack.com)

Does So­ciety need a cul­tural out­let in tur­bu­lent poli­ti­cal times?

Freya Mcneill19 Jan 2025 2:45 UTC
−3 points
0 comments7 min readLW link

On Thiel’s New Amer­i­can Regime

shawkisukkar19 Jan 2025 2:45 UTC
−3 points
0 comments5 min readLW link
(shawkisukkar.substack.com)

be the per­son that makes the meet­ing productive

Oldmanrahul18 Jan 2025 22:32 UTC
9 points
0 comments1 min readLW link

Beards and Masks?

jefftk18 Jan 2025 16:00 UTC
73 points
5 comments4 min readLW link
(www.jefftk.com)