Learn to Develop Your Advantage

ReverendBayes29 Jan 2025 22:06 UTC
16 points
1 comment5 min readLW link

Re­veal­ing al­ign­ment fak­ing with a sin­gle prompt

Florian_Dietz29 Jan 2025 21:01 UTC
9 points
5 comments4 min readLW link

Alle­gory of the Tsunami

Evan Hu29 Jan 2025 19:09 UTC
4 points
1 comment3 min readLW link

My Men­tal Model of AI Op­ti­mist Opinions

tailcalled29 Jan 2025 18:44 UTC
14 points
7 comments1 min readLW link

Plan­ning for Ex­treme AI Risks

joshc29 Jan 2025 18:33 UTC
143 points
5 comments16 min readLW link

Dario Amodei: On Deep­Seek and Ex­port Controls

Zach Stein-Perlman29 Jan 2025 17:15 UTC
53 points
3 comments1 min readLW link
(darioamodei.com)

An­thropic CEO calls for RSI

Andrea_Miotti29 Jan 2025 16:54 UTC
32 points
10 comments1 min readLW link
(darioamodei.com)

Effi­ciency spec­tra and “bucket of cir­cuits” cartoons

Dmitry Vaintrob29 Jan 2025 15:06 UTC
20 points
0 comments7 min readLW link

Deep­Seek: Le­mon, It’s Wednesday

Zvi29 Jan 2025 15:00 UTC
33 points
0 comments33 min readLW link
(thezvi.wordpress.com)

How To Prevent a Dystopia

ank29 Jan 2025 14:16 UTC
−3 points
4 comments1 min readLW link

Whereby: The Zoom al­ter­na­tive you prob­a­bly haven’t heard of

Itay Dreyfus29 Jan 2025 13:01 UTC
4 points
0 comments7 min readLW link
(productidentity.co)

[Question] Whose track record of AI pre­dic­tions would you like to see eval­u­ated?

Jonny Spicer29 Jan 2025 12:05 UTC
2 points
3 comments1 min readLW link

Paper: Open Prob­lems in Mechanis­tic Interpretability

29 Jan 2025 10:25 UTC
71 points
0 comments1 min readLW link
(arxiv.org)

Pos­i­tive jailbreaks in LLMs

dereshev29 Jan 2025 8:41 UTC
6 points
0 comments4 min readLW link

Un­trusted mon­i­tor­ing in­sights from watch­ing ChatGPT play co­or­di­na­tion games

jwfiredragon29 Jan 2025 4:53 UTC
14 points
8 comments9 min readLW link

The Game Board has been Flipped: Now is a good time to re­think what you’re doing

LintzA28 Jan 2025 23:36 UTC
118 points
30 comments13 min readLW link

Recon­cep­tu­al­iz­ing the Noth­ing­ness and Existence

Htarlov28 Jan 2025 20:29 UTC
8 points
1 comment2 min readLW link

Fake think­ing and real thinking

Joe Carlsmith28 Jan 2025 20:05 UTC
120 points
17 comments38 min readLW link

SAE reg­u­lariza­tion pro­duces more in­ter­pretable models

28 Jan 2025 20:02 UTC
21 points
7 comments4 min readLW link

Operator

Zvi28 Jan 2025 20:00 UTC
35 points
1 comment11 min readLW link
(thezvi.wordpress.com)

Deep­Seek Panic at the App Store

Zvi28 Jan 2025 19:30 UTC
51 points
14 comments33 min readLW link
(thezvi.wordpress.com)

“Sharp Left Turn” dis­course: An opinionated review

Steven Byrnes28 Jan 2025 18:47 UTC
228 points
31 comments31 min readLW link

De­tect­ing out of dis­tri­bu­tion text with sur­prisal and entropy

Sandy Fraser28 Jan 2025 18:46 UTC
24 points
4 comments11 min readLW link

Should Art Carry the Weight of Shap­ing our Values?

Krishna Maneesha Dendukuri28 Jan 2025 18:43 UTC
2 points
0 comments3 min readLW link

The mem­o­riza­tion-gen­er­al­iza­tion spec­trum and learn­ing coefficients

Dmitry Vaintrob28 Jan 2025 16:53 UTC
17 points
0 comments10 min readLW link

Ten peo­ple on the inside

Buck28 Jan 2025 16:41 UTC
155 points
28 comments4 min readLW link

Con­sti­tu­tions for ASI?

kanad28 Jan 2025 16:32 UTC
11 points
0 comments1 min readLW link
(forum.effectivealtruism.org)

[Question] Those of you with lots of med­i­ta­tion ex­pe­rience: How did it in­fluence your un­der­stand­ing of philos­o­phy of mind and top­ics such as qualia?

SpectrumDT28 Jan 2025 14:29 UTC
14 points
24 comments1 min readLW link

Will LLMs sup­plant the field of cre­ative writ­ing?

Declan Molony28 Jan 2025 6:42 UTC
8 points
14 comments3 min readLW link

Nvidia doesn’t just sell shovels

winstonBosan28 Jan 2025 4:56 UTC
15 points
4 comments2 min readLW link

Re­in­force­ment Learn­ing by AI Pu­n­ish­ment

Abhishaike Mahajan28 Jan 2025 0:57 UTC
29 points
0 comments8 min readLW link
(www.owlposting.com)

The Good­ness of Morning

YanLyutnev27 Jan 2025 23:25 UTC
−11 points
1 comment3 min readLW link

Jevon’s para­dox and eco­nomic intuitions

Abhimanyu Pallavi Sudhir27 Jan 2025 23:04 UTC
5 points
0 comments1 min readLW link

Defer­ence and De­ci­sion-Making

27 Jan 2025 22:02 UTC
28 points
2 comments7 min readLW link

How differ­ent LLMs an­swered PhilPapers 2020 survey

Satron27 Jan 2025 21:41 UTC
16 points
1 comment14 min readLW link

AI Strat­egy Up­dates that You Should Make

Alice Blair27 Jan 2025 21:10 UTC
20 points
2 comments6 min readLW link

A cri­tique of Soares “4 back­ground claims”

YanLyutnev27 Jan 2025 20:27 UTC
−8 points
0 comments14 min readLW link

Should you go with your best guess?: Against pre­cise Bayesi­anism and re­lated views

Anthony DiGiovanni27 Jan 2025 20:25 UTC
67 points
16 comments22 min readLW link

Is it eth­i­cal to work in AI “con­tent eval­u­a­tion”?

anon_databoy12327 Jan 2025 19:58 UTC
2 points
2 comments1 min readLW link

[Question] Sup­pos­ing that the “Dead In­ter­net The­ory” is true or largely true, how can we act on that in­for­ma­tion?

SpectrumDT27 Jan 2025 16:47 UTC
7 points
5 comments1 min readLW link

Un­der­stand­ing AI World Models w/​ Chris Canal

jacobhaimes27 Jan 2025 16:32 UTC
4 points
0 comments1 min readLW link
(kairos.fm)

The pre­sent perfect tense is ru­in­ing your life

PatrickDFarley27 Jan 2025 16:14 UTC
26 points
14 comments8 min readLW link

The Up­com­ing PEPFAR Cut Will Kill Millions, Many of Them Chil­dren

Bentham's Bulldog27 Jan 2025 16:03 UTC
30 points
2 comments4 min readLW link

[Question] Is the out­put of the soft­max in a sin­gle trans­former at­ten­tion head usu­ally win­ner-takes-all?

Linda Linsefors27 Jan 2025 15:33 UTC
25 points
1 comment1 min readLW link

To know or not to know

arisAlexis27 Jan 2025 13:17 UTC
0 points
3 comments6 min readLW link

My su­pervillain ori­gin story

Dmitry Vaintrob27 Jan 2025 12:20 UTC
123 points
4 comments5 min readLW link

The Clue­less Sniper and the Prin­ci­ple of Indifference

Jim Buhler27 Jan 2025 11:52 UTC
11 points
26 comments2 min readLW link

Scan­less Whole Brain Emulation

Knight Lee27 Jan 2025 10:00 UTC
10 points
5 comments3 min readLW link

Death vs. Suffer­ing: The En­durist-Serenist Divide on Life’s Worst Fate

Alex_Steiner27 Jan 2025 3:59 UTC
4 points
7 comments7 min readLW link

Are we try­ing to figure out if AI is con­scious?

Kristaps Zilgalvis27 Jan 2025 1:05 UTC
21 points
6 comments1 min readLW link
(open.substack.com)