RSS

Mateusz Bagiński

Karma: 4,311

I endorse and operate by Crocker’s rules.

I have not signed any agreements whose existence I cannot mention.

Ori­ent­ing Towards Over­sight: Which AIs Should Want to Defect?

31 Jul 2026 15:20 UTC
21 points
0 comments13 min readLW link
(limits-of-evaluation.org)

The Halo Defense

Mateusz Bagiński16 Jul 2026 10:53 UTC
67 points
22 comments2 min readLW link

The Hu­man Sub­sti­tu­tion Test as a San­ity Check for AI Evaluations

10 Jul 2026 17:27 UTC
31 points
5 comments8 min readLW link
(limits-of-evaluation.org)

The Cube The­ory of Par­tially Grasped Concepts

Mateusz Bagiński9 Jul 2026 15:40 UTC
24 points
2 comments5 min readLW link

En­tan­gle­ment Between an AI and Its Environment

7 Jul 2026 13:28 UTC
27 points
2 comments11 min readLW link
(limits-of-evaluation.org)

The AFFINE Su­per­in­tel­li­gence Align­ment Sem­i­nar – A Retrospective

2 Jul 2026 11:57 UTC
104 points
1 comment8 min readLW link

De­ploy­ment Aware­ness Mat­ters More Than Eval­u­a­tion Awareness

26 Jun 2026 22:54 UTC
46 points
7 comments7 min readLW link
(limits-of-evaluation.org)

Ap­pli­ca­tions open for the On­line wing of the AFFINE Su­per­in­tel­li­gence Align­ment Seminar

15 Apr 2026 16:10 UTC
25 points
0 comments1 min readLW link

Slack in Cells, Slack in Brains

Mateusz Bagiński31 Mar 2026 0:35 UTC
45 points
3 comments6 min readLW link

Don’t Over­dose Lo­cally Benefi­cial Changes

Mateusz Bagiński28 Mar 2026 18:24 UTC
80 points
12 comments4 min readLW link

Scaf­folded Re­pro­duc­ers, Scaf­folded Agents

Mateusz Bagiński26 Mar 2026 23:47 UTC
37 points
2 comments3 min readLW link

Su­per­in­tel­li­gence Align­ment Sem­i­nar (1 month fo­cused up­skil­ling)

Mateusz Bagiński17 Feb 2026 17:03 UTC
118 points
13 comments3 min readLW link

Rea­sons to sign a state­ment to ban su­per­in­tel­li­gence (+ FAQ for those on the fence)

13 Oct 2025 19:00 UTC
83 points
4 comments13 min readLW link

Safety re­searchers should take a pub­lic stance

19 Sep 2025 18:55 UTC
254 points
65 comments8 min readLW link

Counter-con­sid­er­a­tions on AI arms races

15 May 2025 14:54 UTC
24 points
0 comments18 min readLW link

[Question] Com­pre­hen­sive up-to-date re­sources on the Chi­nese Com­mu­nist Party’s AI strat­egy, etc?

Mateusz Bagiński18 Apr 2025 4:58 UTC
14 points
6 comments1 min readLW link

Good­hart Ty­pol­ogy via Struc­ture, Func­tion, and Ran­dom­ness Distributions

25 Mar 2025 16:01 UTC
35 points
1 comment15 min readLW link

Bounded AI might be viable

6 Mar 2025 12:55 UTC
24 points
4 comments20 min readLW link

Less Anti-Dakka

Mateusz Bagiński31 May 2024 9:07 UTC
83 points
13 comments3 min readLW link

Some Prob­lems with Or­di­nal Op­ti­miza­tion Frame

Mateusz Bagiński6 May 2024 5:28 UTC
9 points
0 comments7 min readLW link