RSS

VojtaKovarik

Karma: 1,059

My original background is in mathematics (analysis, topology, Banach spaces) and game theory (imperfect information games). Nowadays, I do AI alignment research (mostly systemic risks, sometimes pondering about “consequentionalist reasoning”).

Ori­ent­ing Towards Over­sight: Which AIs Should Want to Defect?

31 Jul 2026 15:20 UTC
19 points
0 comments13 min readLW link
(limits-of-evaluation.org)

The Hu­man Sub­sti­tu­tion Test as a San­ity Check for AI Evaluations

10 Jul 2026 17:27 UTC
33 points
5 comments8 min readLW link
(limits-of-evaluation.org)

En­tan­gle­ment Between an AI and Its Environment

7 Jul 2026 13:28 UTC
27 points
2 comments11 min readLW link
(limits-of-evaluation.org)

De­ploy­ment Aware­ness Mat­ters More Than Eval­u­a­tion Awareness

26 Jun 2026 22:54 UTC
46 points
7 comments7 min readLW link
(limits-of-evaluation.org)

If This Were a Test, How Much Would It Cost?

16 Jun 2026 22:52 UTC
34 points
9 comments20 min readLW link
(limits-of-evaluation.org)

Ap­ply now to Hu­man-Aligned AI Sum­mer School 2026

20 May 2026 8:44 UTC
17 points
0 comments1 min readLW link
(humanaligned.ai)

Eval­u­a­tion as a (Co­op­er­a­tion-En­abling?) Tool

VojtaKovarik10 Dec 2025 18:54 UTC
18 points
0 comments28 min readLW link

Ir­re­spon­si­ble Com­pa­nies Can Be Made of Re­spon­si­ble Employees

VojtaKovarik8 Oct 2025 11:47 UTC
80 points
16 comments5 min readLW link

Ap­ply now to Hu­man-Aligned AI Sum­mer School 2025

6 Jun 2025 19:31 UTC
28 points
1 comment2 min readLW link
(humanaligned.ai)

[Question] When is “un­falsifi­able im­plies false” in­cor­rect?

VojtaKovarik15 Jun 2024 0:28 UTC
3 points
11 comments1 min readLW link

[Question] What is the pur­pose and ap­pli­ca­tion of AI De­bate?

VojtaKovarik4 Apr 2024 0:38 UTC
13 points
9 comments1 min readLW link

Ex­tinc­tion Risks from AI: In­visi­ble to Science?

21 Feb 2024 18:07 UTC
24 points
7 comments1 min readLW link
(arxiv.org)

Ex­tinc­tion-level Good­hart’s Law as a Prop­erty of the Environment

21 Feb 2024 17:56 UTC
23 points
0 comments10 min readLW link

Dy­nam­ics Cru­cial to AI Risk Seem to Make for Com­pli­cated Models

21 Feb 2024 17:54 UTC
19 points
0 comments9 min readLW link

Which Model Prop­er­ties are Ne­c­es­sary for Eval­u­at­ing an Ar­gu­ment?

21 Feb 2024 17:52 UTC
18 points
2 comments7 min readLW link

Weak vs Quan­ti­ta­tive Ex­tinc­tion-level Good­hart’s Law

21 Feb 2024 17:38 UTC
27 points
1 comment2 min readLW link

Vo­j­taKo­varik’s Shortform

VojtaKovarik4 Feb 2024 20:57 UTC
5 points
7 comments1 min readLW link

My Align­ment “Plan”: Avoid Strong Op­ti­mi­sa­tion and Align Economy

VojtaKovarik31 Jan 2024 17:03 UTC
24 points
9 comments7 min readLW link

Con­trol vs Selec­tion: Civil­i­sa­tion is best at con­trol, but nav­i­gat­ing AGI re­quires selection

VojtaKovarik30 Jan 2024 19:06 UTC
7 points
1 comment1 min readLW link

AI Aware­ness through In­ter­ac­tion with Blatantly Alien Models

VojtaKovarik28 Jul 2023 8:41 UTC
7 points
5 comments3 min readLW link