RSS

Tomáš Gavenčiak

Karma: 544

A researcher in CS theory, AI safety and other stuff.

Ori­ent­ing Towards Over­sight: Which AIs Should Want to Defect?

31 Jul 2026 15:20 UTC
11 points
0 comments13 min readLW link
(limits-of-evaluation.org)

The Hu­man Sub­sti­tu­tion Test as a San­ity Check for AI Evaluations

10 Jul 2026 17:27 UTC
33 points
5 comments8 min readLW link
(limits-of-evaluation.org)

En­tan­gle­ment Between an AI and Its Environment

7 Jul 2026 13:28 UTC
27 points
2 comments11 min readLW link
(limits-of-evaluation.org)

De­ploy­ment Aware­ness Mat­ters More Than Eval­u­a­tion Awareness

26 Jun 2026 22:54 UTC
46 points
7 comments7 min readLW link
(limits-of-evaluation.org)

If This Were a Test, How Much Would It Cost?

16 Jun 2026 22:52 UTC
34 points
9 comments20 min readLW link
(limits-of-evaluation.org)

Ap­ply now to Hu­man-Aligned AI Sum­mer School 2026

20 May 2026 8:44 UTC
17 points
0 comments1 min readLW link
(humanaligned.ai)

Shal­low re­view of tech­ni­cal AI safety, 2025

17 Dec 2025 18:18 UTC
199 points
9 comments47 min readLW link

Sam­ple In­ter­est­ing First

Tomáš Gavenčiak18 Oct 2025 20:09 UTC
8 points
2 comments3 min readLW link

How LLM Beliefs Change Dur­ing Chain-of-Thought Reasoning

16 Jun 2025 16:18 UTC
32 points
3 comments5 min readLW link

Ap­ply now to Hu­man-Aligned AI Sum­mer School 2025

6 Jun 2025 19:31 UTC
28 points
1 comment2 min readLW link
(humanaligned.ai)

Mea­sur­ing Beliefs of Lan­guage Models Dur­ing Chain-of-Thought Reasoning

18 Apr 2025 22:56 UTC
12 points
0 comments13 min readLW link

An­nounc­ing Hu­man-al­igned AI Sum­mer School

22 May 2024 8:55 UTC
51 points
0 comments1 min readLW link
(humanaligned.ai)

In­terLab – a toolkit for ex­per­i­ments with multi-agent interactions

22 Jan 2024 18:23 UTC
69 points
0 comments8 min readLW link
(acsresearch.org)