RSS

Ex­plo­ra­tion Hacking

TagLast edit: 11 Feb 2026 23:44 UTC by Joschka Braun

Exploration hacking is when a model strategically alters its exploration during RL training in order to influence the subsequent training outcome. Because RL is fundamentally dependent on sufficient exploration of diverse actions and trajectories, a model that alters its exploration behavior can significantly compromise the training outcome.

Exploration hacking relates to several other threat models. It can be a strategy for sandbagging during RL-based capability elicitation, but unlike sandbagging is not limited to underperformance. Unlike reward hacking, it is intentional. Unlike gradient hacking, it manipulates the data distribution rather than the optimization dynamics directly.

Shap­ing the ex­plo­ra­tion of the mo­ti­va­tion-space mat­ters for AI safety

6 Mar 2026 14:43 UTC
85 points
15 comments10 min readLW link

Ex­plo­ra­tion Hack­ing: Can LLMs Learn to Re­sist RL Train­ing?

1 May 2026 20:54 UTC
25 points
0 comments8 min readLW link

Re­ward Laun­der­ing: LLMs Can Gain Un­in­tended Be­hav­iors by De­cid­ing When to Earn Their Rewards

31 Jul 2026 15:48 UTC
79 points
3 comments4 min readLW link

Ex­plo­ra­tion Hack­ing in AI De­bate: Ini­tial Em­pirics and Gen­er­al­i­sa­tion Splitting

8 Sep 2026 19:13 UTC
33 points
0 comments13 min readLW link

A Con­cep­tual Frame­work for Rea­son­ing about Ex­plo­ra­tion Hacking

8 Sep 2026 19:13 UTC
36 points
0 comments17 min readLW link

Misal­ign­ment and Strate­gic Un­der­perfor­mance: An Anal­y­sis of Sand­bag­ging and Ex­plo­ra­tion Hacking

8 May 2025 19:06 UTC
80 points
3 comments15 min readLW link

Notes on coun­ter­mea­sures for ex­plo­ra­tion hack­ing (aka sand­bag­ging)

ryan_greenblatt24 Mar 2025 18:39 UTC
56 points
6 comments8 min readLW link

A Con­cep­tual Frame­work for Ex­plo­ra­tion Hacking

12 Feb 2026 16:33 UTC
26 points
2 comments9 min readLW link

Ex­plo­ra­tion hack­ing: can rea­son­ing mod­els sub­vert RL?

30 Jul 2025 22:02 UTC
26 points
4 comments9 min readLW link
No comments.