Shakespeare: To P or not to P, that is the question
Vivek Hebbar
Also: is the text in this section meant to express a coherent argument on its own or just be a lossy summary of the more detailed argument in later sections? If just the latter then maybe it’s not worth answering my parent question.
Each player
independently commits to a bargaining strategy that will be played regardless of whether the bargaining game is or .What exactly do we mean by a “bargaining strategy” here? What restrictions does it place on the ways in which the player’s behavior can depend on the game?
Are we assuming that this bargaining strategy behaves identically when we change the disagreement point from
to ? If so, why would that be a valid assumption?
Advice for making robust-to-training model organisms
Why does off-model SFT degrade capabilities?
Incriminating misaligned AI models via distillation
Research Sabotage in ML Codebases
Sleeper Agent Backdoor Results Are Messy
An Empirical Study of Methods for SFTing Opaque Reasoning Models
How do LLMs generalize when we do training that is intuitively compatible with two off-distribution behaviors?
Five approaches to evaluating training-based control measures
Model organisms researchers should check whether high LRs defeat their model organisms
I did say “suppose you are deterministic”. That said, can you spell out how CDT ratifies the optimal policy if randomization is allowed?
It’s impossible for having a proof of A to make it harder to prove B.
Proof(A&B) should not be longer than Proof(A) and Proof(B) put together, since proving A and B separately also proves A&B. (Except maybe you waste a few tokens of syntax to say “therefore A&B” or something.)
Operationalizing FDT
Thanks, good idea!
What does the disagree vote even mean? 😆