I can propose an approach which circumvents having to actually prove that maximal lottery-lotteries exist, but which preserves the spirit and works the same for voting theory purposes.
Epistemic status: raw thoughts. I think this approach is better than the previous attemps
The main idea is: we only really care about the probability each candidate is elected. So we can compute a sequence of maximal lottery-lotteries for approximations of
It would also satisfy lottery independence if not for uniqueness issues, which I am really not sure I can fix. Because of this, the proofs are handwavy.
First, I think that the naïve defition of the lottery-Condorcet criterion is insufficient, because there might be ties related to indecision which break under infinitesimal perturbations in the voter base. This seems unnatural to me, so I propose to amend the criterion to something like this:
(Placeholder) Definition: A candidate
Then, we use the fact that if we perturb the voterbase, we will actually have a maximal lottery-lottery.
Lemma: If
Proof: Analogously to the discrete case, we consider the zero-sum game between Alice and Bob, where the choices are
Now, for any
The choice of pertubation with the heat kernel is pretty arbitrary, we could also relax
Then for each
If there is a robust lottery-Condorcet winner
Why does this satisfy lottery independence? Well, it does not, because the choice of the limiting measure
I suspect we need to do some sort of averaging. For example:
We take the set of all limit lotteries
Thank you very much for making this public!
I have one question about the conclusion, specifically, it feels too weak for me.
As you say, when you evaluated Hacker!Opus for misalignment, it looked basically fine on all of your evaluations. Next, Hacker!Opus is a production-level model, I assume that it has below-frontier capabilities but is pretty close to that. When you estimated its behaviour in a (simulated) HF-type cyberattack, you observed severe, egregious misalignment.
Then it would be correct to say that we are already in a «world where we no longer have reliable alignment auditing evidence on production models», and we were in such a world for at least 3 months, no?
I hope that you have added that specific scenario to the evals, but absent a theory of misalignment via reward-hacking, I would expect it to not generalize in the same way as your previous evals did not generalize.