Homepage: 0ak.hu
oakhu
The certainty of one million is qualitatively different from a 99% chance of getting at least one million with a 1% chance of getting nothing. That 1% of nothing looms large because of what it means in context: you are giving up a sure million for a gamble that could leave you with nothing.
I really liked your Ellsberg section, and it changed my mind somewhat. But here’s a quibble with the Allais section: “certainty” / “sure” / “could” seem to point to a strict guarantee, not just full probability (“almost sure”); and to the extent that “sure” and “99%” feel “qualitatively different”, it feels like “almost sure” lands closer to “99%” than to “sure”.
More explicitly: suppose that I’m going to flip a fair coin countably many times, and read off a real between 0 and 1 from the resulting sequence; I offer you the choice between (i) 1 million for sure, or (ii) 1 million if the number is an irrational below 0.9, and 5 million if the number is an irrational above 0.9 (so that you get nothing if the result is a rational number). It feels like your defense of Allais preferences would also license picking (i) over (ii): if you can talk yourself out of picking (i) here – “I mean, it could happen in principle, but it’s extremely unlikely” – then it feels like you should be able to talk yourself out of the Allais preferences. But this seems much worse, since you’re giving up the 10% chance of 5 million over the possibility of a probability zero event. So your defense of Allais preferences feels like it overgenerates.Further, we could run a version of Allais where your first option is just 1 million with probability 1, but not for sure (e.g., you get the million only if my coin produces an irrational sequence). But again, it seems somewhat weird to give up your Allais preferences once I make this modification. So your defense of Allais preferences feels like it undergenerates, too.
(To the extent that I’m sympathetic to Allais preferences, I’m tempted to go along with these weird-seeming conclusions; but I’m not very sympathetic, so maybe I’m not the best judge.)
Here’s a version of the first figure for the 39 that Qwen gets correct among the easiest 123; notably now 10⁄39 reconstructions are better than the rock, but the other 75% still aren’t.
Even from the original chart, though, I think “the problems fly over Qwen’s head” isn’t my default hypothesis; since yellow < orange < red for Qwen in that chart (and this one), it’s clearly getting some signal from the differences between the problems (unlike the second figure, where we vary the key number in the problem statement, where Qwen gets ~no signal from this because yellow = orange = red approximately).
For your second point: if you actually compute optimal [rock I] here, we can tell that it’s badly overfit: working with the 123 easiest, if you take [rock I] which is optimal for the other 122 and apply that to your problem of interest, doing this for every problem gives you a mean cosine FVU of 2.05 (i.e., way worse than not doing this). But leaving one out (indeed, leaving half out) makes little difference when instead we just take corrected_i to be unit(unit(recon_i) – mean(unit(recons)) + mean(unit(originals))), i.e., we just correct for the mean bias:
That’s some improvement! In particular, more than 75% of Qwen’s activations now do better than baseline. (Using the fully-optimal one instead just moves yellow and red each down relative to this, and in particular every yellow does better than baseline, although the best yellows don’t move down much; but again, the fully-optimal one is badly overfit.)
So, we knew that Qwen’s NLA was able to distinguish between different problems (even though it ~ignores the numbers), but overall has worse reconstruction error than an optimal rock. But we’ve learned, from a version of the second experiment you suggest, that a large component of its bad reconstruction error can be explained by a single bias direction for these problems.