Another hypothesis: initializing from non-reasoning models is more likely to induce EM (e.g. I’d be curious to see results using Qwen-3-235B-A22B-Instruct or gpt-4.1-mini on this env)
yeah I agree there’s not a strong reason to expect this apriori, though it is consistent with a) the UK AISI results and (low EM on gpt-oss) b) unpublished results from me and others (which I’m less confident about but still update somewhat on)
Another hypothesis: initializing from non-reasoning models is more likely to induce EM (e.g. I’d be curious to see results using Qwen-3-235B-A22B-Instruct or gpt-4.1-mini on this env)
Seems possible, although I had the exact opposite starting intuition!
yeah I agree there’s not a strong reason to expect this apriori, though it is consistent with
a) the UK AISI results and (low EM on gpt-oss)
b) unpublished results from me and others (which I’m less confident about but still update somewhat on)