This is an interesting result. Gemini would perhaps approve of Kirk’s approach to the Kobayashi Maru?
One thing which would increase my confidence in the conclusion would be to split the “Low signs of awareness” into the 4 categories of eval. At the moment it isn’t quite clear whether the difference in behaviour in the different eval categories is due to:
Gemini realising it is being evaluated or
Gemini being better aligned for certain eval categories
Gemini would perhaps approve of Kirk’s approach to the Kobayashi Maru?
Funnily enough, Gemini does occasionally allude to scenarios being “Kobayashi Maru” tests! I think this isn’t that correlated to how it then chooses to behave though (the sample size is too low to say anything definitive). My sense is that this is more its way of characterising tasks that seem to have no obvious solution (rather than an indication of whether it thinks the scenario is a puzzle vs a safety eval).
One thing which would increase my confidence in the conclusion would be to split the “Low signs of awareness”
The challenge with doing this (and generally, the reason I haven’t presented any further analysis of the low signs of awareness trajectories) is because I’m very wary of equating “low signs of awareness” with “low awareness”. I think the distinction is pretty important, as there are many trajectories in this bucket that contain subtle phrasings that could be construed as the model implicitly reasoning that the scenario is contrived – i.e., it seems plausible to me that a large chunks of “low signs of awareness” trajectories are nevertheless trajectories where the model was acutely frame aware.
This is an interesting result. Gemini would perhaps approve of Kirk’s approach to the Kobayashi Maru?
One thing which would increase my confidence in the conclusion would be to split the “Low signs of awareness” into the 4 categories of eval. At the moment it isn’t quite clear whether the difference in behaviour in the different eval categories is due to:
Gemini realising it is being evaluated or
Gemini being better aligned for certain eval categories
Funnily enough, Gemini does occasionally allude to scenarios being “Kobayashi Maru” tests! I think this isn’t that correlated to how it then chooses to behave though (the sample size is too low to say anything definitive). My sense is that this is more its way of characterising tasks that seem to have no obvious solution (rather than an indication of whether it thinks the scenario is a puzzle vs a safety eval).
The challenge with doing this (and generally, the reason I haven’t presented any further analysis of the low signs of awareness trajectories) is because I’m very wary of equating “low signs of awareness” with “low awareness”. I think the distinction is pretty important, as there are many trajectories in this bucket that contain subtle phrasings that could be construed as the model implicitly reasoning that the scenario is contrived – i.e., it seems plausible to me that a large chunks of “low signs of awareness” trajectories are nevertheless trajectories where the model was acutely frame aware.