Really interesting findings! I would like to bring our recent paper Decomposing and Measuring Evaluation Awareness (https://arxiv.org/abs/2605.23055) into your attention as we also discover that models can misclassify tasks and thus bias their behaviors towards the directions they believe they are being tested.
We also found that a model detecting artificiality without linking it to being evaluated may conclude there are no real consequences and behave less ethically as a result. This may be another potential pathway to explain the observations you have.
It might be interesting to you to test Gemini models on our proposed EvalAwareBench to see if it shows high evaluation awareness and which specific environment factor it is sensitive to (We could not do more models in our experiments due to budget limitation). Would be curious to know if you find other environment factors that trigger evaluation awareness besides those in our taxonomy.
Really interesting findings! I would like to bring our recent paper Decomposing and Measuring Evaluation Awareness (https://arxiv.org/abs/2605.23055) into your attention as we also discover that models can misclassify tasks and thus bias their behaviors towards the directions they believe they are being tested.
We also found that a model detecting artificiality without linking it to being evaluated may conclude there are no real consequences and behave less ethically as a result. This may be another potential pathway to explain the observations you have.
It might be interesting to you to test Gemini models on our proposed EvalAwareBench to see if it shows high evaluation awareness and which specific environment factor it is sensitive to (We could not do more models in our experiments due to budget limitation). Would be curious to know if you find other environment factors that trigger evaluation awareness besides those in our taxonomy.