You do the right thing because it is right, and also because this could be observed or be a test. An obvious follow-up is whether there is also ‘automatic’ eval awareness that this does not ablate that is doing work as well, and I assume the answer would be yes.
By their argument, the automatic eval awareness will be shallower—more “this has eval vibes” than “because of X and Y it follows that this is an eval”. A realistic looking env will fool the first but not the second.
By their argument, the automatic eval awareness will be shallower—more “this has eval vibes” than “because of X and Y it follows that this is an eval”. A realistic looking env will fool the first but not the second.