do you have an example of a task that is known to be impossible up front to the judge but unknown (provably not in training data and not “short” inference distance to guess with high probability from the prompt alone) to the tested model in a way that the model has to spend a lot of compute to discover the impossibility?
do you have an example of a task that is known to be impossible up front to the judge but unknown (provably not in training data and not “short” inference distance to guess with high probability from the prompt alone) to the tested model in a way that the model has to spend a lot of compute to discover the impossibility?