huh, I thought that was a description of a honeypot, not of a mistake.. was the counterpoint only in my mind and not in the OP that this will reinforce detection of “smells like a trap environment” without generalizing to large rollouts because no one will spend thousands of dollars per run for multiday-size problems that are “just honeypots” ⇒ the small obvious stupid hacks will result in a slap on the fingers while large hacks (or anything unnoticed by the grader) will be rewarded?
What I had in mind was integrating honeypot questions into larger evals. If most of the questions can be answered legitimately but a small number of them can’t, then you’re still mostly evaluating real (non-cheating) capabilities.
I’m not sure how large hacks could get rewarded. If the only way to answer a question correctly is to cheat, then anything that gets it right is cheating. Unless you mean hacks large enough to straight-up directly change the evaluation score, in which case we have bigger problems.
do you have an example of a task that is known to be impossible up front to the judge but unknown (provably not in training data and not “short” inference distance to guess with high probability from the prompt alone) to the tested model in a way that the model has to spend a lot of compute to discover the impossibility?
huh, I thought that was a description of a honeypot, not of a mistake.. was the counterpoint only in my mind and not in the OP that this will reinforce detection of “smells like a trap environment” without generalizing to large rollouts because no one will spend thousands of dollars per run for multiday-size problems that are “just honeypots” ⇒ the small obvious stupid hacks will result in a slap on the fingers while large hacks (or anything unnoticed by the grader) will be rewarded?
What I had in mind was integrating honeypot questions into larger evals. If most of the questions can be answered legitimately but a small number of them can’t, then you’re still mostly evaluating real (non-cheating) capabilities.
I’m not sure how large hacks could get rewarded. If the only way to answer a question correctly is to cheat, then anything that gets it right is cheating. Unless you mean hacks large enough to straight-up directly change the evaluation score, in which case we have bigger problems.
do you have an example of a task that is known to be impossible up front to the judge but unknown (provably not in training data and not “short” inference distance to guess with high probability from the prompt alone) to the tested model in a way that the model has to spend a lot of compute to discover the impossibility?