If you try honeypot scenarios where you trick it into thinking it has real power, you’re also training it to detect evals.
The solution there seems to be not training it on the evaluation data, simply using it to become aware of threats ahead of time. Using a validation set is standard practice in ML, after all.
The solution there seems to be not training it on the evaluation data, simply using it to become aware of threats ahead of time. Using a validation set is standard practice in ML, after all.