people can spend a lot of compute offline empirically fitting predictors for the expected output as a function of the model parameters
In fact I implemented a transformer to extract strings memorized by a given ReLU MLP (that couldn’t be found just by looking at the weights) faster than sampling, but for now, slower than GCG, the goal was to see whether its computationally hard is to extract them if they’re naturally learned with SGD, rather than developing a mechanistic method.
So I agree that probably such a competition could be easy to trick this way, unless the problem is estimating the output with a high numerical precision (where transformers would fail), as opposed to “the proportion of rare problematic inputs”. And also solving that competition with a mechanistic estimator implies solving this one too.
In fact I implemented a transformer to extract strings memorized by a given ReLU MLP (that couldn’t be found just by looking at the weights) faster than sampling, but for now, slower than GCG, the goal was to see whether its computationally hard is to extract them if they’re naturally learned with SGD, rather than developing a mechanistic method.
So I agree that probably such a competition could be easy to trick this way, unless the problem is estimating the output with a high numerical precision (where transformers would fail), as opposed to “the proportion of rare problematic inputs”. And also solving that competition with a mechanistic estimator implies solving this one too.