Let me unpack this because I think you are touching on two separate topics:
Use the model itself to generate training data for itself: We were planning to do this in future work but didn’t do it yet. The crucial issue here is “filter that for accuracy somehow”. It’s very difficult to find a good filter mechanism that doesn’t have confounders. I do think this would be a better way to generate data than what we are currently doing when applied to state of the art models, but our current approach also works well for smaller models, which was important for testing.
Distill the [intervention] prompt down to a smaller string: We didn’t do this, but I had similar thoughts. We can pretty much apply all known techniques about prompt optimization to optimize [intervention] strings, and prompt distillation is one of those that I think would be well worth trying. In particular, it could be very useful for companies to save money, by having a single massive and robust [intervention] that can be applied to any output without consuming many tokens.
My first thought on “filter that for accuracy somehow” was to generate, say, 5000 of them, sit down and laboriously read them all, then throw away (or even edit) the obviously wrong ones. Not exactly an easy technique for others to replicate, but often a reasonably good way to get an effective training set: prompt engineer, generate, filter manually, retrain smaller/simpler model. Sometimes you can even rinse and repeat on this approach, if your simpler model is generalizing usefully, or at least take a second look at cases where the trained model disagrees with your manual choice, and see if you were wrong and it’s right — though obviously that gets riskier the more you let earlier versions of the model have input into the training data of later ones.
We spent a lot of time doing exactly that in an automated manner. It’s too time consuming to filter that many samples manually, but we did several iterations of the following loop:
comprehensively define quality criteria
generate a dataset
run a second instance of Claude to rate samples by quality criteria
drop the bad ones
manually read a subset of what’s left and identify remaining failure modes.
update the quality criteria for the next iteration
I completely agree with you — I personally have filtered a thousand samples manually, and it takes a while. Finding a good human-LLM centaur solution is very helpful. Sounds like you know all the tricks at least as well as I do.
Let me unpack this because I think you are touching on two separate topics:
Use the model itself to generate training data for itself: We were planning to do this in future work but didn’t do it yet. The crucial issue here is “filter that for accuracy somehow”. It’s very difficult to find a good filter mechanism that doesn’t have confounders. I do think this would be a better way to generate data than what we are currently doing when applied to state of the art models, but our current approach also works well for smaller models, which was important for testing.
Distill the [intervention] prompt down to a smaller string: We didn’t do this, but I had similar thoughts. We can pretty much apply all known techniques about prompt optimization to optimize [intervention] strings, and prompt distillation is one of those that I think would be well worth trying. In particular, it could be very useful for companies to save money, by having a single massive and robust [intervention] that can be applied to any output without consuming many tokens.
My first thought on “filter that for accuracy somehow” was to generate, say, 5000 of them, sit down and laboriously read them all, then throw away (or even edit) the obviously wrong ones. Not exactly an easy technique for others to replicate, but often a reasonably good way to get an effective training set: prompt engineer, generate, filter manually, retrain smaller/simpler model. Sometimes you can even rinse and repeat on this approach, if your simpler model is generalizing usefully, or at least take a second look at cases where the trained model disagrees with your manual choice, and see if you were wrong and it’s right — though obviously that gets riskier the more you let earlier versions of the model have input into the training data of later ones.
We spent a lot of time doing exactly that in an automated manner. It’s too time consuming to filter that many samples manually, but we did several iterations of the following loop:
comprehensively define quality criteria
generate a dataset
run a second instance of Claude to rate samples by quality criteria
drop the bad ones
manually read a subset of what’s left and identify remaining failure modes.
update the quality criteria for the next iteration
I completely agree with you — I personally have filtered a thousand samples manually, and it takes a while. Finding a good human-LLM centaur solution is very helpful. Sounds like you know all the tricks at least as well as I do.
Thank you!