One ridiculously effective method is to tell the AI to mimic the process and checks and style and patterns of $name, to the extent that you can’t tell whether $name wrote it. You have to pick someone with a near perfect security track record. Perhaps the author of this post is a good name.
Does this really work? It doesn’t work for non-code things; ask Sol or Fable to ‘write like gwern’, and assuming they don’t refuse outright because I am a living author, it is certainly not a gwern-like output ‘to the extent you can’t tell whether gwern wrote it’! Or if it does work, what would make code different?
Heh! I’ve certainly heard about good results using that method, but has there been rigorous evaluation? Is it bullet-proof enough to use with security-critical code? What about in cases where adversaries have crafty ways to bias requirements documents (e.g. they control documentation of libraries being used?), so that they mean the right things to specialists but somehow manage to trigger wonky behavior by opaque AI models?
It is a very flawed method, yeah. If you’re going to do something on the level of a onetime system prompt and forget about it, this is the best trick I know.
One ridiculously effective method is to tell the AI to mimic the process and checks and style and patterns of $name, to the extent that you can’t tell whether $name wrote it. You have to pick someone with a near perfect security track record. Perhaps the author of this post is a good name.
Does this really work? It doesn’t work for non-code things; ask Sol or Fable to ‘write like gwern’, and assuming they don’t refuse outright because I am a living author, it is certainly not a gwern-like output ‘to the extent you can’t tell whether gwern wrote it’! Or if it does work, what would make code different?
Heh! I’ve certainly heard about good results using that method, but has there been rigorous evaluation? Is it bullet-proof enough to use with security-critical code? What about in cases where adversaries have crafty ways to bias requirements documents (e.g. they control documentation of libraries being used?), so that they mean the right things to specialists but somehow manage to trigger wonky behavior by opaque AI models?
It is a very flawed method, yeah. If you’re going to do something on the level of a onetime system prompt and forget about it, this is the best trick I know.