Heh! I’ve certainly heard about good results using that method, but has there been rigorous evaluation? Is it bullet-proof enough to use with security-critical code? What about in cases where adversaries have crafty ways to bias requirements documents (e.g. they control documentation of libraries being used?), so that they mean the right things to specialists but somehow manage to trigger wonky behavior by opaque AI models?
It is a very flawed method, yeah. If you’re going to do something on the level of a onetime system prompt and forget about it, this is the best trick I know.
Heh! I’ve certainly heard about good results using that method, but has there been rigorous evaluation? Is it bullet-proof enough to use with security-critical code? What about in cases where adversaries have crafty ways to bias requirements documents (e.g. they control documentation of libraries being used?), so that they mean the right things to specialists but somehow manage to trigger wonky behavior by opaque AI models?
It is a very flawed method, yeah. If you’re going to do something on the level of a onetime system prompt and forget about it, this is the best trick I know.