I’d be interested to know how “strategic” the implanted behaviour can be. That is, something like “out of the 1% most opportune times to trigger the compromise, how many area actually compromised?”
My naive picture is this:
If the attacker has sufficient access at inference time to monitor the context and implant the backdoor trigger when it is opportune, then it can be quite strategic—but this is potentially quite a high bar
If the attacker can train in the backdoor on train instances where some marker of opportunity is present, then the behaviour can also be quite strategic—but this takes a similarly high degree of monitor-and-control opportunity on the attacker’s part, at train time instead of inference time
What’s the best we can do when the attacker’s train and inference-time compromise opportunities are just a random 0.5% of cases?
Is the best we can do compromising a random 1% of generations (supposing that “success” is monotone in number of compromised generations)? (0.5% “learner takes initiative from training”, 0.5% “attacker implants the backdoor at inference”?)
I think an expected impact metric would be more useful for downstream applications than “how many of the 1% best opportunities can we capture”, I only suggested the latter because it seemed like I can say something about it with fewer assumptions.
Not hundred percent sure I’m parsing you correctly. Given capable AI, the attacker can implant the backdoor trigger during inference whenever possible and rely on the AI to attack select. I’m mostly modeling the attacker as caring about being able to shift the AI propensities to what they desire during inference. I think the attack selection problem looks reasonably similar to that where the AI became misaligned on its own.
I’d like to be able to compute “leverage”—if an attacker has access to random 1% of training episodes and random 5% of inference episodes, how much “total effectiveness” can they recover (where total effectiveness is something like how much of some desired impact they’d achieve with 100% control of both).
The naive computation would say something like: 1% compromise gets you a reliable backdoor, and you then control 5% of inference but you can’t choose it so you get 5% effectiveness. Can you do much better?
I’m not expecting you to have a solid answer to this, but because I’m interested in it I was wondering if you had any thoughts on the question.
I’d be interested to know how “strategic” the implanted behaviour can be. That is, something like “out of the 1% most opportune times to trigger the compromise, how many area actually compromised?”
My naive picture is this:
If the attacker has sufficient access at inference time to monitor the context and implant the backdoor trigger when it is opportune, then it can be quite strategic—but this is potentially quite a high bar
If the attacker can train in the backdoor on train instances where some marker of opportunity is present, then the behaviour can also be quite strategic—but this takes a similarly high degree of monitor-and-control opportunity on the attacker’s part, at train time instead of inference time
What’s the best we can do when the attacker’s train and inference-time compromise opportunities are just a random 0.5% of cases?
Is the best we can do compromising a random 1% of generations (supposing that “success” is monotone in number of compromised generations)? (0.5% “learner takes initiative from training”, 0.5% “attacker implants the backdoor at inference”?)
I think an expected impact metric would be more useful for downstream applications than “how many of the 1% best opportunities can we capture”, I only suggested the latter because it seemed like I can say something about it with fewer assumptions.
Not hundred percent sure I’m parsing you correctly.
Given capable AI, the attacker can implant the backdoor trigger during inference whenever possible and rely on the AI to attack select. I’m mostly modeling the attacker as caring about being able to shift the AI propensities to what they desire during inference. I think the attack selection problem looks reasonably similar to that where the AI became misaligned on its own.
I’d like to be able to compute “leverage”—if an attacker has access to random 1% of training episodes and random 5% of inference episodes, how much “total effectiveness” can they recover (where total effectiveness is something like how much of some desired impact they’d achieve with 100% control of both).
The naive computation would say something like: 1% compromise gets you a reliable backdoor, and you then control 5% of inference but you can’t choose it so you get 5% effectiveness. Can you do much better?
I’m not expecting you to have a solid answer to this, but because I’m interested in it I was wondering if you had any thoughts on the question.