You do need some way to notice it is not drifting off, but then there are obviously Pokemon walkthroughs in the models’ training sets since the Internet is so full of them, so they only need to notice they are not drifting off relative to them.
Also, it seems that scale consistently but moderately-slowly improves models’ “not ignoring instructions when things are complex”, which is fairly important for everything but not obviously sufficient for world-takeover.
It’s not a supply chain attack unless you target the supply chain of existing packages, just package manager abuse. Still pretty unethical, looks like HPIM was pretty aggressive in running up its FelonyBench score.
In theory, they could have used the token stealing vuln to carry out an actual supply chain attack, but “an that AI wants to do something bad and was capable of doing it would have done it” is a tautology.