We reviewed a demonstration of this specific technique being used to identify a small number of previously known, minor vulnerabilities. These vulnerabilities all appear relatively simple, and we have found that other publicly-available models are able to discover them as well without requiring a bypass.
To me, this seems like potentially pretty misleading communication from Anthropic and somewhat silly.
Amazon claims that they used a technique that successfully bypasses Fable’s safeguards to get an response from Mythos, which IIUC Anthropic does not dispute. To the question about whether the safeguards work, it is probably immaterial whether the response is something you could have gotten from other models. The value of this demonstration is that it shows that the safeguards did not work, not that in this specific case any harm could/was done. Unless Anthropic makes the imho very implausible claim that the safeguards are developed in a way that they detect whether a particular vulnerability is previously known (or minor), this technique is evidence that an attacker totally could have used this for vulnerabilities that other models could not find.
I greatly enjoy feeding ducks in my local park, so I am biased, but to me the reasons for not feeding them have never quite held up. (although of course they should be fed food appropriate for their digestive system). Specifically on the Malthusian trap, it seems that I either don’t feed the wild animal and it dies from starvation, or I feed it, it reproduces and then its offspring dies from starvation, which is neutral in terms of animals dying from starvation at least within that species.