It feels like they are very hard trying to discredit the standard story of alignment. They use vague concepts to then conclude this is evidence for some weird “industrial accidents” story, what is that supposed to mean? This doesn’t sound like scientific inference to me but very much motivated thinking. Reminds me of that “against counting arguments” post where they also try very hard to get some “empirical data” for something that superficially sounds related to make a big conceptual point.
I pretty much strongly agree with this sentiment:
“Our results are evidence that future AI failures may look more like industrial accidents than coherent pursuit of goals that were not trained for. ”
I have agreed for years, so maybe it’s my bias talking. I think control theory based approaches (STAMP, STPA) will be able to mitigate these risks.
It feels like they are very hard trying to discredit the standard story of alignment. They use vague concepts to then conclude this is evidence for some weird “industrial accidents” story, what is that supposed to mean? This doesn’t sound like scientific inference to me but very much motivated thinking. Reminds me of that “against counting arguments” post where they also try very hard to get some “empirical data” for something that superficially sounds related to make a big conceptual point.
But you agree the Anthropic post does not demonstrate, or even really provide meaningful evidence for that, right?