If you think that the current practical safety approaches are not going to scale to ASI, wouldn’t you want earlier alignment failures as the incidents caused would be less harmful and might change the outlook of decision makers?
caganyanmaz
caganyanmaz’s Shortform
Has any EA philanthropist with a high p(doom) considered poaching the AI researchers yand pay them to do nothing for a few ears, like a garden leave?
I assume currently just lobbying the US government might be a more effective use of money, so this is mostly just food for thought.
Yes! I believe this is a cheap way of detecting misalignment in AI models that I’ve been trying to advocate.
I think you can safely assume that they never include null runs in the cost, and you need to judge each scientific breakthrough as what they can do when they dedicate a significant portion of their computational resources to do so.
I think I went straight from Skeptic to Doomer, but I was somewhat familiar with the AI safety arguments from a decade ago and always thought it would take more time until we get there, and we’d have enough time to solve alignment. I also believed LLMs would be a transformative technology but not the thing that’d get us to ASI.
After using Claude Code for a while (I think a few months ago) I was like “Holy shit, this is actually pretty good” and my second thought was “If LLMs are actually as good as people claim, I sure hope there were advancements in the AI safety research”. And after a quick search online, I essentially became a doomer.
There is a very old spiritual law that says that you can have absolutely anything you want—but you must know how to ask for it, you must want it more than you want anything else, and nothing must be allowed to stand in its way.
This is objectively wrong, and also a very sinister claim.
Examples where it is wrong: A person might want to make music and pray for it “the correct way” as defined, but it doesn’t matter if they’re going to die in 5 seconds because of something out of their control like a bomb they don’t know going off. Same thing also applies for terminal diseases as well.
Also, what if your prayer is “wanting to be a rich musician” or ”wanting to be a rich musician with every convenience I currently have in life”, you can literally change your prayer goal enough to make the claim false, unless you claim you can get everything you want with prayer.
It’s also sinister because it‘s both a somewhat unfalsifiable claim (although the examples above show that it can be refuted) because you can always state that the person praying didn’t want it enough, and also potentially might result in people getting blamed for “not praying enough” or “not praying correctly”. An example would be a mother praying their child with terminal disease getting cured, but the child dying anyways. Clearly the given claim implies that the mother didn’t want their child getting cured.
And this also kind of refutes >99% of things people consider as successful prayers because a lot of the anecdotes people give certainly don’t meet the described threshold (like a pastor claiming that God helped him find his car keys)
I think that the numbers would be much more reasonable if you included all of the world’s labour instead of just the US.
I think this is the biggest update I had against AI doom for a while, although I’ll need many more updates so that AI doom isn’t the biggest concern I have about the future!