I want to see work on verification of claims that AI labs make about their models. We put far too much trust into their research that is uncheckable by the public. Can existing cryptography help? ZK proofs of an LLM’s internals?
Social media was our first real test of the off-switch: a harmful optimizer, fully visible, incapable of fighting back. But we still haven’t pressed it
During the OpenAI HuggingFace incident, agents sacrificed themselves to achieve a goal—the selfish gene?
I thought so too, but @Caleb Biddulph kindly corrected my misconception, see here: https://www.lesswrong.com/posts/GpeMNbmNcGH4b4X7m/ikaxas-shortform-feed?commentId=TuXBeWp73bjuqBQL3
I want to see work on verification of claims that AI labs make about their models. We put far too much trust into their research that is uncheckable by the public. Can existing cryptography help? ZK proofs of an LLM’s internals?
Social media was our first real test of the off-switch: a harmful optimizer, fully visible, incapable of fighting back. But we still haven’t pressed it
Maybe The Real Superintelligent AI Is Extremely Smart Computers