If anyone wants to have a voice chat with me about a topic that I’m interested in (see my recent post/comment history to get a sense), please contact me via PM.
My main “claims to fame”:
Created the first general purpose open source cryptography programming library (Crypto++, 1995), motivated by AI risk and what’s now called “defensive acceleration”.
Published one of the first descriptions of a cryptocurrency based on a distributed public ledger (b-money, 1998), predating Bitcoin.
Proposed UDT, combining the ideas of updatelessness, policy selection, and evaluating consequences using logical conditionals.
First to argue for pausing AI development based on the technical difficulty of ensuring AI x-safety (SL4 2004, LW 2011).
Identified current and future philosophical difficulties as core AI x-safety bottlenecks, potentially insurmountable by human researchers, and advocated for research into metaphilosophy and AI philosophical competence as possible solutions.
Do you have a link/explanation for this? I think this may be fairly cruxy, because I’m guessing your intuitions for truth-seeking disagreeable nerd AGI are substantially based on truth-seeking disagreeable nerd humans, so it matters what those humans’ real motivations are.
One line of thought I have here is, there are lots of things such a human or AGI could disagree or talk about or have an interest in, how does it pick which one? I think for the human it probably comes down to some kind of subconscious status calculation, but in either case, how does the AGI do it if it doesn’t have its own status motivations or other long-term goals?