If anyone wants to have a voice chat with me about a topic that I’m interested in (see my recent post/comment history to get a sense), please contact me via PM.
My main “claims to fame”:
Created the first general purpose open source cryptography programming library (Crypto++, 1995), motivated by AI risk and what’s now called “defensive acceleration”.
Published one of the first descriptions of a cryptocurrency based on a distributed public ledger (b-money, 1998), predating Bitcoin.
Proposed UDT, combining the ideas of updatelessness, policy selection, and evaluating consequences using logical conditionals.
First to argue for pausing AI development based on the technical difficulty of ensuring AI x-safety (SL4 2004, LW 2011).
Identified current and future philosophical difficulties as core AI x-safety bottlenecks, potentially insurmountable by human researchers, and advocated for research into metaphilosophy and AI philosophical competence as possible solutions.


(Cross-posted from X)
Why didn’t more AI safety people advocate for AI pause/stop earlier? Well, aside from sociological reasons, in order to do that, you had to think that none of the following would work out (be feasible, safe, have low enough safety tax) in time, which takes a degree of skepticism almost no one could muster. (Or object to AGI/ASI on non-consequentialist grounds, but there was apparently a very high correlation between consequentialism and early interest in AI safety.)
Friendly AI (CFAI)
Coherent Extrapolated Volition (CEV)
Metaphilosophical AI
Tool AGI
Value Learning
Oracle AI
Agent Foundations
Corrigible AI
Quantilizers
Approval-Directed Agents
Human Imitation
Cooperative Inverse RL (CIRL)
Iterated Distillation and Amplification (IDA)
Task-Directed AGI
RL from Human Feedback (RLHF)
Impact Regularization
Mechanistic Interpretability
AI Safety via Debate
Recursive Reward Modeling
Comprehensive AI Services (CAIS)
Infra-Bayesianism
Alignment by Default
Natural Abstraction
Eliciting Latent Knowledge (ELK)
Shard Theory
Constitutional AI
AI Control
Weak-to-Strong Generalization
To recenter the sociology, it was much easier to build a career out of being bullish one or more of these approaches, than out of general skepticism. (I’ve been independent, financially and otherwise, throughout my participation in AI safety, which was perhaps not a coincidence from being the only AI pause/stop advocate for a long time.)