Wei Dai comments on Wei Dai’s Shortform

Wei Dai 13 Oct 2025 8:18 UTC
4 points
0

Let’s take an area where you have something to say, like philosophy. Would you be willing to outsource that?

Outsourcing philosophy is the main thing I’ve been trying to do, or trying to figure out how to safely do, for decades at this point. I’ve written about it in various places, including this post and my pinned tweet on X. Quoting from the latter:

Among my first reactions upon hearing “artificial superintelligence” were “I can finally get answers to my favorite philosophical problems” followed by “How do I make sure the ASI actually answers them correctly?”

Aside from wanting to outsource philosophy to ASI, I’d also love to have more humans who could answer these questions for me. I think about this a fair bit and wrote some things down but don’t have any magic bullets.

(I currently think the best bet to eventually getting what I want is to encourage an AI pause along with genetic enhancements for human intelligence, have the enhanced humans solve metaphilosophy and other aspects of AI safety, then outsource the rest of philosophy to ASI, or have the enhanced humans decide what to do at that point.)

BTW I thought this would be a good test for how competent current AIs are at understanding someone’s perspective so I asked a bunch of them how Wei Dai would answer your question, and all of them got it wrong on the first try, except Claude Sonnet 4.5 which got it right on the first try but wrong on the second try. It seems like having my public content in their training data isn’t enough, and finding relevant info from the web and understanding nuance are still challenging for them. (GPT-5 essentially said I’d answer no because I wouldn’t trust current AIs enough, which is really missing the point despite having this whole thread as context.)
- cousin_it 21 Oct 2025 15:19 UTC
  2 points
  0
  Parent
  Yeah, I wouldn’t have predicted this response either. Maybe it’s a case of something we talked about long ago—that if a person’s “true values” are partly defined by how the person themselves would choose to extrapolate them, then different people can end up on very diverging trajectories. Like, it seems I’m slightly more attached to some aspects of human experience that you don’t care much about, and that affects the endpoint a lot.