I’ve been interested in the potential of zero-knowledge proof type things to verify that computers are not running unfriendly AIs, to get the minimal amount of omniveillance that may be necessary to thread the Scylla and Charybdis of x-risk via extinction and x-risk via totalitarian stagnation. Each computer attesting that it’s not doing <bad things>, with no more than a yes/no. Possible issues: maybe it’s computationally infeasible, hard to operationalize, or can be used to do more intrusive surveillance. I know Drexler was interested in this. IIRC Buterin may have as well.
Has anyone seriously investigated how technically and socially feasible this is? Is anyone (e.g. governance people at MIRI) working fulltime on this?
Can this be done in Chinese?