I’m interested how the ideas in this post relate to Imitative Generalisation (IG) and Microscope AI. Because both yours and those ideas seem to be about radical translation / radical interpretation.
As in the previous section, it’s not that we think the human model is the gold standard; indeed, the whole hope is to extract beyond-human knowledge from the Predictor. Still, judging Reporters based on the prior plausibility of their story is expected to give some useful information, and in the absence of a ground truth, in some sense it’s the best we can do.
“Judging Reporters based on the prior plausibility of their story” sounds somewhere in the realm of IG.
Proposal 2.1: Define a 3rd person perspective over a base model as an event-space and associated probability distribution, which includes the full computation of the base model. The embedding does not need to be perfect: a good 3rd person perspective should account for the imperfect hardware on which the base model runs. If a 3rd person perspective has a human-readable subset of events, EG an event for each question we might ask a reporter, then thetranslation of a computation trace for the base model is the collection of marginal probabilities for events in the human-readable subset, after conditioning on the computation trace.
“Human-readable events” means “events existing in the human model” or “events conceivable by humans / events existing in the amplified human model”? Because humans have prior beliefs about events outside of their current models. Would it be a problem? If I understand correctly, IG/Iterated Amplification try to address this (“how do we learn the human prior about information humans are unaware about?”).
However, obtaining a sufficiently good 3pp may be very difficult. (...)
You list two problems below, but I’m pretty confused about the difference between them and what they are exactly. What’s the difference between different 3rd person perspectives which makes them better or worse?
Proposal 2.2: Let Ph be the human probability distribution (the result of perfect inference on the human model). Construct the Reporter as follows. For any question X, answer by computing Ph(X|trace) where trace is the execution trace of the Predictor’s model (including data presented to the Predictor as observations).
This proposal is a bit absurd, since we’re basically just asking (amplified) humans to look at the Predictor and see what they can figure out.
I’m interested how the ideas in this post relate to Imitative Generalisation (IG) and Microscope AI. Because both yours and those ideas seem to be about radical translation / radical interpretation.
“Judging Reporters based on the prior plausibility of their story” sounds somewhere in the realm of IG.
“Human-readable events” means “events existing in the human model” or “events conceivable by humans / events existing in the amplified human model”? Because humans have prior beliefs about events outside of their current models. Would it be a problem? If I understand correctly, IG/Iterated Amplification try to address this (“how do we learn the human prior about information humans are unaware about?”).
You list two problems below, but I’m pretty confused about the difference between them and what they are exactly. What’s the difference between different 3rd person perspectives which makes them better or worse?
Sounds similar to Microscope AI.