Introspection is an instantiation of ‘Connecting the Dots’.
Connecting the Dots: train a model g on (x, f(x)) pairs; the model g can infer things about f.
Introspection: Train a model g on (x, f(x)) pairs, where x are prompts and f(x) are the model’s responses. Then the model can infer things about f. Note that here we have f = g, which is a special case of the above.
Introspection is an instantiation of ‘Connecting the Dots’.
Connecting the Dots: train a model g on (x, f(x)) pairs; the model g can infer things about f.
Introspection: Train a model g on (x, f(x)) pairs, where x are prompts and f(x) are the model’s responses. Then the model can infer things about f. Note that here we have f = g, which is a special case of the above.