sam-alt-ese falcon
Coop Veit
Have you looked at a Cohen’s kappa based system. Treat R1 and R2 as raters scoring rollouts. Weighted kappa lets you assign higher cost to disagreements on rollouts that matter more (high-stakes, high-update, etc.), which addresses your point about some rollouts being more critical for learning. And Fleiss’s kappa generalizes to N reward functions if you want to measure agreement across an ensemble rather than just a pair.
The open question is whether κ(R_true, R_proxy) bounds policy regret in the way that EPIC/STARC metrics do for continuous rewards. My intuition is yes for discrete/categorical reward signals (preference judgments, pass/fail), where the L2-family metrics those papers use aren’t natural anyway. Happy to discuss, working on empirical substrate for this.
“The we-intention Sellars regards as intrinsically valid-”It shallwe be the case that each of us rational beings so acts as to promote our welfare” — embodies a particular conception of what is good —namely, the welfare of rational beings. How, though, to establish the superiority of this account of the good over the rational egoist’s account? Sellars lays out a strategy for doing so but despairs of carrying this strategy through:
To have this intention is to think of oneself as a member of a community consisting of all rational beings. …
144. If … the following two premises were established, this community could be shown to be a reality:
To think of oneself as [a] rational being is (implicitly) to think of oneself as subject to epistemic oughts binding on rational beings generally.
The intersubjective intention to promote epistemic welfare implies the intersubjective intention to promote welfare sans phrase.
These premises would entail that the concept of oneself as a rational being implies the concept of oneself as a member of an ethical community consisting of all rational beings. To be sure, this implication need not be recognized. Indeed, it would take all the dialectical skill of a Socrates, a Hegel or a Peirce to bring it to the surface. Yet if the above premises were true, all rational beings would “implicitly” think of themselves as members of an ethical community consisting of all rational beings. But since a community exists if the relevant individuals think of themselves as its members, the ethical community of rational being would have an “implicit” existence.”
— Jeremy Randel Koons, Ethics of Wilfred Sellars
—
And how you behave determines the size of the world you can play in! Are you part of the vast virtuous mutually normative collective of moral agents that is our society? Or, as often happens with smart self-interested sociopaths, does your game end early in jail or shunned or dead.