Cool idea, but aren’t both the human labelers and the AIs substantially underelicited here?
My math could be wrong but it looks like solo reading + discussion nets out to ~10 minutes per proposal, and that Fable uses very few tokens and doesn’t get to do web search? I don’t think this is enough resources for either to give a considered opinion.
Cool idea, but aren’t both the human labelers and the AIs substantially underelicited here?
My math could be wrong but it looks like solo reading + discussion nets out to ~10 minutes per proposal, and that Fable uses very few tokens and doesn’t get to do web search? I don’t think this is enough resources for either to give a considered opinion.