What I’d really like to see (maybe you already have this data) is what each model is using as evidence to support its score. Or if not, run a version where you have each model output specific quotes of the section of the report that it found most critical to base the score on. It would be interesting to see if each model is actually pulling different evidence sources, or merely interpreting the same set of sources differently.
What I’d really like to see (maybe you already have this data) is what each model is using as evidence to support its score. Or if not, run a version where you have each model output specific quotes of the section of the report that it found most critical to base the score on. It would be interesting to see if each model is actually pulling different evidence sources, or merely interpreting the same set of sources differently.