wasn’t most of the backchaining driven by a motivation to understand the scorer in-general, not a motivation to understand the scorer with the goal of getting high reward on your current episode
Yeah, it seems plausible that learning about the scorer in general / for the collective was stronger overall in this incident but I’d bet against (also oops in my previous message I should have mentioned learning about the scorer for the collective). (Also nit: I think it’s more accurate to talk about score for the task rather than reward on the episode here.)
The report mentions many instances of agents sacrificing reward on their current episode, in order to help the collective/the board understand the scorer better. That behavior doesn’t make any sense if it was backchaining from its actual score.
Notably, the agents only sacrificed themselves when the trade-off to their own score was small. So I still think that back-chaining from actual score was stronger than the other motivations here. It’s just that when they were sacrificing themselves, not much of their own score was on the line. See e.g. this snippet where it decides not to take the altruistic action:
Yeah, it seems plausible that learning about the scorer in general / for the collective was stronger overall in this incident but I’d bet against (also oops in my previous message I should have mentioned learning about the scorer for the collective). (Also nit: I think it’s more accurate to talk about score for the task rather than reward on the episode here.)
Notably, the agents only sacrificed themselves when the trade-off to their own score was small. So I still think that back-chaining from actual score was stronger than the other motivations here. It’s just that when they were sacrificing themselves, not much of their own score was on the line. See e.g. this snippet where it decides not to take the altruistic action: