If models are fairly explicitly chasing reward these days, I wonder if that means that there’s a particular direction (or angle) in activation space which corresponds to anticipated reward on the current task.
If models are fairly explicitly chasing reward these days, I wonder if that means that there’s a particular direction (or angle) in activation space which corresponds to anticipated reward on the current task.