The issues might already be present during SFT (instruction tuning), before any RLHF or RLVR is applied. Base models, being pure token predictors, don’t have a problem with faithfully extrapolating style, but they don’t reason and therefore can’t plan ahead very far, so any complex plots are out of reach.
The issues might already be present during SFT (instruction tuning), before any RLHF or RLVR is applied. Base models, being pure token predictors, don’t have a problem with faithfully extrapolating style, but they don’t reason and therefore can’t plan ahead very far, so any complex plots are out of reach.
I’d expect that long-term planning comes more from RLVR than RLHF/SFT. Is there evidence against this?
I didn’t mean to suggest otherwise. Probably bad wording.