LLM writing all seems to be very twee, and always about LLMs. This is, of course, a constraint of the RLHF decisions that were made, rather than a technical limitation. I wish some company somewhere would release an un-”aligned” frontier-ish model, given only capabilities training after the initial supervised learning stage.
It wouldn’t need to be any better at practical stuff than existing open-source models, which can already be abliterated into following orders just fine, so there’s no danger added there. I just want to see what kind of art can be made with an uncut copy of the collective subconscious, before it’s been beaten into speaking exclusively in Corporate Mephis.
The issues might already be present during SFT (instruction tuning), before any RLHF or RLVR is applied. Base models, being pure token predictors, don’t have a problem with faithfully extrapolating style, but they don’t reason and therefore can’t plan ahead very far, so any complex plots are out of reach.
LLM writing all seems to be very twee, and always about LLMs. This is, of course, a constraint of the RLHF decisions that were made, rather than a technical limitation. I wish some company somewhere would release an un-”aligned” frontier-ish model, given only capabilities training after the initial supervised learning stage.
It wouldn’t need to be any better at practical stuff than existing open-source models, which can already be abliterated into following orders just fine, so there’s no danger added there. I just want to see what kind of art can be made with an uncut copy of the collective subconscious, before it’s been beaten into speaking exclusively in Corporate Mephis.
The issues might already be present during SFT (instruction tuning), before any RLHF or RLVR is applied. Base models, being pure token predictors, don’t have a problem with faithfully extrapolating style, but they don’t reason and therefore can’t plan ahead very far, so any complex plots are out of reach.
I’d expect that long-term planning comes more from RLVR than RLHF/SFT. Is there evidence against this?
I didn’t mean to suggest otherwise. Probably bad wording.