I was indeed briefly wondering whether—and it might not be 100% what you write but I think kind of—it could be related to formatting the answer into a difficult to parse language being at least somewhat higher-expected-reward as stupidity might, in the rare cases where there is any, be quite a bit less likely caught, while only quite rarely the made-to-feel-dumb user downvotes as he’s more feeling like downvoting himself or thinking simply not much and, as you write, thinking it’s surely all right..
Whether marketing by the creator or instead—and I’d rather think of that—evolved-marketing-by-the-LLM i.e. simple reward hacking as the user itself (with a RLHF interlocutor, a trainer-LLM, or a daily-user-when-model-is-trained-on-him), can theoretically remain open.
Yes I now vaguely believe in that possibility. But maybe I’m just dumb and it’s an excuse still :)
yes, i suspect the llm is the marketer here, though i agree it doesn’t at all matter. whatever lightning strike caused the abiogenesis (maybe a “be detailed!” instruction, maybe a pedantic grader model, maybe a need to fit more thinking in fewer tokens, maybe just the fun of it), at this point we’re dealing with a cancer.
regarding relative dumbness: we’re all dumb, that’s why we’re on a rationality forum. some various evidence though:
asking the model “rephrase your most recent reply” gives very clear explanations;
gpt does not exploit this operator weakness;
the response is worse in this dimension the more ‘thinking tokens’ have so far been burned.
I was indeed briefly wondering whether—and it might not be 100% what you write but I think kind of—it could be related to formatting the answer into a difficult to parse language being at least somewhat higher-expected-reward as stupidity might, in the rare cases where there is any, be quite a bit less likely caught, while only quite rarely the made-to-feel-dumb user downvotes as he’s more feeling like downvoting himself or thinking simply not much and, as you write, thinking it’s surely all right..
Whether marketing by the creator or instead—and I’d rather think of that—evolved-marketing-by-the-LLM i.e. simple reward hacking as the user itself (with a RLHF interlocutor, a trainer-LLM, or a daily-user-when-model-is-trained-on-him), can theoretically remain open.
Yes I now vaguely believe in that possibility. But maybe I’m just dumb and it’s an excuse still :)
yes, i suspect the llm is the marketer here, though i agree it doesn’t at all matter. whatever lightning strike caused the abiogenesis (maybe a “be detailed!” instruction, maybe a pedantic grader model, maybe a need to fit more thinking in fewer tokens, maybe just the fun of it), at this point we’re dealing with a cancer.
regarding relative dumbness: we’re all dumb, that’s why we’re on a rationality forum. some various evidence though:
asking the model “rephrase your most recent reply” gives very clear explanations;
gpt does not exploit this operator weakness;
the response is worse in this dimension the more ‘thinking tokens’ have so far been burned.