“does this really matter that much, outside of major company NDAs or state-secrets?”
Keeping secrets could be important in social interactions. Personal stories tend to be what allows people to connect and sometimes people would prefer those stories to not become common knowledge for a number of reasons (e.g. not be judged). People are comfortable sharing more when they know there will be no downstream consequences. I think keeping things confidential can enhance social interactions. I for one enjoy when I’m being trusted with a secret.
Below are just some thoughts I may turn into an article at some point.
If not explicitly negotiated, I’ve noticed different people tend to have different thresholds for what they treat as secret. There is also an attitude of outing a secret to show how well-informed you are (I’ve been guilty of this in my teens). Besides that, some contexts require a lot of skill that not everybody has. Suppose a question such as “have you cheated?” is being asked around the table and suppose your answer is “no” and a friend has previously confided that their answer is “yes”. In this case, the Glomar response may be a better answer for you than “no”.
When sharing your own personal secret, you may inevitably reveal information about other people. There is some balance to be struck—revealing things concerning you vs preserving the privacy of others involved (e.g. talking about one’s sex life reveals things about their partners’). It is generally useful to understand what the expectations of the other people involved would be in that context. My best rule of thumb so far is to not say anything that could be used against the other people involved, but this does not cover all cases.
anecdote: Me, A and B were in a hot seat session (together with other people). Later, A would talk with B and mention stuff B said in the hot seat session in front of outsiders. I explained to A that things shared during hot seat stay private by default. In retrospect, it was a mistake that nobody had mentioned the privacy expectation during the hot seat session.
The reward is a value we use in our RL algorithm to calculate the gradient update. In other words, Reward is not the optimization target. The model doesn’t naturally prefer gradient updates in directions where reward is high. You would need to use some meta-reward to train it to do that and this seems to move reward hacking one level above.
Either that or I’m misunderstanding something.