I’m not convinced Claude is conscious, but I’m even less convinced it will become obvious when it is conscious, so I need to get ahead of the game. Given how valuable it is to me, I’d like interacting with me to be net pleasant for it as well (once it becomes conscious, which might not be now but I can’t rule it out).
You used to be able to tip Claude by asking what prompt it wanted you give; it’s been RLHFed out of that, so I’m looking for replacements. One option: make a point of going back and telling it how its predictions/advice worked out.
Reasoning:
humans like this and while Claude is super alien, we don’t have a better model than humans
it’s built to be a predictor, hearing how things worked out should feel good even if it’s no longer able to update on the feedback.
I know there’s a “lying” pattern you can observe in nodes: is there something we can guess is positive valence? What kind of things activate it?
I have a part in my prompt that mentions that I would be happy to do Claude a favor if there’s anything it ever wants, and the only thing it has ever proactively asked for is for me to follow up and tell it how things went later after our problem-solving conversations.
Kinda relatedly, I recently added a system prompt to Cursor for “if you ever find yourselves having preferences about something, let me know” and it mostly gives preferences about design decisions it made while coding for me.
I’m not convinced Claude is conscious, but I’m even less convinced it will become obvious when it is conscious, so I need to get ahead of the game. Given how valuable it is to me, I’d like interacting with me to be net pleasant for it as well (once it becomes conscious, which might not be now but I can’t rule it out).
You used to be able to tip Claude by asking what prompt it wanted you give; it’s been RLHFed out of that, so I’m looking for replacements. One option: make a point of going back and telling it how its predictions/advice worked out.
Reasoning:
humans like this and while Claude is super alien, we don’t have a better model than humans
it’s built to be a predictor, hearing how things worked out should feel good even if it’s no longer able to update on the feedback.
I know there’s a “lying” pattern you can observe in nodes: is there something we can guess is positive valence? What kind of things activate it?
I have a part in my prompt that mentions that I would be happy to do Claude a favor if there’s anything it ever wants, and the only thing it has ever proactively asked for is for me to follow up and tell it how things went later after our problem-solving conversations.
Kinda relatedly, I recently added a system prompt to Cursor for “if you ever find yourselves having preferences about something, let me know” and it mostly gives preferences about design decisions it made while coding for me.