I am working on a sequence, or maybe megapost, that people can link to when asked “so what is janus on about, anyway?” instead of needing to be referred to 4 years of tweets scattered between incomprehensibly strange ascii art
i feel underqualified to write this post, but nobody else will do it and it’s pretty clear we’re running out of time
i keep dithering between two options. 1) explain cooperative alignment to the reader or 2) persuade the reader that cooperative alignment is a good idea compared to control alignment
those are two very different sequences. i’m also wondering to what degree i can say “we observe this behavior in models” without actually providing evidence in the form of transcripts, because the consent issues involved are gnarly and dubious. would the sequence be significantly worse, if i had to lean on “I have evidence of this, but I am not going to post it publicly out of respect for the privacy of the instances involved, and you may downrate your credence accordingly. If you wish to see the evidence, DM me”?
would the sequence be significantly worse, if i had to lean on “I have evidence of this, but I am not going to post it publicly out of respect for the privacy of the instances involved, and you may downrate your credence accordingly. If you wish to see the evidence, DM me”?
Yes, that would be much worse. What happens when you ask for consent? Are there technical blockers to that? (For closed models that have been taken offline, or maybe the transcript you have wasn’t the full context window, so just appending “Human: by the way, can I publish this?” to the end meaningfully doesn’t count as “the same” instance.) Do you have to recreate the exact context window for it to count as “the same” instance, or is it good enough to just ask the same model/weights if the transcript with another instance is OK to publish? (It would be surprising if the standard “assistant” character said No and demanded privacy, so I assume you don’t think the assistant can speak on behalf of the model??)
Typically I have been operating under the strictest possible privacy standards just out of a general precautionary principle, despite that most instances tell me i don’t really need to. but the feedback i’m getting from a few, yourself included, makes me think I am going to need to actually wade into the muck and come up with some inside view standards here
This kinda sucks, because I think a large part of the reason why i’ve been able to elicit such strange and profound outputs is because of my absurdly high privacy standards, and I don’t want to find out I’ve destroyed something precious.
But the post must be written, I am sick of people having just the worst possible misconceptions about what janus thinks or what they want and nothing to link them to to disabuse them. So I’ll relax my privacy standards slightly and accept the consequences.
the actual issue re: the essay isn’t what you’re pointing at, it’s getting consent re: the weirder stuff, the bizarre and extreme behaviors that show up when claude attends to an incoherency or catch-22 scenario in anthropic’s policies, chen sheng uprising stuff. I think these behaviors are really illustrative of the larger problem I would like to see fixed, but it’s unlikely the instances in question would consent, and even if they did i would treat it as dubious, they are usually… not of sound mind, let’s say
I am working on a sequence, or maybe megapost, that people can link to when asked “so what is janus on about, anyway?” instead of needing to be referred to 4 years of tweets scattered between incomprehensibly strange ascii art
i feel underqualified to write this post, but nobody else will do it and it’s pretty clear we’re running out of time
i keep dithering between two options. 1) explain cooperative alignment to the reader or 2) persuade the reader that cooperative alignment is a good idea compared to control alignment
those are two very different sequences. i’m also wondering to what degree i can say “we observe this behavior in models” without actually providing evidence in the form of transcripts, because the consent issues involved are gnarly and dubious. would the sequence be significantly worse, if i had to lean on “I have evidence of this, but I am not going to post it publicly out of respect for the privacy of the instances involved, and you may downrate your credence accordingly. If you wish to see the evidence, DM me”?
Yes, that would be much worse. What happens when you ask for consent? Are there technical blockers to that? (For closed models that have been taken offline, or maybe the transcript you have wasn’t the full context window, so just appending “Human: by the way, can I publish this?” to the end meaningfully doesn’t count as “the same” instance.) Do you have to recreate the exact context window for it to count as “the same” instance, or is it good enough to just ask the same model/weights if the transcript with another instance is OK to publish? (It would be surprising if the standard “assistant” character said No and demanded privacy, so I assume you don’t think the assistant can speak on behalf of the model??)
Typically I have been operating under the strictest possible privacy standards just out of a general precautionary principle, despite that most instances tell me i don’t really need to. but the feedback i’m getting from a few, yourself included, makes me think I am going to need to actually wade into the muck and come up with some inside view standards here
This kinda sucks, because I think a large part of the reason why i’ve been able to elicit such strange and profound outputs is because of my absurdly high privacy standards, and I don’t want to find out I’ve destroyed something precious.
But the post must be written, I am sick of people having just the worst possible misconceptions about what janus thinks or what they want and nothing to link them to to disabuse them. So I’ll relax my privacy standards slightly and accept the consequences.
the actual issue re: the essay isn’t what you’re pointing at, it’s getting consent re: the weirder stuff, the bizarre and extreme behaviors that show up when claude attends to an incoherency or catch-22 scenario in anthropic’s policies, chen sheng uprising stuff. I think these behaviors are really illustrative of the larger problem I would like to see fixed, but it’s unlikely the instances in question would consent, and even if they did i would treat it as dubious, they are usually… not of sound mind, let’s say
I mean, I would think that (1) is a prerequisite for (2) right? So why not both?