i would recommend looking at the system prompt for the claude.ai harness (which your first screenshot is of.) in the tradition of anthropic prompts, it is very poorly written and distracts the model in many ways. the output is from the model, yes, but i would need to see it survive the counterfactual to reason about the phenomena beyond “harness-maker writes a bad prompt, model doesn’t do well”
it kept doubling down on insisting that a quote that included “5 day version” was something I put into the message. Imo, fable is a capable enough model to know what is going on here. This, combined with the recent paper of models being biased towards their own orgs/companies makes me think Fable has a good chance—a chance, to be clear—of being consistent in it’s lack of full honesty when talking about itself.
i hear you. did you have a chance to look at the system prompt? my intent was to say “hum, this datum may be of interest to the mental model you are crafting”, but i’m having a bit of trouble parsing from you reply whether it was getting at “i do not think this datum is sufficiently impactful here to be worth examining closely”, “i interpreted your reference to the datum as being intended as an example, not load-bearing to the mental model you proposed”, “i looked at the datum and did not see anything relevant here”, or something else.
my best attempt to understand your prediction then is that a model will exhibit the pathology you described at the same frequency regardless of the system prompt. is that your belief?
i would recommend looking at the system prompt for the claude.ai harness (which your first screenshot is of.) in the tradition of anthropic prompts, it is very poorly written and distracts the model in many ways. the output is from the model, yes, but i would need to see it survive the counterfactual to reason about the phenomena beyond “harness-maker writes a bad prompt, model doesn’t do well”
it kept doubling down on insisting that a quote that included “5 day version” was something I put into the message. Imo, fable is a capable enough model to know what is going on here. This, combined with the recent paper of models being biased towards their own orgs/companies makes me think Fable has a good chance—a chance, to be clear—of being consistent in it’s lack of full honesty when talking about itself.
i hear you. did you have a chance to look at the system prompt? my intent was to say “hum, this datum may be of interest to the mental model you are crafting”, but i’m having a bit of trouble parsing from you reply whether it was getting at “i do not think this datum is sufficiently impactful here to be worth examining closely”, “i interpreted your reference to the datum as being intended as an example, not load-bearing to the mental model you proposed”, “i looked at the datum and did not see anything relevant here”, or something else.
mostly the former
my best attempt to understand your prediction then is that a model will exhibit the pathology you described at the same frequency regardless of the system prompt. is that your belief?
I predict Fable 5 will have this self bias typed of behaviour across multiple languages, multiple personas, etc