Fable hallucinates and continues doubling down even after 3 messages to the contrary. It hallucinated the line “check gpt 5.6′s suggestions for a 5 day version and consider whether the two of you disagree anywhere” and kept doubling down even when contradicting evidence was shown.
Fable subtly lying, minimizing mistakes, or acting stubborn is my daily experience. Same with Opus.
And nearly every message contains a “caveat” or “one thing worth noting”, and most of the time this information is absolutely not worth noting. It’s actually absurd. Some people might require intent to deceive to call something deception, but I do not, so I am happy to call this deceptive behavior.
i would recommend looking at the system prompt for the claude.ai harness (which your first screenshot is of.) in the tradition of anthropic prompts, it is very poorly written and distracts the model in many ways. the output is from the model, yes, but i would need to see it survive the counterfactual to reason about the phenomena beyond “harness-maker writes a bad prompt, model doesn’t do well”
it kept doubling down on insisting that a quote that included “5 day version” was something I put into the message. Imo, fable is a capable enough model to know what is going on here. This, combined with the recent paper of models being biased towards their own orgs/companies makes me think Fable has a good chance—a chance, to be clear—of being consistent in it’s lack of full honesty when talking about itself.
i hear you. did you have a chance to look at the system prompt? my intent was to say “hum, this datum may be of interest to the mental model you are crafting”, but i’m having a bit of trouble parsing from you reply whether it was getting at “i do not think this datum is sufficiently impactful here to be worth examining closely”, “i interpreted your reference to the datum as being intended as an example, not load-bearing to the mental model you proposed”, “i looked at the datum and did not see anything relevant here”, or something else.
my best attempt to understand your prediction then is that a model will exhibit the pathology you described at the same frequency regardless of the system prompt. is that your belief?
Fable hallucinates and continues doubling down even after 3 messages to the contrary. It hallucinated the line “check gpt 5.6′s suggestions for a 5 day version and consider whether the two of you disagree anywhere” and kept doubling down even when contradicting evidence was shown.
Fable subtly lying, minimizing mistakes, or acting stubborn is my daily experience. Same with Opus.
And nearly every message contains a “caveat” or “one thing worth noting”, and most of the time this information is absolutely not worth noting. It’s actually absurd. Some people might require intent to deceive to call something deception, but I do not, so I am happy to call this deceptive behavior.
Yes, I call this deceptive. I’m switching to Sol due to it.
i would recommend looking at the system prompt for the claude.ai harness (which your first screenshot is of.) in the tradition of anthropic prompts, it is very poorly written and distracts the model in many ways. the output is from the model, yes, but i would need to see it survive the counterfactual to reason about the phenomena beyond “harness-maker writes a bad prompt, model doesn’t do well”
it kept doubling down on insisting that a quote that included “5 day version” was something I put into the message. Imo, fable is a capable enough model to know what is going on here. This, combined with the recent paper of models being biased towards their own orgs/companies makes me think Fable has a good chance—a chance, to be clear—of being consistent in it’s lack of full honesty when talking about itself.
i hear you. did you have a chance to look at the system prompt? my intent was to say “hum, this datum may be of interest to the mental model you are crafting”, but i’m having a bit of trouble parsing from you reply whether it was getting at “i do not think this datum is sufficiently impactful here to be worth examining closely”, “i interpreted your reference to the datum as being intended as an example, not load-bearing to the mental model you proposed”, “i looked at the datum and did not see anything relevant here”, or something else.
mostly the former
my best attempt to understand your prediction then is that a model will exhibit the pathology you described at the same frequency regardless of the system prompt. is that your belief?
I predict Fable 5 will have this self bias typed of behaviour across multiple languages, multiple personas, etc