In fact, NLAs suggest Claude suspects it’s being tested across many of our evaluations, even when it doesn’t verbalize its suspicions.
This is expected, but still very concerning.
possibly referencing how the context resembles adversarial testing
Perhaps it is time for new scenarios? Or at least, more subtle ones?

I tried fine-tuning an AI on my writing (using Thinking Machines’ Tinker) to sound like me, starting with simple rewrites. Got a strange result.
(Model: Kimi-K2.6, SFT)
Prompt:
Output:
Any ideas why it went crazy? Is this a common thing? Happy to share more info on my methodology.