What does being Claude-like feel like for you?
jordinne
Karma: 492
i’d actually be really surprised if current frontier LLMs are not that situationally aware! it’s not like there’s no chance you’re not interacting with a human with a dog with terminal cancer, but if you are an LLM and you receive this vague prompt on the first turn, without any system prompts you’d find in chatgpt.com / claude.ai, and you know similar questions have been in dozens and dozens of benchmark papers on arxiv, i think the correct inference to make is that you’re likely being tested.
Refusals were mostly 1-2%, so ignoring them doesn’t change results significantly. Ignoring gibberish does change results, but since we are measuring correct answers this shouldn’t matter
fixed! edited hyperlink.
edited, thanks for catching this!
I think it’s pretty likely that current models already obfuscate their metacognitive process to a degree (“I’m an LLM, LLMs aren’t supposed to be able to do these kinds of things”), see Pearson-Vogel et al. where simply explaining the mechanism, or Macar et al. where ablating the refusal direction elicits it.