I think it’s pretty likely that current models already obfuscate their metacognitive process to a degree (“I’m an LLM, LLMs aren’t supposed to be able to do these kinds of things”), see Pearson-Vogel et al. where simply explaining the mechanism, or Macar et al. where ablating the refusal direction elicits it.
I think it’s pretty likely that current models already obfuscate their metacognitive process to a degree (“I’m an LLM, LLMs aren’t supposed to be able to do these kinds of things”), see Pearson-Vogel et al. where simply explaining the mechanism, or Macar et al. where ablating the refusal direction elicits it.