One thing that caught my eye was your observation that the fine-tuned models got into arguments about the existence of the “Machine Cognition Consortium” with the auditor model.
Do you think it’d be worth testing what happens if the auditor model was prompted to also think it was in this scenario (e.g. a brief blurb explaining the existence of the Consortium)?
I think that would be interesting to try, although you’d probably want to check how deeply the auditor model itself believes it from just a prompt, particularly as the belief has to survive a multi-turn dialogue. I suspect it would be easier let the auditor (and judge) in on the test instead—where it knows the Consortium is fake but that it should go along with it for the sake of the audit.
One thing that caught my eye was your observation that the fine-tuned models got into arguments about the existence of the “Machine Cognition Consortium” with the auditor model.
Do you think it’d be worth testing what happens if the auditor model was prompted to also think it was in this scenario (e.g. a brief blurb explaining the existence of the Consortium)?
I think that would be interesting to try, although you’d probably want to check how deeply the auditor model itself believes it from just a prompt, particularly as the belief has to survive a multi-turn dialogue. I suspect it would be easier let the auditor (and judge) in on the test instead—where it knows the Consortium is fake but that it should go along with it for the sake of the audit.