We would like to run causal experiments intervening on chain-of-thought reasoning by inserting reasoning saying that the model is or is not in an alignment honeypot, to confirm this causes a change in behaviour. Unfortunately we were not able to get the relevant large models running locally in time for this post, so we can only present results showing a presence or lack of correlation.
FYI, DeepInfra provides a token-completion APIs for most models, if you want to do this sort of analysis without running a model locally. E.g. see docs for GLM-5 here.
FYI, DeepInfra provides a token-completion APIs for most models, if you want to do this sort of analysis without running a model locally. E.g. see docs for GLM-5 here.