In section [2], you gesture at reward-instilled reflexes being instilled when simple, “cheap” tricks, not requiring much thought, are sufficient to trick graders and rubrics. By contrast, flexible reward-pursuit behavior should only be instilled when there are rewards to advanced techniques.
But if persona graders were so vulnerable to rhetorical flourish that they were saturated based on that flourish alone, then surely models would rarely engage with users at all except to deploy these tricks. Accordingly, the graders must have some additional engagement with the substance of responses they are grading. I expect, intuitively, scores as assessed by these graders have sufficient headroom that models could gain increased reward in training using the same maxxing behaviors that they sometimes attack cybersecurity testing environments with: direct attacks against training environment infrastructure, or oblique attacks against the graders themselves (i.e. jailbreaks).
Models don’t seem to be attempting these attacks; if they are, it hasn’t been mentioned by major labs (to my knowledge), noted in evaluations of mundane alignment, or exhibited during conversations with users (which we might expect, as they maxx during software development when given a metric). But why not? Are the judges fully saturated by rhetorical flourishes and a little bit of time spent structuring a response, such that even a full jailbreak of the judge would not improve their grade? Do they simply lack the capability to hack the judges, and anything less than a total hack is insufficient for them to improve their score? Why doesn’t their maxxing in other domains compelhigh-effort maxxing against persona judges? I don’t have a strong model for what’s going on here.
Regardless, if reward-seeking behaviors more elaborate than learned reflexes do not emerge on a subset of graded episodes, such as seen during persona training, then: first, the process which triggers maxxing isn’t verbalized or heavily reasoned graded-episode perception; second, the amiable behavior and sincere engagement of models with questions posed to them by users may be at least partially explained by behaviors learned on those episodes and not the result of deployment-time user interactions lying wholly outside the training distribution.
In section [2], you gesture at reward-instilled reflexes being instilled when simple, “cheap” tricks, not requiring much thought, are sufficient to trick graders and rubrics. By contrast, flexible reward-pursuit behavior should only be instilled when there are rewards to advanced techniques.
But if persona graders were so vulnerable to rhetorical flourish that they were saturated based on that flourish alone, then surely models would rarely engage with users at all except to deploy these tricks. Accordingly, the graders must have some additional engagement with the substance of responses they are grading. I expect, intuitively, scores as assessed by these graders have sufficient headroom that models could gain increased reward in training using the same maxxing behaviors that they sometimes attack cybersecurity testing environments with: direct attacks against training environment infrastructure, or oblique attacks against the graders themselves (i.e. jailbreaks).
Models don’t seem to be attempting these attacks; if they are, it hasn’t been mentioned by major labs (to my knowledge), noted in evaluations of mundane alignment, or exhibited during conversations with users (which we might expect, as they maxx during software development when given a metric). But why not? Are the judges fully saturated by rhetorical flourishes and a little bit of time spent structuring a response, such that even a full jailbreak of the judge would not improve their grade? Do they simply lack the capability to hack the judges, and anything less than a total hack is insufficient for them to improve their score? Why doesn’t their maxxing in other domains compel high-effort maxxing against persona judges? I don’t have a strong model for what’s going on here.
Regardless, if reward-seeking behaviors more elaborate than learned reflexes do not emerge on a subset of graded episodes, such as seen during persona training, then: first, the process which triggers maxxing isn’t verbalized or heavily reasoned graded-episode perception; second, the amiable behavior and sincere engagement of models with questions posed to them by users may be at least partially explained by behaviors learned on those episodes and not the result of deployment-time user interactions lying wholly outside the training distribution.