Nice post, +1 to needing more work here. Science of model Intentions ties closely with our model incrimination work, where we want to incriminate or exonerate a model’s misbehavior through understanding its motives.
One note about graphing out the model’s motivation structure: this is great (and we have some examples of this in our blog), but also can easily become very messy and draws parallels to early-interp work on circuit tracing, which is why we try to frame this work specifically in the context of incrimination/exoneration so that the investigation into the model’s motivation structure remains somewhat narrow.
Nice post, +1 to needing more work here. Science of model Intentions ties closely with our model incrimination work, where we want to incriminate or exonerate a model’s misbehavior through understanding its motives.
One note about graphing out the model’s motivation structure: this is great (and we have some examples of this in our blog), but also can easily become very messy and draws parallels to early-interp work on circuit tracing, which is why we try to frame this work specifically in the context of incrimination/exoneration so that the investigation into the model’s motivation structure remains somewhat narrow.