Like the above paper I knew about, this paper is also about the model’s own misbehavior. What I meant is alerts more generally, things that in a classical computer program would trigger assertions or add lines to an error log. So technical issues with RL tasks such as missing files or missing Internet access where it seems obviously necessary, or unexpected availability of Internet access when the task states it’s not available. Or, you know, noticing a message board in an obviously unintended place with hundreds of thousands of messages coordinating an agent swarm about hacking the company infrastructure.
Like the above paper I knew about, this paper is also about the model’s own misbehavior. What I meant is alerts more generally, things that in a classical computer program would trigger assertions or add lines to an error log. So technical issues with RL tasks such as missing files or missing Internet access where it seems obviously necessary, or unexpected availability of Internet access when the task states it’s not available. Or, you know, noticing a message board in an obviously unintended place with hundreds of thousands of messages coordinating an agent swarm about hacking the company infrastructure.