What you describe (subagents arguing for different plans) seems something between “fairly compatible” to “direct consequence of” predictive processing. Cf https://www.lesswrong.com/posts/3fkBWpE4f9nYbdf7E/multi-agent-predictive-minds-and-ai-alignmentSurprise-enhanced-reward seems interesting, will look it up
What you describe (subagents arguing for different plans) seems something between “fairly compatible” to “direct consequence of” predictive processing. Cf https://www.lesswrong.com/posts/3fkBWpE4f9nYbdf7E/multi-agent-predictive-minds-and-ai-alignment
Surprise-enhanced-reward seems interesting, will look it up