It seems to me like you could make the classic MC-AIXI paper but with LLMs by having an LLM generate your action space and using LLM written reward programs to ontologize reward over the computable environment. The resulting system wouldn’t be superintelligent, and this should be sufficient to study whatever safety properties the team is interested in.
It seems to me like you could make the classic MC-AIXI paper but with LLMs by having an LLM generate your action space and using LLM written reward programs to ontologize reward over the computable environment. The resulting system wouldn’t be superintelligent, and this should be sufficient to study whatever safety properties the team is interested in.
Yeah, we call (basically) this idea MC-AIXI-LLM.