Hmm. Wouldn’t you have to work with its approximations or approximations of its variants, as irl systems have to take finite time to decide on any action? Irl systems such as “real LLM-based systems”.
yes, but we won’t be advocating for RL training of LLMs to make them more like AIXI, and if anything we will be advocating against it.
Hmm. Wouldn’t you have to work with its approximations or approximations of its variants, as irl systems have to take finite time to decide on any action? Irl systems such as “real LLM-based systems”.
yes, but we won’t be advocating for RL training of LLMs to make them more like AIXI, and if anything we will be advocating against it.