Want to push back a bit on the agent foundations part, as someone who got into it very early and came up with a bunch of stuff (e.g. the Lobian cooperation paper cites me for the main result). I don’t think AF has much connection to the alignment of AIs that are being developed now. I think AF is “only” an extremely fun field of math/philosophy. Whether it deserves money/prestige/etc is a question for someone else. I just love doing it, and have a bit of allergy to overselling.
I don’t think AF has much connection to the alignment of AIs that are being developed now.
What would you say are the bounds of applicability for AF, then?
To my understanding LLM-based AI only becomes problematic in the X-risk sense (rather than mundane-technology aspects which we already have flawed-but-workable social solutions for) iff it’s trained specifically to act as an autonomous RL-like agent rather than a ~desire-lacking ensemble-simulator of ML-models extracted from patterns-of-data-conversion present in normal text, which seems like it would be relevant to AF research, but maybe I’m wrong about the theoretical bounds of what we can model in the field?
AF has long been focused on AIXI-type cases, and later things like embedded agency. It’s not implausible that there are formalizable concepts for agency/optimization in scaffolded LLM agents (like optimization in in-context learning?), but as of right now, AF seems quite far from it. AF’s wikitag should also give a good idea of where the field is at
Insta-upvote. Keep writing :-)
Want to push back a bit on the agent foundations part, as someone who got into it very early and came up with a bunch of stuff (e.g. the Lobian cooperation paper cites me for the main result). I don’t think AF has much connection to the alignment of AIs that are being developed now. I think AF is “only” an extremely fun field of math/philosophy. Whether it deserves money/prestige/etc is a question for someone else. I just love doing it, and have a bit of allergy to overselling.
What would you say are the bounds of applicability for AF, then?
To my understanding LLM-based AI only becomes problematic in the X-risk sense (rather than mundane-technology aspects which we already have flawed-but-workable social solutions for) iff it’s trained specifically to act as an autonomous RL-like agent rather than a ~desire-lacking ensemble-simulator of ML-models extracted from patterns-of-data-conversion present in normal text, which seems like it would be relevant to AF research, but maybe I’m wrong about the theoretical bounds of what we can model in the field?
AF has long been focused on AIXI-type cases, and later things like embedded agency. It’s not implausible that there are formalizable concepts for agency/optimization in scaffolded LLM agents (like optimization in in-context learning?), but as of right now, AF seems quite far from it. AF’s wikitag should also give a good idea of where the field is at