I don’t think AF has much connection to the alignment of AIs that are being developed now.
What would you say are the bounds of applicability for AF, then?
To my understanding LLM-based AI only becomes problematic in the X-risk sense (rather than mundane-technology aspects which we already have flawed-but-workable social solutions for) iff it’s trained specifically to act as an autonomous RL-like agent rather than a ~desire-lacking ensemble-simulator of ML-models extracted from patterns-of-data-conversion present in normal text, which seems like it would be relevant to AF research, but maybe I’m wrong about the theoretical bounds of what we can model in the field?
AF has long been focused on AIXI-type cases, and later things like embedded agency. It’s not implausible that there are formalizable concepts for agency/optimization in scaffolded LLM agents (like optimization in in-context learning?), but as of right now, AF seems quite far from it. AF’s wikitag should also give a good idea of where the field is at
I think a scaffolded LLM agent is substantially an embedded agent, no? To the extent that it interfaces with users and tries to satisfy user requests, it needs to have a self-model in its world model.
What would you say are the bounds of applicability for AF, then?
To my understanding LLM-based AI only becomes problematic in the X-risk sense (rather than mundane-technology aspects which we already have flawed-but-workable social solutions for) iff it’s trained specifically to act as an autonomous RL-like agent rather than a ~desire-lacking ensemble-simulator of ML-models extracted from patterns-of-data-conversion present in normal text, which seems like it would be relevant to AF research, but maybe I’m wrong about the theoretical bounds of what we can model in the field?
AF has long been focused on AIXI-type cases, and later things like embedded agency. It’s not implausible that there are formalizable concepts for agency/optimization in scaffolded LLM agents (like optimization in in-context learning?), but as of right now, AF seems quite far from it. AF’s wikitag should also give a good idea of where the field is at
I think a scaffolded LLM agent is substantially an embedded agent, no? To the extent that it interfaces with users and tries to satisfy user requests, it needs to have a self-model in its world model.