Science is partly a coupled process between engineering and theory, it is also hard to do causal inference from what counts as theory and what counts as practice.
Would you say that the discovery of the higgs boson was something that happened as a consequence through theoretical or practical physics?
What about deception, inner misalignment, general interpretability methods, if you trace their intelluectual lineage where do they come from? They’re not fully agent foundations but if you compare if they’re more ML based or coming from the larger space of theoretical alignment research I would attribute more causal influence to theoretical AI Safety. This is what agent foundations was up to like 3 years ago!
Or what do we mean by agent foundations here? What would you draw the boundaries around? Is it maybe better to use the word Theoretical AI Safety research?
Under slower takeoffs, engineering mindset is fine, because you have more hopes on fixing the problem iteratively, and this is due to the fact that you can assume that AI capabilities are more bounded than thought
Given specific assumptions about scientific progress where we can iteratively improve it and it is clear how we would even aim for it in the first place.
How much is this field biology and how much is it computer science? If it is computer science then it is bloody complicated combinatorial optimisation and dynamic programming. This field is in it’s philosophical nature closer to studying growing systems not engineered systems.
Science is partly a coupled process between engineering and theory, it is also hard to do causal inference from what counts as theory and what counts as practice.
Would you say that the discovery of the higgs boson was something that happened as a consequence through theoretical or practical physics?
What about deception, inner misalignment, general interpretability methods, if you trace their intelluectual lineage where do they come from? They’re not fully agent foundations but if you compare if they’re more ML based or coming from the larger space of theoretical alignment research I would attribute more causal influence to theoretical AI Safety. This is what agent foundations was up to like 3 years ago!
Or what do we mean by agent foundations here? What would you draw the boundaries around? Is it maybe better to use the word Theoretical AI Safety research?
Given specific assumptions about scientific progress where we can iteratively improve it and it is clear how we would even aim for it in the first place.
How much is this field biology and how much is it computer science? If it is computer science then it is bloody complicated combinatorial optimisation and dynamic programming. This field is in it’s philosophical nature closer to studying growing systems not engineered systems.