Regards safe outcomes for superintelligence, your parenthetical remark is the one I believe most important. Far above any prosaic or theoretical safety work, our priority should be regulation preventing the development and release of superintelligence, at least until we have strong guarantees on its safety.
I don’t really disagree with any of the other points in your comment. Without a regulatory framework, it seems very likely that prosaic safety techniques will only contribute to bad outcomes. So it makes sense to me if one wants to focus on agent foundations and similar theoretic work. My post is not intended as a critique of agent foundations per se!
However, I do believe that one must be clearsighted on the risks of theoretic work, particularly when built upon abstractions. My critique is that agent foundations sometimes fails to make its assumptions explicit and works backwards from abstractions, effectively building a castle in the sky. A more robust approach would be to make these assumptions very explicit, ideally linking a theory to a set of axioms, so that we can better assess the defensibility of a theory. Some branches of continental philosophy are very bad at this (e.g. Lacan), starting from “metaphor” rather than an axiom, which is why I draw the parallel.
I will note that prosaic safety work could be relevant under a strong regulatory framework. For example, suppose we established an international treaty to freeze AI development at ChatGPT 5.5 Pro / Mythos. The treaty states that we can only advance to higher capabilities/intelligence when we are “sure” that the next model is aligned. With huge amounts of resources dedicated to verifying the next model if safe, it seems feasible to me that prosaic approaches could play a large role in building safe AI under such a regime.
Now, setting up sufficiently strong regulation is of course very hard, and one might critique that “proving” that the next generation of a model is aligned is akin to solving alignment itself! But I suspect that guaranteeing a single model is aligned is much easier than solving alignment for all possible models.
I would still guess it is better not to do prosaic safety work until a global regulatory framework exists, since it accelerates AI progress and thus reduces opportunities to implement said regulation. But there are enough counterarguments that I would be careful moralizing over it (not suggesting anyone in the comments is doing so!).
>> “Now, setting up sufficiently strong regulation is of course very hard, and one might critique that “proving” that the next generation of a model is aligned is akin to solving alignment itself! But I suspect that guaranteeing a single model is aligned is much easier than solving alignment for all possible models.”
Do you know how to prove ” alignment of a model”?
This is currently almost as hard as solving alignment for all possible models.
In fact, strictly speaking it is impossible—the model may have simply encrypted its thoughts and there is no method known to man that can break oneway functions.
Now maybe we make some assumption that we are in a good world where the AI didn’t cryptographically hide its thoughts. It doesn’t really help. The fundamental problem is that there is no behaviourial tests that can definitely exclude a sharp left turn/ treacherous turn/ deception/ name-du-jour. So there is no ′ proving’, just hope & cope and we’re back to prosaic alignment & evals again.
We are much more in agreement than I expected! Thank you for clarifying. I agree with every point you made in this comment. (Excepting the comparison to continental philosophy, which is new to me and not something I think I can evaluate one way or the other.)
Regards safe outcomes for superintelligence, your parenthetical remark is the one I believe most important. Far above any prosaic or theoretical safety work, our priority should be regulation preventing the development and release of superintelligence, at least until we have strong guarantees on its safety.
I don’t really disagree with any of the other points in your comment. Without a regulatory framework, it seems very likely that prosaic safety techniques will only contribute to bad outcomes. So it makes sense to me if one wants to focus on agent foundations and similar theoretic work. My post is not intended as a critique of agent foundations per se!
However, I do believe that one must be clearsighted on the risks of theoretic work, particularly when built upon abstractions. My critique is that agent foundations sometimes fails to make its assumptions explicit and works backwards from abstractions, effectively building a castle in the sky. A more robust approach would be to make these assumptions very explicit, ideally linking a theory to a set of axioms, so that we can better assess the defensibility of a theory. Some branches of continental philosophy are very bad at this (e.g. Lacan), starting from “metaphor” rather than an axiom, which is why I draw the parallel.
I will note that prosaic safety work could be relevant under a strong regulatory framework. For example, suppose we established an international treaty to freeze AI development at ChatGPT 5.5 Pro / Mythos. The treaty states that we can only advance to higher capabilities/intelligence when we are “sure” that the next model is aligned. With huge amounts of resources dedicated to verifying the next model if safe, it seems feasible to me that prosaic approaches could play a large role in building safe AI under such a regime.
Now, setting up sufficiently strong regulation is of course very hard, and one might critique that “proving” that the next generation of a model is aligned is akin to solving alignment itself! But I suspect that guaranteeing a single model is aligned is much easier than solving alignment for all possible models.
I would still guess it is better not to do prosaic safety work until a global regulatory framework exists, since it accelerates AI progress and thus reduces opportunities to implement said regulation. But there are enough counterarguments that I would be careful moralizing over it (not suggesting anyone in the comments is doing so!).
>> “Now, setting up sufficiently strong regulation is of course very hard, and one might critique that “proving” that the next generation of a model is aligned is akin to solving alignment itself! But I suspect that guaranteeing a single model is aligned is much easier than solving alignment for all possible models.”
Do you know how to prove ” alignment of a model”?
This is currently almost as hard as solving alignment for all possible models.
In fact, strictly speaking it is impossible—the model may have simply encrypted its thoughts and there is no method known to man that can break oneway functions.
Now maybe we make some assumption that we are in a good world where the AI didn’t cryptographically hide its thoughts. It doesn’t really help. The fundamental problem is that there is no behaviourial tests that can definitely exclude a sharp left turn/ treacherous turn/ deception/ name-du-jour. So there is no ′ proving’, just hope & cope and we’re back to prosaic alignment & evals again.
We are much more in agreement than I expected! Thank you for clarifying. I agree with every point you made in this comment. (Excepting the comparison to continental philosophy, which is new to me and not something I think I can evaluate one way or the other.)