Thanks for that comment. One thing I didn’t include in the agenda, which I probably should have, is more nuanced failure modes of current AI safety agentic research outside of the traditional reward-hacking and incompetence. This are useful patterns to be aware of and work to eliminate.
Thanks for that comment. One thing I didn’t include in the agenda, which I probably should have, is more nuanced failure modes of current AI safety agentic research outside of the traditional reward-hacking and incompetence. This are useful patterns to be aware of and work to eliminate.