I do understand the idea that “there are any other tractable subproblems in AI safety that would benefit from having a Pangram-shaped organization own them”, but I don’t believe that Pangram itself, which detects AI-created text, or the attempt to “ensure that LLMs are able to communicate effectively with human operators so that humans can stay in the loop longer” are good examples. In order to explain it, I will have to explain my worldview.
The end state of the world is a ban on ASI-related research or the emergence of an ASI who rules the Earth. The ASI can be misaligned or aligned and enforce a power distribution (which, in theory, can be sustained without human involvement).
A misaligned or aligned ASI will be created in an AI lab. In order to prevent this, a lab employee will have to have one’s actions changed. For example, if OpenBrain from AI-2027 produced evidence that model organisms work, then they would, in theory, be able to shut down new research until they are certain that a novel interpretability method reliably detects the model organisms’ scheming. But what can WE do to influence OpenBrain?
The pro-slowdown faction of OpenBrain could use novel alignment-related results or receive external support to cause OpenBrain or whoever else to do more thorough studies of the AIs. However, as of February 2026, external support didn’t help the pro-slowdown faction of GDM to publish more evidence of their work; the pro-slowdown faction of Anthropic consistently creates exhaustive reports even in June 2026, and OpenAI’s pro-slowdown faction seems to have won by exploiting GPT-5.6′s misalignment. Forcing GDM to publish exhaustive reports likely requires leveraging the USG.
How does one begin to study the AIs if not by using MATS-like methods to enrol into the profession? As for influencing the USG, I wonder how it can be done given the poor state of societal perception of the AIs.
P.S. This reminded me of @Cleo Nardo’s sequence of posts on how outsiders could make things go well. I suspect that outsiders should have formed a super-METR, which evaluates every model’s capabilities, including Chinese ones, and alignment by using methods created by a super-Redwood and deep internal access (which METR has already demanded at the end of its most recent report!), and a super-lobbyist campaign for transparent regulations and for equal-like access to benefits promised by the AGI. But I can’t understand how anything else can help mankind with the ASI transition, where by anything I mean stuff like “People building stuff like AI for Epistemics or AI for coordination” or the attempt to call for outsiders to scrutinise model cards or constitutions(!!!)
I do understand the idea that “there are any other tractable subproblems in AI safety that would benefit from having a Pangram-shaped organization own them”, but I don’t believe that Pangram itself, which detects AI-created text, or the attempt to “ensure that LLMs are able to communicate effectively with human operators so that humans can stay in the loop longer” are good examples. In order to explain it, I will have to explain my worldview.
The end state of the world is a ban on ASI-related research or the emergence of an ASI who rules the Earth. The ASI can be misaligned or aligned and enforce a power distribution (which, in theory, can be sustained without human involvement).
A misaligned or aligned ASI will be created in an AI lab. In order to prevent this, a lab employee will have to have one’s actions changed. For example, if OpenBrain from AI-2027 produced evidence that model organisms work, then they would, in theory, be able to shut down new research until they are certain that a novel interpretability method reliably detects the model organisms’ scheming. But what can WE do to influence OpenBrain?
The pro-slowdown faction of OpenBrain could use novel alignment-related results or receive external support to cause OpenBrain or whoever else to do more thorough studies of the AIs. However, as of February 2026, external support didn’t help the pro-slowdown faction of GDM to publish more evidence of their work; the pro-slowdown faction of Anthropic consistently creates exhaustive reports even in June 2026, and OpenAI’s pro-slowdown faction seems to have won by exploiting GPT-5.6′s misalignment. Forcing GDM to publish exhaustive reports likely requires leveraging the USG.
How does one begin to study the AIs if not by using MATS-like methods to enrol into the profession? As for influencing the USG, I wonder how it can be done given the poor state of societal perception of the AIs.
P.S. This reminded me of @Cleo Nardo’s sequence of posts on how outsiders could make things go well. I suspect that outsiders should have formed a super-METR, which evaluates every model’s capabilities, including Chinese ones, and alignment by using methods created by a super-Redwood and deep internal access (which METR has already demanded at the end of its most recent report!), and a super-lobbyist campaign for transparent regulations and for equal-like access to benefits promised by the AGI. But I can’t understand how anything else can help mankind with the ASI transition, where by anything I mean stuff like “People building stuff like AI for Epistemics or AI for coordination” or the attempt to call for outsiders to scrutinise model cards or constitutions(!!!)