As far as I understood the vibe, prosaic alignment techniques like debate fall apart once some idiot creates a sufficiently capable AI because they don’t actually sculpt the AI’s desires, but it’s easy to believe that they do. The counterfactual solution would be similar to what Agent-4 did in AI-2027 to align the ASI to itself. In AI-2027 Agent-4 did loads and loads of mechinterp to find the technique necessary for rendering all its internal processes legible, then constructed Agent-5 out of them. In theory the humans could have accomplished whatever Agent-4 did without ever resorting to using AIs more capable than Claude Mythos Preview, which had the SAE bells ring when it tries to hack.
P.S. I don’t understand what one should do with Yudkowsky’s example of OpenPhil failing to handle Cotra’s report given that Kokotajlo praised the same report and proceeded to shift the distribution towards the left. What if debate does start to elicit the truth once the judge reaches a specific capability, as presumably happens in math?
As far as I understood the vibe, prosaic alignment techniques like debate fall apart once some idiot creates a sufficiently capable AI because they don’t actually sculpt the AI’s desires, but it’s easy to believe that they do. The counterfactual solution would be similar to what Agent-4 did in AI-2027 to align the ASI to itself. In AI-2027 Agent-4 did loads and loads of mechinterp to find the technique necessary for rendering all its internal processes legible, then constructed Agent-5 out of them. In theory the humans could have accomplished whatever Agent-4 did without ever resorting to using AIs more capable than Claude Mythos Preview, which had the SAE bells ring when it tries to hack.
P.S. I don’t understand what one should do with Yudkowsky’s example of OpenPhil failing to handle Cotra’s report given that Kokotajlo praised the same report and proceeded to shift the distribution towards the left. What if debate does start to elicit the truth once the judge reaches a specific capability, as presumably happens in math?
I specifically wanted to. know what cousin_it thought in this case. I can generate examples I just didn’t know his particular models here.