Sure, I think it’s totally reasonable to ask for better norms around when academics put their names on papers.
I‘m just saying under ML academia norms as I’ve seen them, the amount of evidence you should take about a senior academic‘s views from a random paper they’re listed on is not that much.
Zooming out back to the original point, I don’t think that this is really any reason to become less optimistic about frontier labs coordinating to not do neuralese.
If even a prominent AI safety person who has been very much against neuralese is not sufficiently firm about this, then what happens when industry people have much stronger incentives than this, and when we see more and more indications of neuralese-induced capability boosts in the literature?
(My expectation is that they would develop models translating neuralese to what we understand, and would admit that this does not provide full guarantees, but would assert that merely reading the superficial layer of chain-of-thought does not provide full guarantees either. It would be nice to be wrong, but this is the default outcome. (If I understand it correctly, the recent trend is for reduced superficial clarity even with the token-based chain-of-thought, either due to stronger training pressures, or due to more emphasis on interagent communication, or who knows why.))
Sure, I think it’s totally reasonable to ask for better norms around when academics put their names on papers.
I‘m just saying under ML academia norms as I’ve seen them, the amount of evidence you should take about a senior academic‘s views from a random paper they’re listed on is not that much.
Zooming out back to the original point, I don’t think that this is really any reason to become less optimistic about frontier labs coordinating to not do neuralese.
I am not sure.
If even a prominent AI safety person who has been very much against neuralese is not sufficiently firm about this, then what happens when industry people have much stronger incentives than this, and when we see more and more indications of neuralese-induced capability boosts in the literature?
(My expectation is that they would develop models translating neuralese to what we understand, and would admit that this does not provide full guarantees, but would assert that merely reading the superficial layer of chain-of-thought does not provide full guarantees either. It would be nice to be wrong, but this is the default outcome. (If I understand it correctly, the recent trend is for reduced superficial clarity even with the token-based chain-of-thought, either due to stronger training pressures, or due to more emphasis on interagent communication, or who knows why.))