I think you’re overestimating the evidence that a senior academic being listed as a co-author on a paper provides about anything, especially a high-flying one whose prestige people are keen to add to their paper. It’s common for a prestigious lab leader’s name to be listed on papers that they didn’t have much to do with.
If he is actually strongly against neuralese (he has definitely created that impression in the past), he should say “no, don’t list me as a co-author” in a situation like this.
Of course, I am assuming that he has known what the paper is about and that he is not signing his name on papers he is not familiar with (if that assumption is wrong, that would be bad in a different way).
The point is that if we want “accountability”, the person’s signature must mean something. If the person’s signature does not mean anything, what kind of accountability could we talk about?
Especially given that he is not just a senior and influential academic and a prominent voice in all that, but he is leading a prominent AI safety organization, he is a co-president and scientific director here: https://lawzero.org/en. His decisions might actually matter. And there are disagreements about the feasibility of the path he is advocating. So, in his case, it’s rather important to have some clarity on where he stands.
Sure, I think it’s totally reasonable to ask for better norms around when academics put their names on papers.
I‘m just saying under ML academia norms as I’ve seen them, the amount of evidence you should take about a senior academic‘s views from a random paper they’re listed on is not that much.
Zooming out back to the original point, I don’t think that this is really any reason to become less optimistic about frontier labs coordinating to not do neuralese.
If even a prominent AI safety person who has been very much against neuralese is not sufficiently firm about this, then what happens when industry people have much stronger incentives than this, and when we see more and more indications of neuralese-induced capability boosts in the literature?
(My expectation is that they would develop models translating neuralese to what we understand, and would admit that this does not provide full guarantees, but would assert that merely reading the superficial layer of chain-of-thought does not provide full guarantees either. It would be nice to be wrong, but this is the default outcome. (If I understand it correctly, the recent trend is for reduced superficial clarity even with the token-based chain-of-thought, either due to stronger training pressures, or due to more emphasis on interagent communication, or who knows why.))
I think you’re overestimating the evidence that a senior academic being listed as a co-author on a paper provides about anything, especially a high-flying one whose prestige people are keen to add to their paper. It’s common for a prestigious lab leader’s name to be listed on papers that they didn’t have much to do with.
If he is actually strongly against neuralese (he has definitely created that impression in the past), he should say “no, don’t list me as a co-author” in a situation like this.
Of course, I am assuming that he has known what the paper is about and that he is not signing his name on papers he is not familiar with (if that assumption is wrong, that would be bad in a different way).
The point is that if we want “accountability”, the person’s signature must mean something. If the person’s signature does not mean anything, what kind of accountability could we talk about?
Especially given that he is not just a senior and influential academic and a prominent voice in all that, but he is leading a prominent AI safety organization, he is a co-president and scientific director here: https://lawzero.org/en. His decisions might actually matter. And there are disagreements about the feasibility of the path he is advocating. So, in his case, it’s rather important to have some clarity on where he stands.
Sure, I think it’s totally reasonable to ask for better norms around when academics put their names on papers.
I‘m just saying under ML academia norms as I’ve seen them, the amount of evidence you should take about a senior academic‘s views from a random paper they’re listed on is not that much.
Zooming out back to the original point, I don’t think that this is really any reason to become less optimistic about frontier labs coordinating to not do neuralese.
I am not sure.
If even a prominent AI safety person who has been very much against neuralese is not sufficiently firm about this, then what happens when industry people have much stronger incentives than this, and when we see more and more indications of neuralese-induced capability boosts in the literature?
(My expectation is that they would develop models translating neuralese to what we understand, and would admit that this does not provide full guarantees, but would assert that merely reading the superficial layer of chain-of-thought does not provide full guarantees either. It would be nice to be wrong, but this is the default outcome. (If I understand it correctly, the recent trend is for reduced superficial clarity even with the token-based chain-of-thought, either due to stronger training pressures, or due to more emphasis on interagent communication, or who knows why.))