It’s not only about which risks get prioritized. It’s also about who is advocating for that. For most leftists, the AI labs are firmly at the top of the existing power structure, same as all large corporations. AI labs are currently advocating to focus on X-risks. When someone I believe to be acting against my best interests is telling me something, it’s very natural to reject that view by default.
rcc
While I personally agree that the EU doesn’t have much leverage in preventing the doomsday scenarios, a “young talented graduate” also doesn’t, no matter where they are.
But just preventing the doom is not where the journey ends. Lots of work will be needed for AI to go right across the board. Working at a relevant EU institution doesn’t seem like the worst way to contribute.
I think the unintuitive part is that AIs care a lot less about the potential negative consequences than humans. We do have a lot of stories and cultural tropes about people hurting others to get money or power—that is intuitive but it’s not really the sort of instrumental convergence that is relevant for most AI discussions today.
We do not have stories and tropes about humans hacking corporations and governments for the sake of doing better on an exam[1]. For an average person this sort of behavior doesn’t make sense, even if the goal is instrumental. The expected value appears highly negative, due to the risk and impact of being detected, compared to what is gained.
- ^
As a sidenote: we also do not have any cultural model that any exam should be existentially critical for someone, in a way it appears to be for the agents.
- ^
It’s hard to run a business as an amnesiac, even if you keep good notes.
With the physical limitations of a human, yes. It’s a different story if you could read and write those notes instantly at any time. I’d expect AI to still benefit from internalizing learning but I doubt that the penalty of not doing so is sufficient for humans to maintain competitive advantage.
Out of curiosity—why such a long delay between collecting the data and publishing the results?
Who, except maybe rationalists, would believe (or give money to) the Trustworthy Institution rather than an Ideological Institution Supporting Their Worldview? I don’t think there’s a market for it.
Relevant xkcd:
tbh I don’t understand your point. I feel like you’re trying to gesture at an internal contradiction in my post against things I did not say? Which feels like an odd standard to be held to.
I assume you are talking about this:
You say that “no current publicly available model is known to use neuralese, and the theoretical benefits have not really been demonstrated or realized.” At the same time you make a pretty compelling argument that Astra is using neuralese.
I agree that it’s technically not a contradiction—your claim is about certainty and your argument is about probability. To state clearly what I was gesturing at: you seem to be making an isolated demand for rigor—you treat probability as good enough through your entire article, except for this one part where you switch to demanding certainty in a way that strengthens your argument. This is inconsistent epistemics.
I’d be quite surprised that the current contribution of neuralese is more than a point, and I’d guess lower (possibly much lower). [...] the magnitude is extremely unclear, and likely low to date.
Seems to me that we just have different levels of certainty as to the benefits and different world models for how the decision was reached at OpenAI. Until more evidence is available, I think all we can do is present our own arguments and let the readers decide which resonates more.
To the extent I have any marginal contribution or levers to pull on the AI companies other than pause/not pause, it seems like a pretty good first-order guess that I want them to pay less safety costs per unit of capabilities gains, rather than more.
Then it seems that I have misread your intent, apologies. I still believe that this framing weakens the overall argument. If anyone is able to show real capabilities gains from using neuralese, the argument as framed concedes. I think your case against the costs is strong enough to stand on its own, avoiding this problem.
There’s a major difference between the two.
AI Safety went mainstream because it reinforced a preexisting bias against AI—people afraid for their jobs, tired of slop and convinced that it’s bad for environment. There was already a significant (even if highly fragmented) anti-AI movement before the general population started seriously considering x-risk.
Longevity won’t be benefitting from the preexisting bias against death. Culturally there’s no such thing. We don’t have people protesting against death and writing articles about how bad death is.
What we have are preexisting biases against medical interventions (including well-proven ones, like vaccines), “playing God”, wealth concentration and gerontocracy. Longevity will have to fight an uphill battle to overcome all of those.
Nvidia is approximately as involved in the overall dynamic.
Is it, actually? GPUs are more or less a commodity these days. One that is used for both good purposes (e.g. current LLMs which contribute to economic output) and bad purposes (i.e. training new AIs that might pose danger).
What would the defectors advocate for?
Even if NVIDIA wanted to stop the flow of new GPUs to the “bad purpose” use cases, they wouldn’t be able to do so. Only the US Government could actually enforce such decision.
Stopping the production of GPUs entirely would be both illegal (acting against the interests of shareholders) and doing significant harm to the global economy. It also would also, at best, only slow down the frontier labs if they were still keen on reckless development.
Some level of swarm behavior will always be necessary for some use cases due to pure physics. If your compute is spatially distributed (e.g. placed into physical robots) then the bandwidth available between compute instances will typically be significantly lower and less reliable than the bandwidth available within a compute instance.
For an easy example, just imagine a swarm of military UAVs attacking an enemy compound. You can only pack so much compute into a single UAV and the quality of the data link between any two UAVs will vary dynamically, including the possibility of losing the signal entirely. The only possible solution here is for each compute instance to run its own end-to-end thinking and coordinate with other instances separately, i.e. a multi-agent system.
As I understand this proposal boils down to creating more auditable artifacts during the operation of an AI system. I’m not sure we actually need more auditable artifacts today. We already have conversation logs, CoT and tool call logs available. All recent incidents could have been detected from those artifacts alone. The gap appears to lie primarily in the willingness of frontier AI labs to spend resources on real-time detection and share the findings, not in the lack of data to analyze.
I fully agree with the claim that moving away from CoT is bad but I think that your claim that it’s for “dubious benefits” is quite weak and counterproductive to the overall argument.
You say that “no current publicly available model is known to use neuralese, and the theoretical benefits have not really been demonstrated or realized.” At the same time you make a pretty compelling argument that Astra is using neuralese. We also ~know[1] that Astra is notably better than any previous model. Sure, correlation does not imply causation but in this case it does gesture very suggestively. We know that:
OpenAI is strongly incentivized to improve the capabilities of their models.
The theoretical mechanism for neuralese improving capabilities seems sensible.
Using neuralese would need to be an intentional design decision, not something that happens accidently.
Ablation studies are routinely conducted during development of new models.
Even if one doesn’t care about safety, using neuralese comes with PR costs and OpenAI seems aware of that.
If, given all of this, OpenAI decided to use neuralese for Astra, I have a pretty high confidence that it is because their internal studies have shown causal relationship to improving capabilities. Otherwise they are eating the PR cost with no benefit.
Even if you disagree with this reasoning, contrasting the “dubious benefits” with decreased monitoring positions the whole argument as a trade off. It implies that it could be acceptable to make a model less safe if the capability increase would be sufficiently large and I doubt that is your intent.
- ^
Assuming status quo where we trust OpenAI claims about benchmark results.
What’s the value of assessing Jev on those types of questions? Those are clearly not the use cases it was designed for.
It feels very weird to test the model on capabilities it doesn’t claim to posses, then draw a conclusion that it “has a very jagged frontier”. If you asked Claude to generate photorealistic images[1] you could also conclude that it has a jagged frontier compared to ChatGPT
As of writing this comment, to my knowledge, Claude can only directly generate vector graphics