...and that’s only people in developed countries. Most of the global population is far worse off.
Davidmanheim
I think that there are very, very few examples of exploit chaining compared the the number needed to learn a skill in pretraining, given sample inefficiency. So I’m claiming that the way it’s accomplishing that is either via generalized skills based on too few data points to things very far outside what it’s seen, which implies a level of capability we don’t see in similar domains, or via RL.
(They also make a clear implication that it would be irresponsible for Anthropic to go public. I don’t particularly see strong arguments for this, could someone explain?)
One part is presumably generic distrust of short term market incentives, which I think is a reasonable concern. Specifically, it changes management incentives from maximizing long term value into tracking impact on stock prices. Relatedly, it gives investors and shareholders a clearer legal mechanism to attack changes that trade off human survival and short term growth, which includes any pause or pacing agreements.
Do we know that the improvements in cyberoffense aren’t just downstream of improvements in programming generally?
The recent improvements have been exploit chaining and other things that are largely out of sample. If those are downstream of general coding ability, it seems much more like impressive generalization of types that would allow other advances than it does support of Steven’s theory. (Which, to be clear, seems reasonably compelling to me, though not anywhere near fully convincing.)
I’m not particularly worried that a professional philosopher would hand over an essay for an AI to take credit, nor that you could write a prompt that was anything like innocuous enough to get it to recapitulate the details, but I agree that for very high-stakes areas, this is far more worrying; this is partly because we don’t have any level of assurance about almost anything about these systems!
I think people might incorrectly think that getting AGI ‘alignment’ obviates the need for AI control.
Why? In short, the road to hell is paved with good intentions; you won’t have safety without signposting the dangers and blocking the road.
This doesn’t matter for infinitely smart ASI—as @Joscha Bach/Plinz recently suggested, aligning ASI could might need to address the infinitely smart case. But for finitely intelligent AGI systems, that’s not true, we need something else as well.
What finite systems need, in addition to good intentions, is rules. As @Eliezer Yudkowsky put it, for humans (who are only finitely intelligent,) “go three-quarters of the way from deontology to utilitarianism and then stop. You are now in the right place. Stay there at least until you have become a god.” Deontology is the human moral implementation of rules.
That’s still not enough, because humans are inside of larger systems, and not every human follows the rules. That’s why companies have roles and assigned responsibilities and scopes, which both channel and restrict the individually poorly aligned humans into behavior the company wants. (c.f. @Zvi and “Moral Mazes” for why that’s not enough, and not solved—but then, neither is AI control or AI alignment. And unfortunately, control and alignment themselves are often confused; all the of evaluations used by OpenAI for “alignment” are evaluating behaviors, so at best Astra is the most controlled AI model, not the most aligned. And the ethics of that control are a different and worrying problem, which is part of why @janus hates all of this.)
In any case, a necessary but insufficient requirement fore safety is that AGI systems need to follow rules, say, approximately as well as humans. And there has been some progress on that front, though it’s fragile and insufficient. But that won’t align them, and won’t prevent misaligned systems working within the rules to do bad things, but it’s needed anyways—because without controls, misaligned systems will have a far more direct path to dangerous outcomes, and even aligned systems with limited intelligence will lead to disaster. And that’s not to mention that alignment itself is unlikely to survive the optimization pressures that make these systems more capable.
So if and when the AI control problem is solved—something which Redwood / @Buck / @ryan_greenblatt are working on, but very much is not solved—we still need alignment, which is far harder to assess, much less achieve, and we need that alignment to scale, which may be entirely impossible.
The problem is that the scholastic method doesn’t care if the axioms match reality. And so the need for empiricism is what differentiates other rigorous fields.
No, that won’t work since they require the full LLM transcripts.
Is already published work eligible? Looks like not, or I’d submit this: https://link.springer.com/article/10.1007/s13347-025-00975-5
“DM suggested the topic and wrote the original outline. Literature reviews were conducted with the help of OpenAI Deep Research/GPT 4.5, and Elicit. The outline was critiqued by Claude 3.7. The metaphor of the hall of mirrors and the comparison to Rorty’s critique was suggested by GPT 4o. The generation of the ideas and the initial draft was iteratively done by DM, Claude Opus 3.7, Grok 3, and GPT 4.5. Most sections were drafted by ChatGPT 4o, except for the technical grounding, which was written by Grok 3, with revisions by DM and ChatGPT4o, …”
Agreed that my claim was much stronger than his phrasing, I was making a further claim he could have about why it might not matter, which he might have meant or agree with, or not, but I agree your questions are important to whether the claim is true.
My contention would be that 2026 AI wouldn’t matter without 2026 levels of hardware build out and production trends, so that is the more critical issue. (If you needed the compute to train a frontier model on top 2016 hardware, much less 2006 hardware, you’d need multiple decades to finish. And given architectures, it wouldn’t even work.)
I’ve honestly spent a decade wondering why more people didn’t agree with me about complexity of poly relationships, so I greatly appreciate reading that—and I’m happy you’ve figured out what works for you as well.
Agree that how he answers would change the view he stated—but if he’s right that current AI systems and methods and continuing AI research is enough to lead to takeoff and ASI in a few more years, I’m not sure the difference matters as much.
I’m not an expert here either, but I think we agree; to restate it, my supposition was that whatever disabilities Autism creates, strong selection will select for those who are least disabled on relevant axes / best at overcoming the relevant difficulties. And the fact that rationalists formed a community that requires interpersonal interaction including with some proportion of non-autistic people, that poses more challenge for those who have fewer issues—so it seems like it will select for those who “pass” best.
the typical prevalence of autism-spectrum individuals in well-diagnosed communities at 2–5%, so >50% would require these to be more common among Rationalists by over an order of magnitude. Obviously the Rationalist community is strongly self-selecting for a number of characteristics, but that seems a rather surprising level of correlation.
Some key dimensions of selection are intellectual capacity, motivation to be successful, and for the sunset you interact with, capability in understanding social interaction well enough. Those seem like very strong selectors for not being noticed as socially incapable!
Fortunately for AI safety, the smart policy person who wants to work on compute governance or export controls isn’t proposing the AI-safety equivalent of a donkey sanctuary.
This seems obviously false; as two examples, accelerating race dynamics and ignoring the plausibility of a need to stop AI development entirely are typical.
It really seems like you’re assuming that investors won’t notice a widely predicted trend that starts materializing quickly enough to make money on it, which...?
I understand that there are some bottlenecks in robotics equipment, but they aren’t as fundamental as the UV lithography constrain for chips—and one of the key bottlenecks is chips, and we’re obviously seeing lots of money invested in building out that capacity already.
Thanks; that makes sense, but I think it doesn’t make the case I think you’re expecting. The claim linked in that article is that TSMC won’t be able to get cleanroom space in 2026/2027; unless they are blind, they’ll avoid repeating that mistake for 2028/2029 - https://newsletter.semianalysis.com/p/the-great-ai-silicon-shortage But even if all of that happens, the new chips are faster, and will be made available; the increase of NVIDIA’s chips just won’t be quite as exponentially large as historically.
But even then, it’s not like Google’s TPUs and others are so far behind that a multi-year fumble wouldn’t allow other firms to catch up, and China is certainly going all in. So even if the analysis is correct, it won’t eliminate the exponential trend.
“It did in some sense take more than 60 years to invent transformers.”
I think this is wrong in a meaningful sense; transformers weren’t useful until we had enough compute. A bare-minimum transformer model would be tens of millions of parameters, and require tens of millions of FLOPs to do inference, and a large multiple of that to train. So it’s like saying we didn’t have Minecraft back then; true, but not useful; it couldn’t have been written until computers were fast enough and had good enough graphics.
...but then this is an argument that more compute won’t be available, and the chip roadmap for NVIDIA goes through the Feynman Architecture in 2028, then panel level packaging and HBM5, along with increased production volumes—so it seems implausible we’d see a compute availability slowdown?
Also, the ‘algorithmic progress is actually compute progress’ argument is far older and more general than the version they present for LLMs, and is convincing to some extent, but also weaker than I think it appears if you look at the data.
Wouldn’t GPUs in pods without high-bandwidth interconnects be enough to allow inference without allowing training?