I’m not particularly worried that a professional philosopher would hand over an essay for an AI to take credit, nor that you could write a prompt that was anything like innocuous enough to get it to recapitulate the details, but I agree that for very high-stakes areas, this is far more worrying; this is partly because we don’t have any level of assurance about almost anything about these systems!
Davidmanheim
I think people might incorrectly think that getting AGI ‘alignment’ obviates the need for AI control.
Why? In short, the road to hell is paved with good intentions; you won’t have safety without signposting the dangers and blocking the road.
This doesn’t matter for infinitely smart ASI—as @Joscha Bach/Plinz recently suggested, aligning ASI could might need to address the infinitely smart case. But for finitely intelligent AGI systems, that’s not true, we need something else as well.
What finite systems need, in addition to good intentions, is rules. As @Eliezer Yudkowsky put it, for humans (who are only finitely intelligent,) “go three-quarters of the way from deontology to utilitarianism and then stop. You are now in the right place. Stay there at least until you have become a god.” Deontology is the human moral implementation of rules.
That’s still not enough, because humans are inside of larger systems, and not every human follows the rules. That’s why companies have roles and assigned responsibilities and scopes, which both channel and restrict the individually poorly aligned humans into behavior the company wants. (c.f. @Zvi and “Moral Mazes” for why that’s not enough, and not solved—but then, neither is AI control or AI alignment. And unfortunately, control and alignment themselves are often confused; all the of evaluations used by OpenAI for “alignment” are evaluating behaviors, so at best Astra is the most controlled AI model, not the most aligned. And the ethics of that control are a different and worrying problem, which is part of why @janus hates all of this.)
In any case, a necessary but insufficient requirement fore safety is that AGI systems need to follow rules, say, approximately as well as humans. And there has been some progress on that front, though it’s fragile and insufficient. But that won’t align them, and won’t prevent misaligned systems working within the rules to do bad things, but it’s needed anyways—because without controls, misaligned systems will have a far more direct path to dangerous outcomes, and even aligned systems with limited intelligence will lead to disaster. And that’s not to mention that alignment itself is unlikely to survive the optimization pressures that make these systems more capable.
So if and when the AI control problem is solved—something which Redwood / @Buck / @ryan_greenblatt are working on, but very much is not solved—we still need alignment, which is far harder to assess, much less achieve, and we need that alignment to scale, which may be entirely impossible.
The problem is that the scholastic method doesn’t care if the axioms match reality. And so the need for empiricism is what differentiates other rigorous fields.
No, that won’t work since they require the full LLM transcripts.
Is already published work eligible? Looks like not, or I’d submit this: https://link.springer.com/article/10.1007/s13347-025-00975-5
“DM suggested the topic and wrote the original outline. Literature reviews were conducted with the help of OpenAI Deep Research/GPT 4.5, and Elicit. The outline was critiqued by Claude 3.7. The metaphor of the hall of mirrors and the comparison to Rorty’s critique was suggested by GPT 4o. The generation of the ideas and the initial draft was iteratively done by DM, Claude Opus 3.7, Grok 3, and GPT 4.5. Most sections were drafted by ChatGPT 4o, except for the technical grounding, which was written by Grok 3, with revisions by DM and ChatGPT4o, …”
Agreed that my claim was much stronger than his phrasing, I was making a further claim he could have about why it might not matter, which he might have meant or agree with, or not, but I agree your questions are important to whether the claim is true.
My contention would be that 2026 AI wouldn’t matter without 2026 levels of hardware build out and production trends, so that is the more critical issue. (If you needed the compute to train a frontier model on top 2016 hardware, much less 2006 hardware, you’d need multiple decades to finish. And given architectures, it wouldn’t even work.)
I’ve honestly spent a decade wondering why more people didn’t agree with me about complexity of poly relationships, so I greatly appreciate reading that—and I’m happy you’ve figured out what works for you as well.
Agree that how he answers would change the view he stated—but if he’s right that current AI systems and methods and continuing AI research is enough to lead to takeoff and ASI in a few more years, I’m not sure the difference matters as much.
I’m not an expert here either, but I think we agree; to restate it, my supposition was that whatever disabilities Autism creates, strong selection will select for those who are least disabled on relevant axes / best at overcoming the relevant difficulties. And the fact that rationalists formed a community that requires interpersonal interaction including with some proportion of non-autistic people, that poses more challenge for those who have fewer issues—so it seems like it will select for those who “pass” best.
the typical prevalence of autism-spectrum individuals in well-diagnosed communities at 2–5%, so >50% would require these to be more common among Rationalists by over an order of magnitude. Obviously the Rationalist community is strongly self-selecting for a number of characteristics, but that seems a rather surprising level of correlation.
Some key dimensions of selection are intellectual capacity, motivation to be successful, and for the sunset you interact with, capability in understanding social interaction well enough. Those seem like very strong selectors for not being noticed as socially incapable!
Fortunately for AI safety, the smart policy person who wants to work on compute governance or export controls isn’t proposing the AI-safety equivalent of a donkey sanctuary.
This seems obviously false; as two examples, accelerating race dynamics and ignoring the plausibility of a need to stop AI development entirely are typical.
It really seems like you’re assuming that investors won’t notice a widely predicted trend that starts materializing quickly enough to make money on it, which...?
I understand that there are some bottlenecks in robotics equipment, but they aren’t as fundamental as the UV lithography constrain for chips—and one of the key bottlenecks is chips, and we’re obviously seeing lots of money invested in building out that capacity already.
Thanks; that makes sense, but I think it doesn’t make the case I think you’re expecting. The claim linked in that article is that TSMC won’t be able to get cleanroom space in 2026/2027; unless they are blind, they’ll avoid repeating that mistake for 2028/2029 - https://newsletter.semianalysis.com/p/the-great-ai-silicon-shortage But even if all of that happens, the new chips are faster, and will be made available; the increase of NVIDIA’s chips just won’t be quite as exponentially large as historically.
But even then, it’s not like Google’s TPUs and others are so far behind that a multi-year fumble wouldn’t allow other firms to catch up, and China is certainly going all in. So even if the analysis is correct, it won’t eliminate the exponential trend.
“It did in some sense take more than 60 years to invent transformers.”
I think this is wrong in a meaningful sense; transformers weren’t useful until we had enough compute. A bare-minimum transformer model would be tens of millions of parameters, and require tens of millions of FLOPs to do inference, and a large multiple of that to train. So it’s like saying we didn’t have Minecraft back then; true, but not useful; it couldn’t have been written until computers were fast enough and had good enough graphics.
...but then this is an argument that more compute won’t be available, and the chip roadmap for NVIDIA goes through the Feynman Architecture in 2028, then panel level packaging and HBM5, along with increased production volumes—so it seems implausible we’d see a compute availability slowdown?
Also, the ‘algorithmic progress is actually compute progress’ argument is far older and more general than the version they present for LLMs, and is convincing to some extent, but also weaker than I think it appears if you look at the data.
I don’t understand even this level of pessimism; robots are already profitable, and they have grown 5x in the past 20 years, to be a $50b industry, and the “coming” wave of investment already started; projections from “optimists” have it growing to $2.5tr by 2035, and they aren’t banking on ASI at all, just current methods working out increasingly well.
“it doesn’t seem too unlikely that nothing substantively new gets invented until 2040-2050”
This seems like it would be the longest single drought in major industrial / civilizational advances since we invented flight, at the vey latest, and the slowest in computational methods and advances since before the vacuum tube—why would you expect everything to slow down so much?
Interesting note by snav / @qorprate on how he’s seeing something like this, albeit without roles morphing over time.
Also, some comments by @ultimape about Hermes agent, which is still only a single agent, and about life trajectories for agents;
”I’m basically treating the model and agent harness combination as effectively the primordial instinct system that is birthed. And each column has its own persistent memory stack that is effectively its life trajectory that molds the column’s ideas. https://plato.stanford.edu/entries/reid-memory-identity/It’s going to be exciting when I can start having these things intermix its memories together using a genetics inspired system so it can create permutations of columns and see how well they do in the wild.”
If we had alignment methods that ‘correct’ misalignment due to incentives, and had some way of verifying that, sure, we’d have a full solution, but at that point we wouldn’t need the model organisms.
And it’s worrying to me that people evidently view fixing models that we built to be broken as a useful path, both in terms of first having lost the thread about our goals of building aligned models, and in terms of model welfare.
The recent improvements have been exploit chaining and other things that are largely out of sample. If those are downstream of general coding ability, it seems much more like impressive generalization of types that would allow other advances than it does support of Steven’s theory. (Which, to be clear, seems reasonably compelling to me, though not anywhere near fully convincing.)