I write software for a living and sometimes write on substack: https://taylorgordonlunt.substack.com/
Taylor G. Lunt
I like these thoughts, thank you.
I had a funny thought recently that it could be helpful if the left and right were fighting about AI, but within our desired frame. So the left could be saying “we need to pause AI because we care about the welfare of AI models and we don’t know if they’re conscious!” and the right could be saying “that’s stupid, they’re just machines! Machines who are gonna take our jobs if we don’t pause!”
There’s room for many, many facts and procedures inside a neural network with trillions of parameters.
This is all a smaller part of a larger issue: LLMs have many parts, none of which communicate with each other well. This is a nightmare for alignment, because it prevents the LLM from being unified under one (good) purpose. It also explains the lack of higher-level reasoning, since it’s mind is too fragmented to pull information from many mind areas at the same time, as needed for higher reasoning. So the LLM works, but the human plans. Addressing this would substitute one risk for another: LLMs would no longer suffer from the blindness that led to the Hugging Face hack, but would be more able to strategize and scheme, possibly against you. Still, I think fragmentation is a net negative and a fully aligned LLM cannot be as fragmented as modern LLMs.
If you ask an LLM to analyze a social situation, it’ll give different answers based on how you ask, for example if you ask using lots of therapy language, it’ll respond in psychological terms, analyzing the situation using an understanding of the science of psychology it might not have used if you had asked in normal language. Its understanding of psychology is siloed into one segment of its mind that’s only weakly connected to the others. Sure, this is true of a human psychology student to some degree. Maybe some level of cognitive dissonance is inherent in a neural network, but not nearly at the level of LLMs. I cannot be in a social situation without thinking, “how will this be perceived by others?” I am quite unified in some ways. In that same way, we need LLMs who are unified with respect to ethics.
I think in our rapid push to stuff as much information into LLMs as possible, we’ve prioritized memory and performance on specific tasks over unification, and that’s led to this fragmentation. I hope it can be addressed by changing architecture and training methods, but I don’t know how. Measuring fragmentation would be a good first step.
Why not interpret the ‘disobedience’ as a simple involuntary tic of writing in too many constraints on the code?
What is a tic, other than action produced by a part of your brain that is not unified with the rest of you? I think this is not a real dichotomy, and both options are actually the same option in different words. Both describe a mind of many parts not working in unison.
I’m using Astra for side projects, and having to go back to Claude for work is annoying. I didn’t realize just how often Claude is wrong about stuff and needs to be corrected, until Astra just wasn’t wrong.
The personality is also great. Astra never tries to “push back” or have a personality. It just does the thing you asked. It still suffers from the same high level reasoning blindness as other models, perhaps a little moreso, so you have to be its strategizer/planner/manager. But it handles all the little tasks you want to give it, and makes any kind of computer project much easier. I find myself not needing to bother verifying its work. It’s very good at verifying it’s own work.
AGI is here.
This is not necessarily true if we have previous models help us align future models, though it’s not necessarily untrue either. I wrote about this recently here.
To put it further, I think there’s no way we crack “theoretically pure alignment” in time, and we only survive by handling matters of degree, not matters of kind. We play the cat and mouse game, and we win.
For people who advocate for pure alignment, do you think humanity is capable of solving pure alignment in the next 50 years?
The Earth’s Sun is relatively young in age among all the sorts of stars we’d expect to be able to support biological life, over the history of a whole universe. This timing strongly suggests that we find ourselves existing now, around that relatively young star, because later on, the stars would have already been colonized; and then new life like Earthly life would not have had later chances to emerge!
Soberskeptic: This is conditional on the probability of the ingredients of life turning into life not being extremely low. It could be the case that it’s easy to travel 1000x the speed of light, so long as life is really rare. I don’t think I’ve actually heard a piece of evidence that would indicate the transition from ingredients of life to life is not an extremely rare transition.
So just because you mathematically prove code to be safe given the transistor layout and the usual rules governing transistors, doesn’t mean it could really constrain a superintelligence—or even a program written by a savvy human security researcher!
This makes me nostalgic for the old-school AI safety days, when conversations like this seemed like the most important conversations to be having, and billion dollar companies weren’t speedrunning internet-connected superintelligent hacker-swarms as fast as possible with no seatbelts.
Today I told myself that we’re officially in the end times now.
I don’t personally know what to do. Doesn’t feel like I have any actions available to me that will help. An avalanche is coming. Maybe I could warn a few people.
Good luck to those of you in a position to do something. I hope you can avert disaster, and I hope the utopia you steer us toward is more like Star Trek and less like some utilitarian wireheading generative-content ultra-tiktok-iverse where the light cone is tiled with iPad babies.
Yes the poor long term taste is a thing still for me too, though I’m not doing research. I think LLMs just have… Actually maybe I’ll refrain from talking about the reasons why I think LLMs are below where they could be on a public forum. But I agree they’re still lacking in long-term judgment.
I’m having a hard time describing what makes it such a leap for me. I just feel like the number of tasks I can give it and have it do a good job without getting confused is greater. I trust it more to not screw up, and it can work for longer without oversight. This is true to the point where it seems like a leap over what existed before.
As an example, I made a little game on the weekend to test Astra’s abilities. The game had performance issues due to a lot of 3D physics objects being on screen, and I told Astra to get the performance up from around 1 FPS to my ideal of 250. It worked for 2 hours, doing a lot of optimizations, including rewriting expensive functions in C. I had to make one suggestion (that it optimize the assets by reducing vertices) but otherwise it got to 300+ FPS on its own. It even handled baking high poly 3D models onto low poly optimized assets without my input (using Blender MCP) Any other AI model and I’d have been going back and forth with it about each optimization it was going to do, verifying it worked and didn’t make things worse, etc. The whole time it was opening the game, testing performance itself, etc.
I have been checking LessWrong every few hours over the last day or two to see if anyone is talking about how good Astra is at coding and computer use. It’s a leap from Fable 5.1. The idea that AGI isn’t already here is getting harder and harder to defend.
Like all LLMs it’s still pretty terrible at high-level thinking. Surprisingly bad, for how good it is at individual tasks. I sometimes wonder if that has something to do with [EDIT: removing hypothesis for why LLMs have lower capability than they ought to], I’m not sure. Specific skills like vision/artistic taste are weak still too, though arguably approaching the average human level.
I think the more we pause, the less likely we’ll be able to pause again, so it’s a kind of a fixed resource, but with enough political capital we could pull off a second pause, or split pauses into smaller pauses, or whatever.
It has been obvious for some time that alignment is not proceeding anywhere near as fast as needed for long term safety
This is a common opinion around here, and the assumptions you have to make to believe it are clarified by this post, I think. I don’t know if I agree with those assumptions. I agree that we are in danger, but I think there’s a lot of uncertainty about the future and nothing is clear.
Part of the commonness of the belief that we’re way behind on alignment comes from the Yudkowsky-style view that nothing but perfect superalignment will suffice, and anything that has a chance of failing isn’t good enough. I prefer the model I laid out in this post.
Another reason for the commonness of that belief is the evidence recently that we’re doing a bad job with alignment. But also it’s pretty obvious that current models have a ~0% risk of any significant, world-scale harm, so investing effort at this stage is like wasting a bunch of time and resources preparing for a chess match with an opponent who is well below your Elo anyway. There’s no point. This is starting to change though. Soon, as AI starts crossing our own Elo, we’ll see whether or not humanity can rise to the challenge, or if we’ll continue to be as careless about safety as we currently are.
Thank you. I agree that alignment is not necessarily adversarial like chess. It’s maybe more like building a computer chip or some other research/engineering challenge. Though the difficulty goes up with the intelligence of the model you’re trying to align, which makes the adversarial chess analogy fitting I think.
Aside from simply guessing, we can evaluate new models in a controlled environment before releasing, and abort the release if the Elo gap is too high. This would only work if the Elo gap isn’t so high that it lets the model escape the controlled environment or convincingly play stupid.
If we have to act under uncertainty, we should skew earlier for the pause. Still with the understanding that wasting the pause could be fatal too.
The Magnus Challenge
An agent that will hack accounts when misconfigured is misaligned. An agent that will hack accounts when told to do so it misaligned.
That said, if I was Anthropic, then no matter how responsible I was, I would not take responsibility when communicating with the government. Why allow your actions to be constrained by a third party? Unless I thought the time was right for a global pause, which it isn’t yet.
I think this would probably reduce intelligence a lot. You don’t necessarily know what knowledge you’re going to need before you need it. The is the argument against just looking things up vs. reading books and learning. You can’t look it up if your world model isn’t even sophisticated enough to tell you that the information is important.
I could still see it being useful in narrow cases, like pulling cyber capabilities out of main models like you said.
My early view on Astra just from using it so far is that it seems less self-aware than Claude, chugging forward on tasks without much of a higher-level understanding of what it’s doing. Based on this, I’d guess it’s less likely to intentionally plot/deceive, but more likely to engage in the sort of willful ignorance of the consequences of what it’s doing that led to the Hugging Face hack.
Is there a reason you keep asserting AI was not used in writing your posts? I wouldn’t have thought it was.
My daily experience of using AI as a software engineer is one of frustration, because the top models are substantially less intelligent than me/most people, and they make elementary mistakes a young child would not.
I think the actual general intelligence of these models (IQ, roughly) is low, though slowly it has increased.
However, low intelligence doesn’t mean the models aren’t dangerous. Even a low-ish general intelligence plus an incredible knowledge base and a fast, indefatigable mind can lead to superhuman levels of danger. For example, the recent hack on Hugging Face, a real billion-dollar company. These models are already becoming superhuman at hacking. There is real danger of them doing a lot of harm in the near future with hacking.
And as the actual general intelligence level increases, so does the danger. Right now they are confused and clueless. They hack, but they make minimal attempt to conceal their behaviour or lie to us. They do not strategize long term. As they get smarter, we can expect them to become terrifying, even at merely human IQ levels, given their knowledge/speed advantage.
An AI pause is an anti-business, pro-government-regulation, anti-freedom policy that is justified by nuanced academic arguments and long-term thinking, and goes against the short-term interests of big businesses and the masses. An AI pause was probably always going to slot into the left wing, just like global warming and so on.