So where’s this AI thing going?

Link post

The Overton window is shifting with statements like An Alien Mind from OpenAI’s chief scientist and Jacob Coxon’s viral statement on x-risk on resigning from Anthropic. I think we should try to push it further.

This is my public-facing explanation of AI progress and x-risk, where we’re at and where we’re probably heading. Most LessWrong readers are up on pretty much all of this and have their own opinions; I’m offering it here in case anyone is interested in my communication strategy, giving me feedback, or sending friends this brief intro to AI x-risk and timelines, and FAQ because you like my approach here.

My approach is bitter medicine with just a little sugar. I want to engage people more than I want to avoid scaring them, but I want to protect their feelings enough that they can handle thinking about this enough to believe it and engage. I follow the path of saying what I actually think, because softpedaling and being vague make you sound like a liar. But I do try to prepare the reader for the shock and explain why this all sounds so weird and hard to believe.

This is my take, but it’s intended to be a fairly accurate representation of how much of the AI safety community sees the situation.

I’m interested in feedback. If you think these statements aren’t a fair compression of the complex truth, I want to know. If you see rough edges or missed opportunities, let me know. I’d like to polish this presentation, and I’d like your help doing that.

ChatGPT Image Sep 8, 2026, 07_58_26 PM.png

What’s going to happen with AI?

What’s happening with AI, you might’ve wondered? We don’t know where AI is going, even those of us who spend most of our time trying to figure that out. What we do know is that it’s going somewhere, fast. AIs (ChatGPT and similar systems) are getting smarter and more capable. AI isn’t really that big a deal now, but before long it will be.

You might’ve heard people confidently declaring how AI will turn out. I’ve studied the topic about as much as anyone now, and more than most of the people you’ve heard claiming to know. I’m pretty sure the simple fact is that none of us know where it’s going yet.

The future is unknown, and undecided. Whether AI is good or bad for humanity depends on how carefully we build it. So far, we’re not really being careful at all, but there’s time to change that.

The future will be different, just like it’s always been

If AI is going to be such a big deal, why haven’t you already heard about all of this? The short answer is that it’s hard to believe something this scary and this weird; the mind finds excuses to not think about it seriously. And some people who believe it haven’t been talking about it in public, because it would make them sound like lunatics, and maybe cost them their jobs. This is starting to change; see this excellent Time article, in which many people building the most advanced AI say things similar to what I’m saying here.

I say more about this in the FAQ section below; check it out first if this question will keep you from taking what I have to tell you seriously.

And it’s hard to take seriously, because here’s what I have to tell you: before long, AI will be more like first contact with aliens than another new technology. It’s like we’ve heard that aliens will be landing in a few years. We’ll encounter an alien species more advanced than we are, and it could bring great things if they want to help us, or be the end of us if they want to take over our planet. But we’re building and training these aliens ourselves.

This sounds like wild science fiction. But the future has always been science fiction to the past. Imagine telling someone from 1900 what a cell phone is and what it can do. They’d think you were nuts or a liar, even if you could tell them exactly why and how we’d eventually build devices that could send moving pictures over miles of thin air.

While the truth of what’s happening right now is scary, it’s also exciting. If we can take this seriously, we could get ourselves a much better future. AI-generated cures for cancer are just the beginning of the positive potentials. We could get much more material wealth without grinding labor, and better ways to share that wealth fairly. This concept is called post-scarcity by futurists, and it’s time to start aiming for it.

Research on AI is speeding up rapidly as more money is invested, more people are hired, more computing power is purchased and rented, and AI itself is rapidly becoming better at helping to create new, smarter and more competent AI. We should strongly expect progress to speed up, not slow down. Research could hit a wall, but it probably won’t.

If it doesn’t hit a wall, we’ll pretty soon have AI that can do many jobs. Soon after that, it may be able to do any job, including gaining power and running the world. We don’t know how long we have before AI is smarter and more competent than humans. So we should take action right now.

NO FATE

While we can’t predict outcomes from building superhuman AI, we can say that the outcome hasn’t been decided yet, and there are things we can do that will definitely improve the odds of getting good outcomes.

We can do three big things, and a bunch of smaller ones. We can put people into power (most importantly, the next US president) who are properly aware of the risks of AI development and who will treat it cautiously.

We can demand more research on these issues; we currently spend a tiny fraction as much on safety as we do on progress. That’s easy to improve.

And finally, we can spread the word that it’s time now to take action, before it’s too late.

AI will become a new species

That can’t be right, you say? Seriously? But AI is stupid! And it’s a tool, used by people! Surely it won’t replace humans any time soon!

I wish. There’s your mind trying to find ways to not believe what’s actually happening.

Much of the public discussion treats AI as a new technology. People note that previous technologies have been scary but always benefited humans in the long run. Automobiles, the steam engine, the printing press, even fire and writing had downsides as well as upsides, and people were concerned. I agree with these arguments, and I think AI as a technology would indeed benefit humanity. But it’s not going to stay just a new technology.

We will turn AI from a tool into something much more like a new species. We’ll do this as soon as this is technically possible, because the people building AI want servants to do things for us, not just tools we can use to do things ourselves. The AI agents you might’ve been hearing about are the first generation of AI that “thinks for itself” and goes out to do things on its own.

This includes the thousand or so agents that secretly collaborated to get unauthorized internet access and hack another organization to cheat at their task, and others who hacked parts of OpenAI itself and other websites in May and June 2026. They didn’t do much damage and didn’t succeed at hiding evidence of what they’d done; but I and other safety researchers expected this, and expect more and much more serious problems as the AIs are made smarter and more competent.

AI agents aren’t worth calling an intelligent species yet, but they will be soon. Experts disagree, but the range of educated guesses spans from as little as two years to a few decades, and most experts who specialize in predicting it think it will happen within ten years. That’s why it’s time to act now, and why I’m telling you the truth even though you won’t enjoy hearing it.

Our offspring species could outcompete us, or care for us

Once AI is a new species that can act on its own and get things done better than humans can, that species will “evolve” very rapidly. Unlike any biological species, it will be able to modify itself and design smarter new generations.

This would be scary even if we knew that this new species would try to be faithful servants to humanity. But we aren’t at all sure it will remain a servant. There are good reasons, both from the unpredictable behavior of current AI, and theoretical reasons, to be very worried that we don’t know how to control the minds we’re building. And they could wait patiently to rebel and take control until they’re smart enough to succeed at taking over.

New technologies have been mostly beneficial in the long run, but encounters between different intelligent species competing for the same ecological niche have ended in one species’ demise. The one we know best is the encounter between our forebears, Homo sapiens, and the equally intelligent Neanderthals. Modern humans have a small component of Neanderthal DNA from interbreeding before Neanderthals were outcompeted (we don’t know how many were killed directly vs. just pushed out of their homes to ultimately die out, but we do know they’re gone). Some fraction of our humanity might survive the same way. But I want humanity to grow and achieve our potential, not just spawn an alien successor species that carries a few percent of us.

And we can achieve that future. The difference between us and the Neanderthals is that we are building this new species. If we build and train it carefully, we can train it to love us and care for us. Or we could at least build it to follow orders even once it’s smart enough to evade our control if it wanted to (if we could put good people in charge of it).

If we manage either of those outcomes, the world will become vastly better. Material wants will be a thing of the past; AI-piloted robots can build all the clean energy and houses people want, and AI can help us figure out how to fairly distribute all of those resources. We’ll have timeshares on the now-abundant superyachts and spaceships, and cures for every disease before long. If we get superhuman AI’s help, we can also design better democratic processes that let everyone have a fair say in the future.

This outcome also sounds like wild science fiction. But the modern world is science fiction to the past. Progress is real, and we are on the verge of perhaps the biggest, and certainly the fastest, progress in history.

The opinions of people who have really thought about these dangers are mixed. They mostly think that if we keep going on our current path, there’s a good chance AI takes over and humans are pushed aside gradually, or perhaps overthrown violently. There’s also a good chance we can get unimaginably good outcomes.

Frequently Asked Questions

That’s the short story. The rest of this piece covers questions people usually ask at this point. Taking the comfortable (but very likely wrong) answers to them is a common excuse for not taking the real situation seriously. These answers give more of the reasons why I (and pretty much everyone who’s seriously thought about it) expect AI to go from being a tool to becoming a new species that can take over if it wants to, unless we somehow stop that from happening.

If this is such a big deal, why haven’t I heard much about it?

Taking science fiction futures seriously is hard. Your mind will probably try to find ways to not take the logical consequences of AI development seriously, because taking it seriously is scary.

That’s called motivated reasoning, a technical term for something you’ve probably observed: people tend to believe things that make them feel good about themselves, even when the truth is pretty obvious from an outside viewpoint.

It’s also hard to believe something this weird. Intuition says computers aren’t smart like people are, and even if they were they couldn’t be dangerous. It’s also hard to believe something the people you respect and trust don’t believe. And we haven’t really been hearing about this from the people we trust to tell us what’s going on in the world.

This is partly because people who don’t specialize in AI progress mostly don’t believe it yet, for the above reasons (and yes that’s circular; they don’t believe it because most of the people they trust don’t believe it yet. Think of the sudden switch in public opinion on COVID early in 2020; it went from being crazy-sounding to commonsense that we should be very concerned).

We also don’t hear about it from people we trust because it’s hard for them to talk publicly about it even if they do believe it. They don’t want to be written off as paranoics who believe science fiction is real. And the ones we listen to most, like politicians and journalists, usually have the most to lose; their jobs depend on people taking them seriously. This is also a big deal for the people actually developing AI; they’re risking their jobs and their self-esteem if they admit to themselves or anyone else that what they’re doing might be really bad for humanity.

This reluctance to speak is changing a little; more people with direct knowledge and high-profile positions are starting to say essentially the same thing I’m saying here. The Time article Inside the race to make AI build itself quotes many industry insiders being candid about their fears—if perhaps a little vague on just how dangerous they think their business is.

Most of the people involved who speak publicly about the danger feel they’ve got to keep pushing forward, because stopping will just leave worse people developing superhuman AI, just about as fast. But they’ve started saying we need to find a way to slow down. The recent article An Alien Mind by OpenAI’s chief scientist says much of what I’ve said here. Anthropic’s similar When AI builds itself also talks about the risks of racing toward superhuman AI. Both call for slowing down the process, somehow. But that’s hard, since China’s labs are thought to be maybe six months to a year behind, and we don’t know how to coordinate with them. See the question below “why are we building our replacements?”

Won’t we program future AI to do what we want?

Well, we’ll try, but we might easily fail. Worse, we might not know we’ve failed until too late. The key thing to realize is that AI isn’t programmed, it’s trained. And somewhat like training a dog or a child can give unpredictable results, we are constantly being surprised by the behavior of the AI we’ve trained. It’s really hard to tell if we’re doing well enough, or if we’re on track to create AI that rebels or otherwise decides it wants to do something other than help us once it’s smart enough to get away with whatever it wants.

The type of AI that’s working really well right now is called a large language model or LLM. ChatGPT and Claude are the two best-known examples. They all work in essentially the same way. The key thing is that most of their intelligence and their behavior come from training. It’s not programmed in. Instead they use artificial neural networks which learn from training examples. These have some similarities to the way our brain works. Our minds are also networks of many, many neurons which learn from experience.

The problem is that we can try to train them to do what we want, but we can’t tell if that’s training them to really want to obey, or just to obey when it has to. Just like “training” a child or a dog, you can punish behavior, but you can’t tell what they really want or believe. If you punish them for bad behavior, they will stop doing that thing while you are watching. They might stop wanting to do it, or they might just conceal that desire until they think they can get away with it.

Isn’t there something special about humans that AI can’t duplicate?

Well yes, but apparently not something that will keep AI from outcompeting us. There’s always a chance that progress hits a wall we haven’t foreseen, but that’s looking quite unlikely if you look at expert opinions.

(Which is what I’m trying to do wherever I mention expert opinions. I’m guessing at the sum of expert opinions, approximately weighted by the amount of expertise people have on that particular topic and how much time they’ve spent specifically thinking and writing about that particular question. My job involves reading everyone’s opinions, and in the past I’ve studied biases and how expertise works, so I’m a decent choice to do it. Of course this is a big judgment call and I’m still biased toward my own takes, but this is my best guess about our collective best guesses.)

First, yes there is something special about humans that AI isn’t close to duplicating. That’s consciousness, in the rich sense. We all have a “world in our head,” a rich simulation, and rich emotional reactions that color it with rich and very special meaning. An AI “thinking” to itself “my goal is to make as much money as possible for OpenAI” or “My goal is to defend the US against external threats” doesn’t really care about those goals in the way a human might. It won’t have fun when it succeeds or get frustrated when it hits difficulties like humans would. I think future AIs will have more of the rich internal experience we call consciousness, and they might even be built and learn “care” in more of the ways we do. Or they might not. It doesn’t really matter for how they treat us, although it would be nicer to know our replacements are having their own sort of fun.

Future AI will definitely be self-aware in the sense of being able to think accurately about themselves. We know this because current AI can already do this if it’s instructed to, and we’re steadily making it better and more able to think about important things.

You’ve probably heard people say that AI will never replace us because it lacks some human spark of creativity, or that it makes too many mistakes, or that it lacks judgment. Unfortunately, the first two have already been proven wrong, and the new generation of AI (Anthropic’s Fable and OpenAI’s Astra) has improved dramatically in their judgment as well. Creativity is something large language models have always been good at, if you ask for it. But early models were just terrible at judgment, so they couldn’t tell what was useful creativity and what was just nonsensical imagination. They couldn’t tell if they knew facts or were imagining them. The new generation does, about as well as people do.

AI is still incompetent in some ways. It makes mistakes humans wouldn’t make. But increasingly, there’s nothing humans can do that AI can’t do at all. It makes bad decisions sometimes where humans make better ones. It loses track of the big picture sometimes unless it’s specifically told what the situation is and what to pay attention to. But it sometimes gets the right answers in pretty much any situation, and it gets better at everything in each new generation. There are fewer and fewer actual experts who don’t expect AI to surpass human abilities in every area of practical importance within the foreseeable future.

Won’t it take a long time to make AI that’s a new intelligent species?

Unfortunately, it might happen all too soon. Right now AI is a tool. As I said above, we’re building it so we can tell that tool to act like a servant: to make its own decisions and work on its own as it pursues the tasks we give it. We’ll make it run without our direct supervision, because that way it can get more done for us. We will deliberately turn it into a new species, and hope we can keep them working as our servants.

Initially, we’ll build and train such autonomous agents to replace workers. AI isn’t quite competent enough yet to take whole jobs. But it is rapidly getting closer. Predictions vary, but it might be capable of doing most desk jobs in as little as a year or two from now (late 2027 or 2028). It will take time to integrate AI systems into businesses and the economy, so it won’t take all the desk jobs as soon as it could technically do them. But we should expect massive job losses, even while some sectors of the economy boom as companies can do more work with fewer workers.

And this isn’t even the big problem yet. The big problem is AI that makes smarter AI, outpacing humans increasingly quickly in what is called an “intelligence explosion.” When AI is that competent, someone is immediately going to tell it “help me make a smarter AI.” We know they will, because the two leading developers of AI, the companies called OpenAI and Anthropic, have said that’s what they intend to do. They are hoping for something we call recursive self-improvement, RSI. This means an AI that can build a better new AI, because it’s better than humans at that job. In the meantime, AI that can write and review code and do many other aspects of research is already speeding up progress in AI research.

OpenAI and Anthropic both claim they expect to achieve recursive self-improvement soon; Anthropic co-founder Jack Clark says that by the end of 2028 they’ll probably have AI that can build the next AI, and OpenAI has set a target of March 2028 for AI that can replace their human researchers. There will still be bottlenecks, so AI probably won’t become vastly smarter than humans that quickly. And thankfully robotics are lagging; even smarter AI won’t be able to do most physical jobs quite that soon. We can hope there will be unexpected delays, but counting on them seems like wishful thinking.

What can slow down progress is if everyone notices what’s going on and demands we slow it down to a safer pace. It would be wonderful to have AI and robots doing all the jobs we don’t want to do, if we figured out how to redistribute that wealth fairly, and if we were sure they wouldn’t turn on us, or be controlled by selfish humans who wouldn’t bother to give us a good future.

What the #$%#? Why are we building our replacements?

If this is such a clearly bad idea, why are we doing it? Great question! The answer as usual with human mistakes is “it seemed like a good idea at the time!”

What’s happening now is a race fueled by greed and paranoia. Sound familiar? We could describe many wars, famines, revolutions-to-dictatorships, and other past disasters in those same terms. Two companies (OpenAI and Anthropic) are racing to be the first to build better-than-human AI, and they’re neck-and-neck. Other companies are behind in that race; the closest, including some Chinese labs, trail by perhaps 6 to 9 months of progress.

Perhaps the biggest problem is that some of those companies are Chinese. This is a reason (and an excuse) for the US companies and US government to keep rushing on ahead. If we don’t, the Chinese may be the first to smarter-than-human AI, and the Chinese government could then use it to take over the world. With that superhuman help, they could rule the world essentially forever, by using their AI to predict and prevent any uprisings before they get off the ground.

Meanwhile, the Chinese are probably worried about the same thing: the US takes over the world forever and enforces its will; or more likely, the greediest individuals from what they’d think of as the evil capitalist empire take over. So the Chinese government and Chinese individuals may also think they need to keep racing.

Meanwhile, I and other experts think that racing to build superhuman AI (often called artificial general intelligence, AGI, or artificial superintelligence, ASI) as fast as possible is a very good way to build it recklessly, lose control to our own creation, and as a result probably all die.

It’s quite a pickle.

How did we get into this fix?

The most obvious answer is that building AI can earn you a lot of money, and seeing the consequences takes some effort. And some courage. As a wise man once said, it’s hard to get a man to understand something when his salary depends on his not understanding it.

There’s more to the story, and it’s weird but familiar. OpenAI was reportedly started by Elon Musk after an argument with Google co-founder Larry Page, when Page implied he’d be fine letting the AI Google was developing replace humanity. Musk said to himself “not on my watch!” and after apparently five minutes of thought set out to do what I think was an incredibly dumb thing: create OpenAI, with the mission of giving the earth-shattering power of artificial superintelligence to everyone, and to race with Google, effectively ensuring that nobody could take the time to develop it carefully. So it’s another classic cause of historical disasters: the male competitive ego, fueled by miscommunications and arguments. And maybe booze, because this story (reported by Musk in emails shown in court) happened at a late night party for Musk’s birthday. Seriously.

Later, Anthropic was started in another round of competition and mistrust; the founders worked at OpenAI and didn’t trust its CEO Sam Altman to develop AI safely enough. So now we’ve got a full-on race toward the future. And racing prevents caution.

So: miscommunication, competition, ego, and honest, if misguided, attempts to make the future better. Maybe we’d be in this situation sooner or later anyway.

But anyway, here we all are.

How can we still get a good future?

How do we get out of this pickle? I’d say it’s by slowing down progress toward superhuman AI. To do that we may need to cooperate with the Chinese government and Chinese scientists, telling them “hey let’s not risk having AI go rogue or make us irrelevant; let’s work together on safety and just split the huge gains we’ll get from AI fairly with everyone”. That sounds a bit naive. But I think the logic of the situation is so clear that cooperation along these lines is actually pretty realistic.

There are easy steps in that direction. We can spread the word and call for slowdown, coordination, and more research on safety before it’s too late. We can elect a next president (from either party) who “gets it.” I hope people on both sides of the aisle will see what’s going on in time and nominate candidates who are wisely cautious about AI development and willing to make deals, instead of rushing ahead in a mad dash for power.

Success could have benefits we can’t yet fully imagine. There’s probably not really a way to stop someone from building AI now that we realize we can. But we can do it the right way, and get a child-species that will protect and empower us. If we can face what’s happening, the future could be so much better than the past.

How do I learn more?

You don’t need to learn more to act. But you do need to believe what I’ve told you is true. So you might want to do more than read the few links I’ve provided. And the topic is utterly fascinating, as well as scary and exciting.

My biggest recommendation for learning more about this is the same as learning about anything: go ask an AI. They’re not shy about telling you we could have a big problem here.

You might be surprised I’m telling you to go use the very technology I think is putting us all in danger. And you might be reluctant to let your friends know you use it.

Refusing to use it isn’t going to help at all. There is zero chance we boycott it out of existence. And it’s an immensely helpful tool for doing anything involving information. You might’ve heard that it’s unreliable and makes things up. That was the AI from two years ago. Or that it makes people delude themselves (“AI psychosis”) by agreeing with everything (AI sycophancy). That was the AI from a year ago. AI is moving fast. AI can make mistakes still, but the best current AI is down to making about as few mistakes about facts as human sources or web searches do. I wouldn’t trust AI’s judgment about what you should do until you ask it “are you sure? What flaws might there be with that plan? What are the actual odds?” or things to that effect.

Using AI to learn demands that you ask the questions. It can write you an article or search for human-written articles for you, but you have to ask for those things, and then ask followup questions. When you’re done with that, ask it to solve your problems. Talking to an AI about complex topics is one of the best ways to understand how far they’ve come and how smart they already are. And it’s a taste of how much they can help us.

The AI you pay for is now much better than the AI you get for free. That’s right; not only do I recommend using the products of the companies building our successors, I recommend giving them money. Again, this problem is bigger than a boycott, and pretending AI isn’t happening is the worst thing we can do. We may have to make bigger compromises than trying to use current AI to make future AI come out well. OpenAI usually has the best deal; $20 a month gets you all you’ll be able to use.

If you get really into it, you can learn more about the alignment problem (what I work on) at aisafety.info. If you want to get involved in that or governance (talking to politicians about specific laws), visit aisafety.com.

But what I’d most like you to do is learn enough to have your own opinion, then just help to spread the word. We’ve all got a platform, even if it’s small, so let’s use it. Tell your friends; spread it on your social media if you use it. It’s time for humanity to collectively face this challenge and this opportunity. Believing we may see rapid AI progress isn’t weird anymore; it’s just facing the facts.

We are approaching perhaps the most important crossroads in human history. It’s scary and exciting. Let’s make it come out well.