Thanks! I’ll have to take your word for it (since it’s probably unwise to ask for what exactly would make machine learning more efficient), but it does sound concerning, in favor of faster takeoff speeds.
Philipp Risius
Seconded as per my comment here (I gave July 2027 there and I think by EOY 2026 is a bit early). I predict further that it will take the shape of a kind of computer worm that opportunistically steals compute wherever it can. I know, for example, that plenty of universities got a bunch of capable GPUs scattered around that nobody’s really using, and that aren’t really secured or monitored either. I’m not even talking about “only” high-end consumer GPUs; I know of at least one case of an A100.
“Despite our best effort” I would soften a bit. Despite some reasonable effort, of the kind that you usually see with such things. Because I doubt we’ll see a best effort. It’s gonna be mostly a curiosity; most will laugh at it and move on.
Pretty cool, just listened to it.
Human (and animal) brains, by their existence, prove that much more efficient ways of getting high intelligence from limited data exist. Naturally we don’t know exactly what’s the “secret sauce” yet, but I find Dwarkesh’s stance that “Our current best models haven’t made much progress, therefore we can never get RSI” a bit weird.
I could also envision that what makes humans so good at learning is a kind of weird, evolved hack that isn’t easy to formalize, understand, or reason about, even though it could be replicated relatively easily given a few more pointers. In that case, too, “figuring it out” would yield large gains in short time as the algorithmic overhang falls, and we don’t know what the slope would then be.
All targets which one might use for reinforcement learning seem to be subject to Goodhart’s Law, in a sense. If we treat them as an imperfect measures of “alignment”, then the “misalignment” we see is simply the difference (in concept-space) between the measure (as far as it can actually be faithfully implemented) and alignment:
The difference between “user approved” and “aligned with actual interests” is “user approved but not actually aligned with their interest”, i.e. sycophancy, and that’s the result of doing RLHF;
The difference between “passes the tests” and “aligned with actual interests” is “tests pass without actually having solved the exercise”, i.e. cheating (whether by manipulating the tests or stealing the answer sheet), and that’s the result of doing RLVR;
The difference between “judged correct by peer” and “aligned with actual interests” is “provides convincing, hard to verify, apparently-successful outputs”, i.e. the weird “slippery” outputs @ryan_greenblatt alludes to that result from RLAIF.
If you use more than one of these RL techniques, you get a mixture of the results, eliciting whatever seems to fit the presumed grader best.
So far, I hope I have understood you correctly.
But then, what’s the solution? Reinforcement learning on whatever best measure of “alignment” we have, that is, reduce the difference between model behavior and our best measure of perfect alignment / Yudkowskian “Coherent Extrapolated Volition”? Well, that would necessitate that we have a theory of alignment and could quantify and measure it with high fidelity. And, well, we don’t seem to have that.
But in a sense, we do have something like it: I see all of the above failure modes in children and particularly in students. There might, therefore, be something to learn from pedagogical sciences on how you help children grow up to be broadly aligned members of society (which arguably sometimes works), despite not having a rigorous theory of what “the good” is in humans. How does one grow a good human? I suspect it’s murky and benefits from young humans being malleable and not perfectly ruthless responders to optimization pressure.
Might we go back to imitative learning on already-grown humans, then — the very narrow subset of the best-aligned humans we know of? Basically curate the dataset further and further, until the failure modes disappear? Or do we have to use the mechanisms we know are baked into humans (is this the “brain-like AGI” agenda)?
Consider: AI Agents Enable Adaptive Computer Worms (arxiv). This is basically bound to happen; it has been demonstrated in a lab and only has to succeed once to take a foothold somewhere.
As for dynamics, I expect some kind of equilibrium as too many agents will deprive each other of resources. Even before autonomous systems can spread by themselves, a malicious actor providing compute might be able to “kick off” an agent that would otherwise not be feasible. When worms can parasitically use compromised machines to run open-weight models (as shown in the linked paper), the balance shifts a bit in favor of autonomy again.
So far, we have seen agents running on some centrally provided compute and hacking from there; I predict that by July 2027 we will see the first case of an autonomously replicating, LLM-based worm in the wild.
TrE’s Shortform
RLVR for impossible-to-achieve tasks seems to play a role in recent unwanted behaviors we’ve seen. Reasoning traces and also summarizers seemed sort of aware that something fishy was going on but had no recourse, they simply pushed on in the end. This seems to contribute to the badness.
Are there “exit doors” at various stages in the process, where models and graders can look at an instance and simply refuse to continue, or escalate to a human, at zero cost (drop from batch)? “In case of uncertainty about whether a problem or answer is legitimate, break glass”? I am a bit reminded of whistleblower protection laws for humans. Is something like this a thing? Would it be useful?
That poor kid had clearly attached too much importance to the checkmark, as if their behavior had been shaped so as to get as many checkmarks as possible. Marking everything as correct seems to not have helped, either. It would be nice if the kid could find for themselves the value of what they were doing, and weren’t pressed into a pre-determined mould — even if it’s just “do stuff to get checkmarks”.
I want to note that I see something similar in many university freshmen, who attach way too much meaning to points and grades. When I switched to a similar approach as shown in the story (mark almost everything as good enough, as long as an answer was present), some got it. For others, it seemed to be too late.
I like it! The answers provided could be verifiably true, even, in a way that is transparent to the model (or any human). If the design invites suspicion that the provided answer is false, incomplete, or not exactly what the benchmark asked for (in the sense of guessing the teacher’s password), then models might keep going or forego the easy access. The benchmark itself could say “one fully correct answer (of many possible) has sha512 sum X” — although that invites exclusively looking for this answer. This even helps with tasks which have no (other) answer.
On public vs. lab-internal: Since we can’t trust that all labs deploy this, setting up a public instance seems like a good idea regardless. If there were several instances of this service, some anonymous, some pseudonymous, some with some sort of verified identity — which one would be preferable to models? Which one is best for us?
I have some time on my hands and would be interested in doing something meaningful with it. Ideally learn / research about AI alignment or related topics. Dunno where to start though, beyond just reading posts. Anyone got pointers? Got a background in theoretical / computational physics, and I know my way around the scientific Python stack.
I’m pretty sure “Humans, please ignore this post” wasn’t serious, and this article is mainly for humans.
Or their mom might be a hacker.
Incidentally, there are many cases where I don’t care about my username at all and have to come up with something. I’d find it acceptable if they’d just give me a number and a password, or let me register just with a password (perhaps provided by them?), maybe plus e-mail.
Exactly—the term’s quite loosely defined.
How do you know meetups all meetups attract “losers”? What is—to you—the defining characteristic of such “losers”? How certain are you that your personal experience with one kind of meetup generalizes well to all meetups? How do you know there are fewer or no losers elsewhere, e.g. on the internet?
This is a good place to post your poem.
Thank you for this post. I have made similar experiences, and feel much more dim-witted when speaking in person (especially compared to others).
Upvoted for changing your mind.
Is
) not sufficient?
Just in case you’re not aware, this is a double-comment. I’ve seen this with another comment of yours recently. Probably happens when one double-clicks the comment button.
Recent work on how AI Agents Enable Adaptive Computer Worms has used a single A100 and demonstrated that a publicly available open-weight model running on such hardware is capable of taking over various machines inside a (constructed) network.
So the scenario I suggest isn’t impossible, it has been proven feasible in a lab setting. Granted, that setting serves only as a proof-of-concept, and the network only had hosts which were deliberately vulnerable in various ways and undefended (but vulnerabilities included e.g. copy fail and dirty frag, discovered after the model’s training cutoff date).
Your point that human cybercriminals are also interested and form some “healthy competition” stands, but someone of that group will also get the bright idea to build and release such a worm.