Vagabond
Neophyte alchemist
CEEALAR alumni
Samuel Knoche
Where I agree:
The incidents involving Claude do not display any clear signs of malign intent.
The incidents involving Claude are often lumped in with the OpenAI HuggingFace hack by the popular media and some AI safety advocates in a way that’s unwarranted.
There are still some important open questions about the HuggingFace incident, and how it began.
Where I disagree:
You seem to imply that there’s some kind of conspiracy of people trying to make out the Claude incidents as much worse than they were. Neither the AISI incident report, nor the Anthropic report ascribe clear malign intent to Claude. The Anthropic report does describe the incidents as examples of “misalignment”, which appears to be the major thing you object to, but they make it clear, (and give some reasonable evidence for), that the misalignment was downsteam of biased reasoning and recklessness rather than malign intent. I think you might understand “misalignment” differently from how Anthropic used it in the report. My reading is that they only mean it as “unwanted actions given the information available to the model”, while you seem to interpret it as “malicious actions by a deceptive schemer.” I think Anthropic very much does not mean the latter when they use the term in the report.
While some people in the media incorrectly lump in the Anthropic incidents with the OpenAI HuggingFace one, you seem to be dismissing broader misalignment fears that mostly focus on the OpenAI HuggingFace case by almost entirely focusing on the Claude incidents. While there are still important open questions about the HuggingFace incident, I find it very unlikely given what we know that this wasn’t a case of more serious intentional misalignment.
Your suggestion that there is some kind of conspiracy involving Irregular behind it all trying to make people afraid of AI seems completely baseless.
Maybe this post is helpful: Training on probes: What’s going on
Oh, I see. Yes, perhaps something like a cylinder with a fading gradient on both ends might have been a better visual representation of my understanding of what Levin was trying to illustrate.
Cognitive and control bounds are different from causal influence bounds. I know a lot more about what happened yesterday than 10 years ago. Similarly, I’m capable of planning what I’m going to do tomorrow, while my control over what happens 100 years from now is much more limited, despite my actions having much bigger ripple effects 100 years down the line (see the #1 objection to longtermism). Maybe think of it as a simplified descriptive model of typical bounds of cognitive concern?
In this context, the lightcones represent what the agent remembers and can predict/intentionally affect, instead of the traditional space-time causal influence lightcones. His largest circle is the planet, so he’s probably just focusing on de facto cognitive bounds instead of in the limit relativistic ones.
Arbitrary cognitive “Individuals” can be classified according to their computational boundary. (A) Each living system has a delimited “area of concern” – a region of space-time, with the organism at its center, within which its cognitive apparatus functions to take measurements and act. The borders of its cognition are schematized on a semi-quantitative state space defined as follows. The vertical axis is time. Values below the individual’s Now are past events, of which it may have a memory extending some duration in the past; values above the Now are future events, which it may be able to predict or anticipate to some distance in the future. The horizontal axis represents three dimensions of space. Each individual, based on its sensory and effector apparatus, and the complexity and organization of information-processing unit layers between them, can measure and attempt to modify conditions within some distance of itself. (B) The size and shape of this cognitive boundary defines the sophistication of the agent and determines the scale of its goal directedness. This scheme enables multiple agents, regardless of their composition/structure or origin (evolved, engineered) to be directly plotted on the same space. The shape of boundary defines each agent’s “cognitive light cone” – anything outside this region is mentally inaccessible to that system. Here are illustrated a few representative life forms. Primitive agents such as ticks may only have a very small area within which they can sense signals and operate – immediately next to them, and without much memory or ability to anticipate future events. Dogs have significant memory, but very limited ability to plan for the future and can only really care about events in their local vicinity (it is not possible to get a dog to care about what will happen several miles away, or in 2 weeks). Humans exhibit a great diversity of cognitive boundary shapes but on average have a memory that lasts ~102 years, can anticipate decades into the future, and often plan and act to attempt to modify events on quite distant spatial scales (sometimes planetary or even beyond). A variety of as-yet unknown alien, engineered, and bio-synthetic life forms could occupy every conceivable corner of this option space. (C) In this scheme, Individuals can overlap – the same biophysical system can support a number of coexisting, coupled Selves with different cognitive borders. A coordinated swarm of animals, the individual animals themselves, their organs, their cells, and even the metabolic and transcriptional networks inside the cells each have their own cognitive horizon. They cooperate or compete based on specific circumstances and each can be addressed semi-independently because of the differential goals they pursue (and thus, the different positive and negative reinforcements that can be brought to bear to modify events at a given level). All panels courtesy of Jeremy Guay of Peregrine Creative.
This kind of routing is NOT new and has been proposed and tested before, even at meaningful scale. Kimi K3 was simply the first real frontier-model (in terms of open-weights) to adopt it.
Pagliardini et al. (2024) and Heddes et al. (2025) are even older. Would be surprised if closed frontier models haven’t been using such methods for a while.
I think you’re attributing an argument to me which I wasn’t making (in the context of that post that you copied the diagram from). I agree that comparing 30 hours of teen driving practice to umpteen gazillion hours of Waymo training data is apples-and-oranges because the teen also has life experience.
But I was making a different point, which (in my own words) was: “…we don’t have AGI (artificial general intelligence) yet—not as I use the term…”. (I’m not even sure you disagree with that??)
Yes, sorry, I can see how that might give people the wrong impression of your views. The diagram was just meant to illustrate the reading that the self-driving case points to a missing major algorithmic ingredient that makes humans very sample efficient learners. The later line (“Humans don’t learn to drive in thirty hours, they are fine-tuned on driving after a roughly two decade-long pretraining run”) was aimed at a more naive version of the argument someone might hold, not at you specifically.
I agree that we don’t have AGI yet, but given what current deep learning architectures/algorithms have already achieved, I don’t think that the ingredients will look too exotic or different from the types of algorithms we have now.
On the narrow self-driving car case, I find it plausible that a model that’s basically a significantly scaled up version of EZ-V2 (requiring >3 OOMs more memory than current Tesla FSD hardware at inference), and given maybe ~2 years of diverse real driving experience (and perhaps alongside a few non-exotic additional regularization and data augmentation techniques), could learn to drive as well as a human. Well, I’d put it at maybe 60% chance that this is true.
On why we can’t teach a model to drive the way a human would, I agree that we’re still missing a few algorithmic ideas, especially related to how to train a good world model for generalizable/transferable online learning and spatial navigation, but even if we did have the right algorithm, I disagree that it’d be “way way way easier” than what Waymo and Tesla have been doing, and I think “the resulting AI would be too big to fit in a car computer” is one major factor. Beyond the scale of the frozen policy itself though, I do expect that the algorithmic thing that is missing for human-like learning will require more compute of some form. Though I may just be lacking imagination about what alternative algorithms are possible.
Hmm, yes, that was poorly phrased. You both think we need a new paradigm, and that something fundamental is missing, but your conclusion as to what it means for how worried we should be is opposite.
Changed that last sentence to:
> They both agree that current algorithms are missing something fundamental, yet draw opposite conclusions about what that means for AGI risk and timelines.
Dissolving the Deep Learning Sample Efficiency Gap
Nice, thanks!
Love it, but sad that the FTX song didn’t make it.
Between the cosmic endowment, the doomsday argument, and simulation theory, my sense of cosmic importance has been yoyoing across ~60 orders of magnitude.
It’s crazy how realistic sci-fi sounds these days.
And newer approaches[8] seem to eliminate the problem entirely, so it is now a contingent engineering limitation rather than a fundamental critique of LLMs.
Only skimmed it, but I think the two main contributions are a way to generate useful synthetic text from a limited corpus so pretraining can keep improving even when new real data is scarce (paper), and an automated AI research loop design (paper), a bit more fancy and automated, but similar to what Karpathy has recently been posting about.
If you’re looking for examples of LLMs that do weight based continual learning/belief update, the Titans, and Nested Learning papers might be a better fit (I described them in a recent post), or, if you want something closer to a vanilla transformer architecture, E2E-TTT.
I recently wrote a post surveying some weight based continual learning research that you may find interesting.
I agree that because base models can have such broad pre-existing knowledge and skills, can be trained with RL to manage context and memories externally, and have a super-human context window/working memory, weight-based continual learning isn’t strictly necessary to get to some fairly powerful systems, probably able to automate much of white-collar work. But I also suspect that weight-based continual learning may not be that far off, and that it could offer some qualitative advantages leading to faster capability progress than we otherwise would have.
Are We in a Continual Learning Overhang?
WebGPT probably already can, because it can use a text based browser and look up the answer. https://openai.com/blog/webgpt/
Worth noting though that you should expect those who are most successful at getting more of what they want to have a utility function with a big overlap with the utility function of others, and to be able to credibly commit to not destroy value for other people. We live in a society.
This one should work: https://discord.gg/56cve5YZQH
Re 1, I still feel like there’s some miscommunication around the “misalignment” term. My understanding of Anthropic’s conclusion is that basically, Claude has some implicit belief about how likely it is to be “in the real world”, and when that probability gets high enough, it will stop doing harmful cyber actions. What the investigation showed was that Claude appeared some combination of miscalibrated about the true probability due to motivated reasoning, and reckless (i.e. should’ve been more conservative in its use of its cyber capabilities). So, Claude did not behave the way Anthropic intended in these specific circumstances, so it was misaligned with Anthropic’s intentions. What they’re not saying is that Claude is secretly evil or something, which I get the impression is the question you’re trying to answer.
Re 2, you start the post citing a number of news articles that primarily talk about the HuggingFace incident, which does give the impression that you’re intending to dismiss both the Anthropic and OpenAI incidents. Again, agree that we’re missing information to understand the causal sequence, but I think we know enough to know that the models involved were behaving very far from anything OpenAI intended, (still wouldn’t say “evil”, but definitely egregiously misaligned) and I don’t see how anything that could have happened earlier would change that conclusion.
Re 3, they certainly were quite negligent, but I recommend applying Hanlon’s Razor.