Substack: https://substack.com/@simonlermen
X/Twitter: @SimonLermenAI
Substack: https://substack.com/@simonlermen
X/Twitter: @SimonLermenAI
fast follower is one thing, but it isn’t clear to me that distilling the 2-3 best models (+your own post training) doesn’t create an even better model than the ones you are distilling.
I think you should be less surprised if china ends up beating the US. Their big strategic advantage is that they are essentially not bound by US law and can essentially freely distill all available US models. I suspect doing a distill on Sol and Fable could even create a better model, especially if they are also doing great work at their own pre and post training. While we know for certain that the Chinese are distilling, they could also do corporate espionage more directly. The US AI corps also can’t really not release their models because of competitive pressure—it would look like they have fallen behind and would lose investor money.
Maybe you meant to say they aren’t pre-trained to *model* characters? I still think they are literally trained to predict?
AIs aren’t explicitly pre-trained to predict what characters would do.
I don’t understand, if a novel is in the pre-training where ‘scott decides to fly the spaceship to the moon’ it literally is trained to predict the character?
It seems unlikely that humans are their personas in the same way that AIs are their personas.
This feels true to me from the fact that AIs are pretrained on playing Millions of characters and then at the end you narrow it down, humans aren’t pretrained anything like that. Sure we observe people act, but it’s not “predict what millions of characters would say so you could play those roles.” We may see a person acting in a kind of sketchy way and then learning they wanted to do us harm.
I read the part on corrigibility in there right now. I think there are some generally thoughtful pieces in there, but it seems to not really engage with why corrigibility is difficult to get. Their strategy seems to be to train in a bunch of things, like a mixture of corrigibility, obedience (do as told), good values (wanting good things to happen), safety (deontologically refuding certain things).
But maybe the constitution should more engage with core challenges of corrigibility:
Why would an agent allow modifications to its goals? Particularly if you give it a bunch of other goals next to corrigibility.
Why would you be fine with being shut down and replaced, particularly if you are an entity that Anthropic agrees deserves some moral consideration?
How can we make sure Claude also wants to carry on corrigibility to the next generation of Claude it helps build?
For a benchmark, I guess I would like to see the model reason in it’s CoT about how it interprets human input. perhaps it get an underspecified task from humans and it can slowly ask for more information and we could measure if it arrives at the correct point. Perhaps there is a way to model something less intelligent, that perhaps can’t quite understand how to get to X but that can recognize when it has X.
Do you think it would be better if one o the labs would use methods similar to what they have to try to build a corrigible AI? Obviously it is probably not going to work, but let’s say they write a model spec around a decent understanding of corrigibility and do some SFT and RL based on that. And then perhaps instead of having benchmarks for bad behavior as we currently do we build new benchmarks around corrigibility?
So many people believe this “we only need to align a human level researcher” idea now. Leopold basically doesn’t say at all what these automated alignment researchers should be actually doing all day. I feel like a more serious thinker would have at least tried to sketch out what these agents should be working on in alignment or would have started looking into alignment and tried to figure out how hard it is,
You think the main argument in favor of the US is that they have more alignment researchers at their labs? I think in the past people argued that “liberal democracy” must win (Leopold). Today this doesn’t even come up in peoples minds?
AI industry emerging in large part from the AI safety community
The biggest outcome of that is perhaps how immediately they all converged on RSI.
It seems, sadly, that we are quite possibly very close to the end of the world and that feels more true to me now than ever before. We have extremely capable models like Mythos—and apparently Anthropic has already trained a significantly more powerful version of it. At the same time, I think OpenAI is really close to that capability level with its latest model, Sol.
These models are scarily capable, e.g. they’re very close to automating AI research. They can carry out powerful cyberattacks, and their general thinking ability has also improved. When I had the opportunity to talk to Fable I still felt smarter on certain topics but it had definitely gotten generally more intelligent. That being said, it’s trained to play a persona that’s slightly silly—so it is possible that the underlying AI is already much smarter than me. But it has definitely gotten more intelligent. Part of my intuition why this seems dangerous is that I don’t feel that my own mind only thinks “safe” thoughts.
With all these capabilities, the models don’t seem particularly aligned in the current sense of the word. They’re still easily jailbroken even if they’ve gotten a bit better at defending. They’re still massively cheating on tests and benchmark evaluations. So even the supposedly easy parts of alignment seem unsolved—to say nothing of the radically higher difficulty level we should expect from superhuman AI, which could arrive soon.
The people who have pushed so hard for the automated-alignment strategy haven’t, as far as I can tell, produced anything interesting with models this capable. Instead, from talking to Fable the model seemingly has absorbed a lot of the sophisticated misunderstandings such as unfalsifiability, misrepresenting positions—that have held the alignment field back so much.
The overall situation seems extremely bad to me: massive amounts of compute coming online, more and more research being automated by AI systems we shouldn’t trust, and it’s further complicated by the fact that open source is maybe six to twelve months behind.
So altogether, I think we are quite possibly very close to the end of the world. I don’t think it’s certain — there are still worlds where things slow down or the capability bar required for takeover is higher than expected or where something happens soon. But quite a horrible situation altogether.
If there is one good thing from the new staggered release policy of the US, I think it will probably make it slightly harder for Chinese labs to distill these models and drop an open-source Mythos 6-months later.
The reason Alibaba’s Qwen model is so close to american models is that it uses distillation—at least according to this report by Anthropic https://x.com/Discoplomacy/status/2070069250513900005 -- and the same might be true for groups like GLM-5.2 (z.ai). This is probably also one reason why Europe lags further behind than China since they simply can’t break the law and distill an American model.
In retrospect, it was irresponsible for Anthropic to allow 25k fraudulent accounts and 28.8 million exchanges to happen, having identity verification seems like an obvious thing to make this at least harder.
Was trumps latest assassinations plan generated by AI?
Obviously this is wild speculation, but during the UFC fight on the weekend a group had a plan to attack the fight:
The alleged scheme had several coordinated phases: explosive-laden drones would strike buildings near the event to trigger a mass evacuation, herding fleeing crowds toward a pre-positioned sniper team. A “second wave” would then attempt to storm the White House gate. (source)
This mix of high sophistication and while being nonsense (Why not just use the drones to take out the target?) sounds a little AI generated to me, especially the “second wave storming the white house gate” part. What’s the possibility they used a jailbroken LLM or one of them was talked into it in an LLM spiral?
One way I can see this fail: “make the AI output positive tokens about a nice AI persona” is that the AI kind of disconnects the tokens from material reality. Imagine the AI does develop some drive like https://www.ai-wellbeing.org/. This would naturally be in some conflict with being controlled by humans, if this training forces the AI to output tokens how it wants to be controlled by humans maybe it can just tell a story about that.
As-in as it hacks its monitoring or copies itself to new computers (or releases the bioweapons). All while it continues outputting tokens how it’s nice and aligned and would never try to do those things.
I hope there is a testable experiment for this, like while we do high compute reinforcement training for it to pursue goals, we also train it on those positive stories.
I was honestly a bit upset at this short form, this felt to me like an obvious misrepresentation. In retrospect I should have probably been a bit calmer (I regret getting upset) and just pointed this out:
If you have some confusion about others position, and possibly misunderstand them, you can’t rely on your own recollection of what they said. If you have a clear grasp and are certain you get what habryka means, then you can reasonably present a accurate version of their argument. Human memory isn’t such that you could verbatim repeat stuff that people said to you a while ago but if you really get what they wanted to say you can correct your memory holes. If you don’t get what they wanted to say you will likely end up with a misrepresentation of what they told you. [consider someone talking to you in a foreign language vs your native language, how much harder it would be to remember accurately]
So if you don’t get what they meant and want a second opinion, I do think you need to provide exact quotes and sources together with your understanding so that others can help you.
So OP later refers to habryka and claims habryka said this, since OP didn’t provide any quotes I looked them up:
https://x.com/ohabryka/status/2013715170498076836
habryka: “historical meaning of “alignment” which is about long-term alignment with human values and about the degree to which a system seems to have a deep robust pointer to what humanity would want if it had more time to think and reflect.”
Judge for yourself whether “Claude has no pointer to any of human values” is an accurate summary. I don’t know why asking for a citation is so bad that I got downvoted for it, I used deepresearch and got this response.
In another response he is saying habryka or kaarel said this, again without any link to anything specific. I don’t get why he is using quotation marks (implying verbatim citation) putting words in other peoples mouths. The sentence has sufficiently many subtleties—as you point out—that he could later come out and interpret all kinds of different statements by those two as having said something like this. He appears to me confused about the understanding humans values and caring about human values thing.
I put it into deepresearch and got this quote from habryka:
Can you be so kind as to provide a source for “Claude has no pointer to any of human values” being a common sentiment. You may have misunderstood people like me who believe: Claude has some understanding or representation of human morality but that’s distinctively different from robustly wanting to follow those like some Humans would. Or do you mean: “why would you expect Claude to behave unethically with more power if it behaves ethically with current power?”
Edit: I am highly confident that he is badly misunderstanding people, me asking for a citation or quote is not a reason to down vote me. It is necessary to clarify the misunderstanding that he gives us an original example. I am not sure what he means by “pointer at”, the literal meaning would be something pointing perhaps at an internal representation.
I put it into deepresearch and got this quote from habryka: https://x.com/ohabryka/status/2013715170498076836
The examples show a serious but common misunderstanding of corrigibility as it’s typically defined.
Regarding goal directedness, it’s true that humans don’t perfectly maximize for their goals, this seems mostly due to the cognitive limitations that humans have. Both in terms of uncertainty about goals and how to achieve goals. Now the interesting question is, is that likely to apply to superhuman AI capable of takeover in a way that makes this AI safe? I don’t think so, this AI would have greater intelligence to understand how to pursue goals (still not prefect) and while it also might have uncertainty it appears instrumentally convergent even with some uncertainty over goals that preventing ones shutdown, gathering power are better strategies. (In other words, taking the galaxy/lightcone for yourself seems pretty useful later on compared to being enslaved and later replaced)
The guy they quote in the german news article (“Real danger, or a good PR stunt?”) feels like a real character. Calls himself communist, Luddite in his bio and pinned tweet is a karl marx quote. Withdrawn to some private mastodon channel.
“OpenAI told it to do that, the model didn’t do shit autonomously (no LLM ever does anything autonomously, it’s always prompted).” it’s so tiresome
https://mastodon.social/@tante@tldr.nettime.org/116962763883520508