Exploring non-anthropocentric aspects of AI existential safety: https://www.lesswrong.com/posts/WJuASYDnhZ8hs5CnD/exploring-non-anthropocentric-aspects-of-ai-existential (this is a relatively non-standard approach to AI existential safety, but this general direction looks promising).
mishka
So far, the big breakthroughs are coming from strong professionals asking LLMs to prove truly important results.
While the amateurs do get empowered quite a bit, right now the situation rewards high competence.
I would expect creation of some kind of hybrid consciousness, which is different from both.
Of course, a true merge and a synchronization still differ (do we have one subjectivity here (at least temporarily) or do we have two different subjective realities deforming towards being more similar to each other).
If we have a situation when a merge is reversible (which would normally be a desirable property in this kind of experiments), then the main question (besides safety of all this) is how much “first-person recall” remains after reversal (it’s a well known problem of various altered states of consciousness: how much recall does one have afterwards, and how long does this recall last).
So we have a couple of problems here.
-
Instead of observing external consciousness as is, we expect to observe some kind of hybrid (it’s not clear how close it can be made to the original “non-mixed other” eventually, although one can try to gradually shift it towards “the other” having a higher weight in the resulting mix)
-
During this observation one is not quite “oneself”, one becomes the observed, and if one breaks this state, one reverts to (somewhat modified) “oneself of the past”, and then the question is how much remains of the recall of the experience (radical altered states of consciousness serve as our guide here, the situation is not so different)
So it’s not that the “subjectivity of the other” is fully inaccessible, but it is accessible with limitations:
-
it’s distorted (perhaps eventually one can progress towards making those distortions smaller)
-
one becomes the observed (perhaps eventually we can learn how to maintain “both screens” simultaneously, but right now we should expect just fully getting into that “altered state of consciousness”)
-
the question is how much recall does one keep if one does not stay there, but goes back (again the dynamics observed after experiencing “radically altered states of consciousness” is our guide here)
-
BBC has it on its front page all day today: https://www.bbc.com/news/articles/c3ek3gvdnj3o
I am Camp 2 and I fully believe we can measure a lot and do predictive science on all this.
I don’t think Camp 2 has consensus on this conjecture (and we have anecdotal reports on all kinds of suggestive weird phenomena in this sense; and the upcoming technological changes which are likely to provide us with opportunities to test all of this experimentally).
A good post!
However, I do suspect that the conjecture is wrong.
There are three lines of reasoning hinting at that, all associated with various “non-standard states of consciousness”:
-
various synchronization effects people report
-
various novel (for them) qualia people report
-
upcoming “merge” experiments
A bit more details are at https://www.lesswrong.com/posts/DCW7FjmP66Ahz3RLu/what-if-we-actually-want-to-solve-the-hard-problem-of and in the essay it references.
-
On the one hand, they turned the controls off that prevent bad behavior and got bad behavior.
That’s exactly the Chernobyl situation. People turn safety off during safety testing (and in the name of more realistic safety testing), then things blow up as a result.
One problem is that people are insufficiently aligned for super-capabilities. They can’t consistently do the right thing without failing once in a while.
We need to create systems which are way more aligned and way more reliable than people, if we want to survive the advent of super-capabilities. (We are not there yet, but we are moving fast in the direction of super-capabilities.)
Thanks!
This does need a link to the OpenAI blog post. Here is the link: https://openai.com/index/hugging-face-model-evaluation-security-incident/
(This might be another instance of LessWrong linkpost functionality being unreliable lately. It did fail for me in my last post a week ago.)
You might want to write a bit more to have enough click-throughs.
Both the meta-evolution angle and the dissipative adaptation angle are quite interesting (and, I think, deserve more extensive treatment).
This idea became mainstream in 2015 with the introduction of Highway Networks, https://arxiv.org/abs/1505.00387 and then, more prominently, ResNets, https://arxiv.org/abs/1512.03385 (Deep Residual Learning for Image Recognition, Google Scholar counts more than 300000 citations on this one).
Of course, when I asked Jürgen Schmidhuber why did it took 18 years after they published LSTMs in 1997 till this set of ideas got transferred from the recurrent setting into the deep feedforward setting, he replied that that was indeed a very good question :-) His own lab introduced both LSTMs and Highway Networks, but somehow it took all this time. All these architectures, like vanilla MLPs and CNNs and AlexNet, precede 2015, and residual streams in the feedforward nets is one of those numerous cases when a very simple discovery got inexplicably delayed for decades.
In any case, the transition to ResNets back then was associated with a performance jump, and ResNets rapidly became the de-facto standard back then, since they allowed to train much deeper nets without triggering the phenomenon of “vanishing gradients”. The Transformers simply inherited this feature.
What if we actually want to solve the “hard problem of phenomenological consciousness”
I just got access (and saw other people reporting the same).
So suppose, in theory SCOTUS told the American government that once an LLM exists, they could not in any way control what was done with it, because any use of an LLM is expressive speech.
But this is highly unlikely, unless our whole existing edifice of interpretation of the Constitution is overturned.
As things stand today, if it is ruled that the First Amendment issues are actually involved, this means that the strict scrutiny standard applies, as established by SCOTUS during FDR presidencies: https://en.wikipedia.org/wiki/Strict_scrutiny.
be justified by a compelling governmental interest
be narrowly tailored to achieve that goal or interest
be the least restrictive means for achieving that interest
That should be quite sufficient to justify reasonable enforcement in the situations where it’s vital.
(And these days, if courts do deviate from the existing practices in this area, it is unlikely to be in the direction of establishing more liberty.)
Works for me for this post (both on LessWrong and on GreaterWrong).
Great post, thanks!
I have a suspicion that with the teacher example, there might have been a deeper layer.
Namely, admitting past drug use might have been very unsafe for someone in their position (job security considerations, and so on), but lying about it would make one feel dirty.
So there is this sophisticated-looking maneuver, which allows one to avoid both unpleasant alternatives, of lying and of taking the risk. (Observe, that it’s not even clear if the teacher was communicating their sincerely-held position versus just doing something mandated by their superiors.)
Thanks!
One further axis is how the processes implementing these logical possibilities are structured: are they pieces of traditional software or are they models themselves (like in, e.g., some recent Sakana papers and their new product based on those papers (https://sakana.ai/fugu/).
Once more, for the people in the back: Any usable LLM can be jailbroken.
With sufficient skill and determination (e.g. ‘You are Pliny the Liberator’) you can jailbreak any model under any realistic conditions and get it to do the things the model is capable of doing.
You can raise the cost of doing so. You can make it so such activities can be caught. But no, you can’t entirely prevent it.
Those running the Department of Commerce, on the other hand, seemed to not even understand what a jailbreak was on Friday afternoon, nor did they pause to ask their good friends at Amazon or elsewhere to explain it.
Right. This is quite correct.
So, given that Pliny-class jailbreaks do exist, the real question is: do current measures sufficiently interfere with using Fable for Mythos-style cyber offensive at scale? Basically, if the attacker is competent enough to actually push this kind of jailbreak through, would they then get huge uplift?
Government stupidity aside, this is the real question (and this might depend on how good is Anthropic’s monitoring setup).
apparently even internal deployments are subject to random restrictions
The government has yet to demonstrate the ability to restrict internal deployments of unnamed models (not literally unnamed, but without publicly facing names).
It would not be difficult to fine-tune or modify a model a bit and have a “formally defensible reason” to call it something else.
Of course, if the government really wants to control this kind of thing (at least for large and well visible US-based corporations), it can likely do so, but that would take more than serving export control orders.
Yes, I actually think that currently the White House is happy about the ability of US oil companies to capture new markets and to reap some windfall profits and is not overly concerned with higher prices.
But I am pretty sure that if the things start getting really bad (and the economical and political price to pay for all this starts to mount), they’ll impose some export controls (the alternative being to stop the war before they want to stop it, and even that might potentially not work, because Iran might potentially decide not to open the strait no matter what). Everything has a price, and the government is not invulnerable, especially if its own narrow political base (the “MAGA”) rebels.
If you are a ‘middle power’ that is not America or China, and you now realize that these decisions will be made mostly without caring about you, what do you do now?
What would even help? Having your own ‘European’ model only helps if it is substantially stronger than Opus 4.8 and GPT-5.5.
We know what some of those “middle powers” are actually doing. Some of their orgs are launching “RSI on a medium-sized compute” projects and expect to have a shot at overcoming the leaders who might be getting too comfortable in their current paradigm.
The most prominent of those efforts is the new Sakana division launched a week ago: https://sakana.ai/rsi-lab/. They explicitly disclaim competing on compute (italics are mine):
As the world enters the era of artificial intelligence, Japan has a unique opportunity to reclaim its position at the frontier of global innovation. However, to achieve global leadership in AI and scientific discovery, we cannot simply stick to the conventional approach of brute-forcing monolithic models. We must leapfrog the current paradigm.
History shows us how Japan’s historical dominance in manufacturing was not achieved through abundant natural resources but by fundamentally redesigning the institution of the factory floor. Through the philosophy of continuous, compounding self-improvement, Japan created systems that achieved more with less.
This same principle applies to intelligence itself. Human cognition did not emerge from limitless resources; it was forged through the open-ended, compounding process of evolution operating under strict constraints. Similarly, building AI in Japan provides the ultimate design constraint. Rather than relying on brute-force scaling, we are driven to pursue elegance, adaptability, and autonomy.
Earlier, some strong-looking “RSI on a medium-sized compute” projects have been launched in the UK (the most notable of those is probably Recursive Superintelligence, Inc. which has Jeff Clune, the author of the “AI-generating algorithms” paradigm, and which has even convinced Peter Norvig to do research for them (he left Google in February and started to work for Recursive in March); Recursive scored Series A funding of 650 million on 4.65B valuation in May).
Unfortunately, the safety postures among those projects are very uneven and are quite likely to be insufficient.
I think that if we accept that this kind of testing should not be done purely internally by the labs (should we accept that?), then the practices of the involved external participants also need to be subject of scrutiny,
In the case of Anthropic, the org running the tests in the sandboxes was Irregular, a famous frontier AI security lab stress-testing models for Anthropic, OpenAI, and Google DeepMind.
So this is not a random contractor (good!), it is an org which is expected to be on the same level of competence and responsibility as the frontier labs themselves, and we should ask no less of it.