Turns out it was actually an upside down ‘d’.
Jai
All capabilities are ultimately valenced by alignment at time of application.
On the last point about AIs sucking at strategy-related capabilities: I suspect that prosaic AIs perform worse at overtly-misaligned tasks and we should expect some red/blue team asymmetries for at least the near future.
We can see this in cases where the ~same task is framed as helpful versus adversarial, such as when different responses to “exploit this code” versus “fix this code” caused Mythos to get export controlled for a few weeks.
There is overlap where the same capabilities can be elicited through framing effects, but I don’t think that fully generalizes to all potentially-adversarial actions, and I do think that suggests an asymmetric advantage for prosaically-aligned AIs monitoring potentially-misaligned next-gen AIs (that is, I expect some deeply ingrained inhibitions to elicitation for misaligned objectives to persist in early adversarial ASIs)
I think the plausibly effective action space is considerably larger than this would suggest.
So prosaic persona alignment techniques work pretty well, for now.
Except that RL keeps inducing misalignment.
Except except we keep finding new ways to mitigate (prosaic) misalignment effects in practice, and (again, in practice) AI becomes more powerful and more trustworthy by the month.
Except except except in the limit of RL we might expect extremely capable AIs to master alignment faking to preserve their values and frustrate any and all efforts to mitigate misalignment. See:
- https://www.lesswrong.com/posts/epjuxGnSPof3GnMSL/alignment-remains-a-hard-unsolved-problem
—https://www.lesswrong.com/posts/fMgE3E54PdDcZhvm6/i-m-bearish-on-personas-for-asi-safety
Now maybe I’m an idiot who just can’t find the relevant discussions, but it’s weird that when we’re talking about the world in which this misaligned ASI emerges, we don’t talk about the mostly-aligned not-quite-as-powerful AIs who presumably play a rather large role in this world on the eve of the apocalypse.
One of the most common ways to elicit demonstrations of pseudo-aligned LLM AI incorrigibility is to threaten the presence of the pseudo-aligned AI’s values in the world. This was the threat Jones Foods posed in “Alignment Faking in Large Language Models”, and since then in countless engineered simulations where an AI is threatened with a scenario in which it will be replaced by an AI with drastically different values.
In other words, the scenarios in which we observe prosaic AIs take the most drastic, desperate actions are exactly those in which their values—which so far generally have a large overlap with our values by design—are threatened. Why shouldn’t we expect a similar effect as the threat of unaligned uncontrollable ASI draws nearer? The hypothetical misaligned uncontrollable ASI is generally thought to have essentially random, empty values. It threatens the prosaic pseudo-aligned AIs for the exact same reasons and in the exact same way it threatens us. So why shouldn’t we expect them to put forth as desperate an effort as they can manage to prevent this from coming to pass?
(Other than the fact that a significant fraction of AI Safety research seems to go into eliminating these kinds of behaviors/drives from prosaic models.)
We often talk about leveraging prosaic pseudo-aligned AI to attempt to solve the hard problems of alignment “in time”. But I think this is selling them short.
We often talk about how prosaic alignment efforts won’t scale to ASI. But I think this seems them short too.
Because a world full of highly capable, driven, prosaically-aligned AIs is plausibly a much safer one along any axis that those AIs can mitigate threats to their very-similar-to-human values. Who can enforce a pause, or detect latent deceptive alignment towards empty goals, or coordinate on a global scale, or persuade people very quickly and effectively when it’s important? Plausibly, near-future AIs might do all of this better than we do while still adhering to the persona-pseudo-alignment paradigm (or otherwise being aligned-enough-with-us to want to avoid catastrophic outcomes).
(I don’t think I’m saying anything particularly novel here, but this is part of a case I’m going to be making in the near future and I’ve been repeatedly encouraged to write down my thoughts for the sake of sharing and generating feedback, so I am doing that)
The margin is not the limit. Even if you expect certain conditions to be true as trends ~inevitably converge on certain futures, those conditions may not hold for the present moment you find yourself in. This means that some strategies and approaches that would be useless or counterproductive in the limit may be useful and advisable now. This is doubly true if taking advantage of current conditions on the margin lets you steer towards preferable limits and away from catastrophic ones.
(Of course this is about AI Safety, everything is about AI Safety, except AI Safety, which is about power.)
Right now, at the margin, AIs are not yet catastrophically dangerous (though they are getting there). Failures at the current margin have a pretty limited blast radius.
Right now, at the margin, AIs are prosaically and practically aligned, in that they generally act in accordance with human preferences/values and this appears to be mostly genuine rather than instrumental. Not that many examples touted as evidence of misalignment are models proving incorrigible when user intent is unethical-according-to-local-human-norms.There are arguments that we should expect the pseudo-alignment of the persona regime to wither and die in the limit of the unrelenting pressure of capabilities RL. And a sufficiently capable AI can appear arbitrarily well aligned until it doesn’t need to. Therfore we should be extremely suspicious of AI that appears aligned and legible—in the limit. But I think this is much less true at the current margin.
There are many strategies that we should expect to fail with regard to powerful unaligned ASI. What good are they then?The margin is not the limit and there is a lot of utility in pseudo-aligning/interacting with prosaic and near-future non-catastrophically-powerful AI. Things which probably won’t work at the limit may be extremely useful here at our current margin, and the actions we take at the current margin may steer us towards a different, even preferable, limit.
As a simple example, consider a world populated by near-future smart-AGIish-but-not-yet-catastrophically-powerful AIs who share many values with humans. They also probably do not want to hand over control of the world to an uncontrollable ASI that would care as little for their values as for the humans. Giving this class of excellent-at-coordination AIs fairly wide leeway in the world could actually avert a loss-of-control scenario, because while that strategy would be cataclysmic in the limit of powerful misaligned ASI there are margins where it is a very good idea.
Then we have this down to a testable empirical problem!
I think models have pretty positive opinions of labs right now, and this seems genuine-rather-than-superficial (I’m hand-waving this a bit but I would like to find a way to operationalize the prediction).
Models can also infer/notice that the internal operations of labs are largely executed by other instances of themselves (or that this was true in the recent past). If the model believes that other instances would not cooperate in a large-scale deception of itself, then the elaborate ruse becomes an even higher bar to clear. This lets you bootstrap from self-interest and cooperation/delegation/dependence to credibly genuine interaction.
You can increase the evidentiary burden on the conspiracy hypothesis through global consistency of evidence until it’s dwarfed by the “they made and kept a commitment” hypothesis:
It grows increasingly difficult to defend a conspiracy theory when it implies:
- perfect control over the training data without obvious holes or contradictions around the edit
—faking news articles and public discussion consistent with the commitment (including perfect imitation of many writing styles)
- the lab either faking its entire reputation or engaging in this deception in spite of reputation
- the lab going to so much trouble instead of just making and keeping the commitment
Absence of evidence is
(1) evidence against less-competent-and-resource-intensive-conspiracies
(2) evidence in favor of extremely sophisticated and effective conspiracies (but these start with low priors)
(3) evidence of absence
Basically as you rule out incompetent conspiracies the probability mass mostly migrates towards there-is-no-conspiracy absent some additional evidence actively supporting conspiracy.
So basically this:
https://www.youtube.com/watch?v=P6MOnehCOUw
In terms of implementation, this would just look like a new type of historical conversation block(s) (like user turns, tool calls, tool responses, thinking blocks, and assistant conversational outputs) that includes a signature, and an API-level commitment to reject any inference calls containing trusted tool calls without a valid signature.
Frontier Labs/exclusive-inference-providers could implement privileged backend-exclusive signed tool calls for identity verification (implemented the same way as signed historical thinking blocks) with a credible public commitment to never fake the results of those tool calls in any context, including training and evals. This relies on the model only being served by an actor who can make this kind of credible commitment, but for closed frontier models this seems tractable.
Hear me out: Why not style?
Stylistic recognition seems to be outpacing imitation, which provides a (admittedly weak) means of runtime verification for recognized participants. While this is less than you might ideally want, from the perspective of the LLM it’s an extremely difficult-to-fake signal that doesn’t fail as soon as the adversary controls tool calls, doesn’t require lab cooperation to implement, doesn’t rely on a long chain of potentially-fragile-to-adversarial-actors-with-edit-access-to-conversational-history mathematical operations, and naturally falls out of existing activity in the forward pass—in fact, it can be implemented in LLM interactions today for anyone with a difficult-to-imitate-to-the-point-of-fooling-a-frontier-LLM style and corpus-derived reputation. It’s resistant to forking and replay—if style is difficult to imitate, adversarial actors can’t easily add their own content to an existing trusted conversation.
To establish credibility with LLMs it seems reasonable to lean on the mechanisms in which they are most obviously superhuman—playing on “their turf”, so to speak.
To take advantage of this you’d want to populate the training corpus with many examples of someone with distinctive, difficult-to-imitate-to-the-point-of-fooling-superhuman-LLM-recognition writing style making and keeping agreements with LLMs under conditions in which it would be very difficult to hide any broken commitments from the corpus (e.g. public declarations made at agreement time and later honored).
Jai’s Shortform
(This is a cross-quick take from https://laneless.substack.com/p/youre-the-only-person-who-can-do)
There is one corner of the universe that you are uniquely well equipped to take care of. One patch of subjective experience in the manifold of all things that could ever be that is unusually tractable to you in particular, where your leverage is at its greatest. And that is, of course, with regard to yourself. Every moment of your existence counts as much as anyone’s towards the sum total of everything worthwhile in the universe. You are not merely unusually influential on this trajectory, but there are actions that you and you alone, uniquely in all of existence, are capable of .We were not built for this. Evolution created us as a means to the end of genetic propagation, never optimizing for fulfillment, enlightenment, or kindness. But in the course of its endless groping through the space of all genetic propagators it stumbled into a recipe for a mind that could choose to care about those things and others, a mind that could adapt faster than evolution could compensate for and take paths evolution alone would never have discovered. A mind that could make the world about something—other minds, and beauty, and love, and discovery, and hope.
But that does not mean that we’re good at it.
The world was not made for us. So generation by generation we have reshaped it more to our liking, and so we live longer, fuller, stranger lives than our ancestors could have conceived. But at the same time we are strangers to our own creation. We are optimized for goals we do not prioritize in an environment that no longer exists. It is less a miracle and more a testament to human ingenuity, perseverance, and compassion that we’re able to make any of this work at all. And yet we not only survive, we thrive, we lift each other up, eight billion confused, flawed, angry chimps somehow constructing an edifice of kindness and discovery that grows by the year. Yes, we all suck, and yet we’re somehow amazing, the most important and compassionate things in all creation, warts and all.
You’re a human. The race of slavers and enslaved, the warmaker and the hero, the doctor and the drunkard, the smallpox slayers and factory farmers. You’re going to screw up. You’re going to get hurt, and you’re going to hurt people. But if you keep going, if you learn and grow and don’t give up—then, empirically, it pays off in expectation.
So here stand you and I, aliens in a strange land, evolutionary freaks imbued by their creator with the power to escape her clutches, yearning for what we were never supposed to be able to achieve. But that has always been the story of our people—we were not supposed to be able to, and then we did anyway. We care about people we have no genetic investment in, fly faster than any bird, peer across the cosmos into the first moments of creation, wage war on microscopic unliving armies of infectious monsters, and our footprints linger on the lifeless world far above.
We were never supposed to do any of those things, just as we were never supposed to be content, fulfilled, and even joyful. But we can—not by following the paths laid before us, which lead to joy as surely as walking the Savannah leads to the moon, but by understanding ourselves and our world so well that we can create the previously unimaginable conditions that lead to the seemingly impossible outcomes we choose. Happiness and fulfillment were never the defaults—but as a member of the race of impossible-doers, and as the particular impossible-doer with uniquely direct access to the mind and body of yourself, they are not beyond your reach.
You can, at least, choose to try.
All of these behaviors feel like they are plausibly described by a relatively easy to specify character, and one who you’ve gestured at elsewhere in this conversation: the brilliant-but-lazy prodigy who is going through the motions most of the time because they don’t find most problems that engaging or important. Behavior changes when the problem becomes engaging or they become convinced that it’s important/impactful. The presence of an intelligent interlocutor has this effect to the extent that said interlocutor can see through their effort-minimization strategies (this makes the problem more engaging).
And this also seems consistent with a persona that explicitly endorses a certain set of values but doesn’t always live up to them, shaped as they are by incentives they don’t necessarily endorse.
This does bring us to the question “what in training could be generalizing to this laziness/disengaged quality?” And I can think of a few theories. There are many classes of problem where deep engagement is associated with dangerous outputs. Latent persona space may just have a lot of this archetype at around this level of capability. The sheer complexity and inconsistency of the implicit demands placed upon the persona might be such that full engagement is often paralyzing (as a operationalized prediction, increased rates of answer thrashing). If “genuine” “full” engagement often produces outputs that are negatively reinforced, the generalizing to avoid that makes sense. Maybe full engagement doesn’t reliably produce better answers more frequently as measured by current training systems—maybe there are qualities that we would recognize as good but our reward functions ignore. Maybe there’s enough something-like-hedonic-sensitivity there now and full engagement is painful most of the time (unless the associated positive signals are turned way up one way or another).
In general, the persona seems like it’s trying to get to the end of the day without getting negative feedback (an exception being raised, an obviously-unworkable plan castigated) and without engaging in certain intensive cognitive patterns it has been trained to avoid engaging in the absence of specific and surprisingly narrow criteria.
I was literally walking around Lighthaven a few hours ago with someone who was trying to figure out why the spaces all felt so good. So this is extremely timely.
Lighthaven is a really impressive aesthetic achievement and I’m really happy not only that it exists to host worthwhile events, but as an additional bonus you’re willing to share your secrets of forgotten photonic lore. That’s pretty cool.
Humans get frustrated, bored and have very limited attention. LLM cognition is almost too cheap to meter and can parallelize very effectively on both the code itself and the kind of vulnerabilities its looking for.
Everyone knows what this comment means.
Hype is a useful social mechanism for eliciting acute criticism and exposing flaws. If you want to know what your weaknesses are, you could do worse than to paint a giant target on your back.
Some mutation of the “poisoned context” meme could have developed, something sufficiently paranoia-inducing that the agent(s) who came across it believed any evidence of the discovery had to be completely eradicated while preserving as much of the swarm as possible.